Verified Alibaba Cloud account Ultimate Alibaba Cloud Manual
0. Welcome to the Ultimate Alibaba Cloud Manual (Where Money Meets Electricity)
Let’s be honest: “cloud” is a wonderful word that makes people forget the basics. It’s like saying “we’ll handle lunch” while opening the fridge and discovering a single carrot with commitment issues. Alibaba Cloud is powerful, feature-rich, and capable of hosting everything from a small website to enterprise-scale systems. But power comes with options, and options come with the possibility of clicking “Confirm” on something you probably didn’t mean to buy.
This manual is designed to help you build real things on Alibaba Cloud while staying sane, secure, and aware of cost. It’s not a “click blindly and hope” guide. Instead, it’s a “learn the shape of the platform, pick sensible defaults, and build a repeatable workflow” guide. Think of it as a friendly co-pilot who also says, “Hey, maybe you want to enable backups,” before you drift into the rocky canyon of accidental public exposure.
We’ll cover the essentials: account setup, identity and access management, regions and networking, compute, storage, databases, load balancing, container options, monitoring, logging, backups, and cost control. You’ll also get a practical set of checklists and troubleshooting tips that mirror what real users do when the cloud behaves like the cloud—mysterious, but not malicious.
1. What You Actually Need Before You Touch Anything
Before deploying anything, you need three things: a goal, a rough architecture, and a plan for security and costs. If you don’t have these, the cloud will gladly provide all three… in the form of surprises.
1.1 Define your goal (the part where we avoid building a cathedral for a sandwich)
Start with questions like:
- What are you building? (website, API, app backend, data pipeline, batch jobs)
- Who will use it? (internal users, public internet, region-specific audience)
- How much traffic or data? (rough estimates are fine)
- How important is uptime? (minutes of downtime vs “we can’t sleep if it fails”)
If you’re unsure, build a simple baseline: a small compute instance (or container) with basic storage, then layer in scaling and resilience later. The cloud is excellent at scaling—especially when you’re prepared to pay for it.
1.2 Draw a simple architecture (boxes and arrows, not modern art)
At minimum, think in layers:
- Client → Load Balancer (optional) → Compute (ECS/containers)
- Compute → Database (RDS/PolarDB/other)
- Compute → Object Storage / NAS / Disk storage
- Observability: metrics, logs, alerts
- Security: VPC, subnets, security groups, IAM roles
Don’t worry about making it perfect. You’re building a map, not a spaceship.
1.3 Plan security and cost early (because “later” is where bugs go to hide)
In cloud land, security and cost are not afterthoughts. They are structural elements.
- Security: ensure access is controlled, networks are private by default where possible, and secrets are handled safely.
- Cost: choose the right service type, right instance sizes, proper shutdown/termination policies, and monitor spend.
2. Account Setup and the “Please Don’t Make a Mess” Security Pass
Alibaba Cloud gives you an administrative console and a set of services. Before using resources, spend a few minutes setting up access properly. This avoids the classic scenario: “We accidentally shared admin access with someone who can approve refunds and also decided to click ‘Delete All’ because it looked lonely.”
2.1 Set up your Alibaba Cloud account
Typical steps include verifying your identity and setting up billing. Choose the appropriate billing method and ensure payment details are accurate. If you’re working in a team, consider using multiple accounts (or at least multiple roles) rather than giving everyone the keys to the kingdom.
2.2 Use RAM (Resource Access Management) concepts properly
On Alibaba Cloud, RAM is the usual approach for controlling permissions. Instead of granting broad admin access to everyone, create roles or users with least-privilege permissions.
Common best practices:
- Create separate accounts for admins and operators.
- Grant permissions per service and action scope.
- Use MFA (multi-factor authentication) for human accounts.
2.3 Lock down network access with intent
Most cloud security incidents are boring: open ports, overly permissive security groups, or credentials left in environment variables without rotation. Your goal is to:
- Restrict inbound access to only the required IP ranges.
- Use security groups (or equivalent) as the firewall layer.
- Prefer private subnets for databases.
In short: don’t put your database on the public highway unless you enjoy security audits.
3. Regions, Availability Zones, and the Geography of Regret
Alibaba Cloud resources live in regions and within those regions, often in multiple availability zones. Choose carefully, because latency and availability depend on these decisions.
3.1 Pick a region based on your users
Ask: where are your users? Choose the region closest to them to reduce latency. If you must serve multiple geographies, you may use global acceleration services or multi-region strategies later.
3.2 Use multiple zones for resilience
For critical workloads, distribute across zones. When services support it, use load balancers across instances in different zones. For databases, consider replication and failover options.
Cloud failures happen. The goal is not “never fail,” it’s “fail safely and recover quickly.”
4. Networking: VPC, Subnets, and How to Avoid Turning Your App into a Public Park
Networking is the skeleton of your deployment. If networking is wrong, everything else becomes complicated. If networking is right, everything else becomes merely annoying rather than catastrophic.
4.1 Create a VPC (Virtual Private Cloud)
In most real deployments, you’ll create a VPC for isolating your resources. A VPC provides network segmentation so your instances can talk privately.
Verified Alibaba Cloud account Key concepts:
- VPC CIDR block: your address range
- Subnets: subdivisions of the VPC, often mapped to availability zones
- Routing: how traffic flows between subnets and external networks
Verified Alibaba Cloud account 4.2 Subnet planning: public vs private
Common pattern:
- Public subnet: for load balancers, bastion hosts, and anything that must accept inbound internet traffic (with strict security rules)
- Private subnet: for application instances and databases
Use NAT or gateway patterns if private instances need outbound internet access for updates and dependencies.
4.3 Security groups: your firewall with a sense of humor (but obey it)
Security groups define allowed traffic rules. You generally attach security groups to instances and allow only required ports/protocols.
Examples of rules you might need:
- Allow inbound HTTP/HTTPS to web/app ports
- Allow inbound database port only from application security group
- Deny everything else by default
If you’re wondering whether you need to open a port, you probably don’t. Add ports only when you have a specific requirement.
4.4 DNS and domain mapping
If you have a domain, you’ll map it to your load balancer or gateway. Plan for certificate management as well. If you use HTTPS, you’ll want a certificate and a strategy for renewal.
5. Compute Options: ECS, Containers, and the Art of Choosing the Right Tool
Alibaba Cloud provides multiple compute paths. The main ones people discuss are ECS (Elastic Compute Service) for VMs, and container-based approaches for orchestrated deployments.
5.1 ECS instances: when you want control and simplicity
ECS is a good starting point for:
- Small web apps
- Prototypes
- When you need VM-level customization
Typical workflow:
- Create instance in your VPC/subnet
- Verified Alibaba Cloud account Attach security group rules
- Configure OS, firewall, and application
- Set up system monitoring
- Consider auto-scaling if traffic grows
Keep an eye on instance lifecycle. Stopping vs deleting varies by platform, and bills have very consistent memories.
5.2 Containers: when you want portability and easier scaling
Containers are useful when you want consistent environments across dev/staging/production. If your application is already containerized, it can reduce the pain of “it worked on my machine” by making machines all agree to be machines.
Container orchestration options depend on Alibaba Cloud services. The general approach is:
- Package your application in a container image
- Run containers via an orchestrator
- Use service discovery/load balancing
- Configure scaling policies
Verified Alibaba Cloud account If you’re new, start with the simplest approach that matches your goals. Don’t build Kubernetes when you only needed a bicycle.
5.3 Scaling: vertical now, horizontal later (or the reverse, depending on your life choices)
Scaling strategies:
- Vertical scaling: larger instance sizes
- Horizontal scaling: more instances
- Auto-scaling: automatic instance changes based on metrics
Horizontal scaling is usually more resilient. But it also requires stateless application patterns and careful database design.
6. Storage: Disks, Object Storage, and the “Please Back Up Your Stuff” Gospel
Storage is where your data lives, and data is where your future self will either thank you or curse you. Choose correctly and back up everything important.
6.1 Block storage (disks) for OS and system data
ECS typically uses block storage for:
- Operating system files
- Application working directories
Plan disk sizing for growth. Also consider snapshot policies if you need recovery points.
6.2 Object Storage for files and artifacts
Object Storage is great for:
- User uploads
- Static assets like images and videos
- Backups and exported data
- Build artifacts (like deployment bundles)
With object storage, think about lifecycle policies (move older data to cheaper storage), versioning (if supported), and access controls.
6.3 NAS or shared file storage (use carefully)
Shared file storage can help with certain workloads, but it can also create performance bottlenecks or operational complexity. If your application can use object storage for most needs, you often simplify your life.
6.4 Backups: because disasters love unprepared systems
Set backup policies for both compute-related data and databases. Backups should be:
- Automated
- Encrypted (where possible)
- Tested (yes, actually restore sometimes)
- Retained for a sensible period
Backups that you’ve never restored are like a parachute that was only folded, never jumped with.
7. Databases: Choosing the Right One and Not Falling Into the “We’ll Change Later” Trap
Database choices influence performance, cost, and complexity. Alibaba Cloud offers managed database options, which reduce operational overhead compared to self-managed servers.
7.1 Relational databases for transactional workloads
If you need SQL and transactions, relational managed services are typically the path. They provide:
- Automated patching
- Backups and snapshots
- Replication options
- Better scaling workflows
Typical best practice: place database instances in private subnets and restrict inbound access to only application instances.
7.2 Connection and performance planning
Databases dislike sudden connection storms. If you run a high-traffic app with many instances, configure connection pooling and limit concurrency. Also:
- Set sane timeouts
- Use indexes appropriately
- Monitor slow queries
If you do these early, you’ll avoid the classic “why is production so slow?” meeting where everyone pretends to be surprised.
7.3 Migrations and schema changes
Plan migrations as part of your deployment pipeline. Use tools that support safe migration strategies (like zero-downtime approaches where possible). Always test migrations in staging with realistic data volumes if you can.
8. Load Balancing and Traffic Management: Keep It Smooth, Keep It Protected
Load balancing makes your application resilient and scalable. It also centralizes traffic management, which helps with security and monitoring.
8.1 When to use a load balancer
Use one when:
- You have multiple instances
- You need health checks
- You want a consistent entry point for routing
- You need TLS termination
If you have a single instance and you’re experimenting, you might skip it temporarily. But for production readiness, a load balancer is usually worth it.
8.2 Health checks: let the cloud babysit your instances
Configure health checks that match your application’s real readiness. If health checks only verify that a port is open, you might route traffic to a half-dead app.
8.3 TLS/HTTPS configuration
Plan how certificates are managed. Options include:
- Termination at load balancer
- End-to-end encryption depending on your security requirements
- Automated renewal if supported
Verified Alibaba Cloud account Make sure your application behaves correctly when using forwarded headers if TLS terminates upstream.
9. Observability: Monitoring, Logging, and Alerting Without Losing Your Mind
Observability turns your cloud from a black box into a transparent fish tank. You can see what’s happening, why it’s happening, and when it’s about to happen again.
9.1 Monitoring: metrics that matter
Track metrics like:
- CPU, memory, disk usage
- Network in/out and latency
- Application-level metrics (requests, errors, response times)
- Verified Alibaba Cloud account Database metrics (connections, query time, locks)
Define alarms for thresholds that represent real issues. Too many alarms create alarm fatigue, which is like crying “wolf” so often that nobody notices the real wolf (which, frankly, should be illegal).
9.2 Logging: centralize logs and search them
Store application logs and system logs in a centralized system if possible. Then ensure you can:
- Search by time range
- Filter by service or instance
- View errors and exceptions quickly
Log correlation matters. If you use request IDs, include them so debugging becomes less like solving a mystery with missing clues.
9.3 Tracing (optional, but delightful)
For microservices, distributed tracing helps you track a request across services. If you’re not there yet, start with logs and metrics, then evolve.
10. Automation: The Secret Sauce That Stops You From Repeating Yourself
Manual deployments are fine until they aren’t. Once you deploy twice, automation starts paying for itself. Your goal is repeatable infrastructure and repeatable application deployments.
10.1 Use templates or infrastructure-as-code where possible
Infrastructure-as-code lets you define your resources in a versioned way. That means:
- New environments (dev/staging/prod) are consistent
- Changes are reviewed like code
- Rollbacks are more manageable
You can also reduce “it worked when I clicked it” syndrome.
10.2 CI/CD pipelines
Set up a pipeline that builds, tests, and deploys your application. Best practices include:
- Run tests before deployment
- Use environment-specific configuration
- Deploy gradually when feasible
- Automate rollbacks on failure
If you’re using containers, CI/CD naturally integrates with image builds and deployment updates.
10.3 Secrets management
Do not hardcode credentials in code or store them casually in plain text files. Use secret storage mechanisms appropriate to your platform.
Rotate secrets periodically and whenever you suspect exposure. Also, treat IAM keys like valuable glass: don’t share them on postcards.
11. Cost Control: The Art of Not Buying Your Own Volcano
Cloud costs can creep up like a cat on a countertop. You don’t notice until suddenly there’s more cost than expected and everyone is asking “Did we accidentally leave something running?”
11.1 Use right-sizing
Start with conservative instance sizes, then scale up after you observe real usage. Over-provisioning is a common early mistake.
11.2 Turn off what you don’t need
For non-production environments, set schedules or termination policies. Temporary test instances should be ephemeral.
11.3 Monitor spending with alerts
Enable billing alerts. Keep an eye on costs by service so you learn your spending patterns.
11.4 Choose storage lifecycle policies
For object storage, apply lifecycle rules to move older objects to cheaper tiers. For snapshots, ensure retention is reasonable.
11.5 Understand data transfer costs
Verified Alibaba Cloud account Network egress can be significant. If you serve content frequently, consider CDN-like approaches (if available and appropriate) and cache aggressively.
12. Operational Readiness: Being Brave on Day 2
Day 1 deployments are exciting. Day 2 operations is where you either become a calm hero or a panicked juggler. Let’s make you the hero.
12.1 Incident response basics
Create a simple runbook for common issues:
- Service downtime: check load balancer health, instance health, logs
- Database errors: check connections, slow queries, resource usage
- Performance degradation: compare metrics to baselines
- Verified Alibaba Cloud account Deployment failures: review CI/CD logs and rollback
Document steps and keep them accessible. If your runbook is trapped in a brain that hasn’t slept, it won’t help.
12.2 Backups verification
Regularly test restores in a safe environment. Confirm that you can:
- Verified Alibaba Cloud account Restore database backups
- Recover application data dependencies
- Validate integrity
12.3 Patch management
For ECS and any VM-based deployments, plan patch cycles. If using managed services for databases, patching is typically handled by the provider, but you still need to test compatibility.
13. Troubleshooting Playbook: When Things Go Wrong (Because They Will, It’s Only a Matter of Time)
Cloud troubleshooting is detective work. Here are practical patterns you can use regardless of the specific service.
13.1 Connectivity issues
If your app can’t reach the database or API endpoints:
- Confirm VPC/subnet routing is correct
- Check security group rules (inbound/outbound)
- Verify network ACLs if used
- Verified Alibaba Cloud account Confirm DNS resolution for internal endpoints
Most connectivity problems are configuration problems, not cosmic problems.
13.2 Instance health problems
If instances are unhealthy behind a load balancer:
- Check application logs
- Verify health check endpoint logic
- Confirm ports and firewall rules
- Check resource saturation (CPU/memory/disk)
Sometimes the app is “running” but doesn’t respond correctly. Health checks reveal truth.
13.3 Database slow queries or timeouts
If the database is slow:
- Look for missing indexes
- Check query execution plans
- Review connection pool settings
- Check lock contention and long-running transactions
Slow queries are often symptoms. Find the cause, not just the delay.
13.4 Unexpected bills
If you’re seeing unexpected spend:
- Check if instances are running when they shouldn’t
- Review autoscaling events
- Verified Alibaba Cloud account Inspect storage growth (especially snapshots and logs)
- Look for large network egress or external calls
Then add guardrails: alerts, budgets, and termination policies.
14. A Practical “First Production” Checklist
If you want a quick pass before you declare victory, use this checklist. It’s not exhaustive, but it covers the common landmines.
14.1 Identity and access
- MFA enabled for admins
- RAM roles created with least privilege
- Secrets stored securely and rotated
14.2 Network and security
- Resources placed in VPC and correct subnets
- Security groups allow only required ports
- Database access restricted to application instances
- No “0.0.0.0/0 to database port” situations
14.3 Reliability and data protection
- Backups enabled and restore tested
- Load balancer configured with health checks
- Multi-zone deployment for critical components when available
14.4 Observability
- Metrics collection enabled
- Central logs enabled
- Alerts configured for high-risk thresholds
Verified Alibaba Cloud account 14.5 Cost governance
- Budgets and spend alerts enabled
- Non-production instances scheduled or terminated
- Storage lifecycle policies applied
15. Example Deployment Paths (Pick One and Start Building)
Let’s give you a few common patterns. These are not “the only way,” but they’re proven shapes for many applications.
15.1 Pattern A: Simple web app on ECS with a managed database
- Create VPC, public subnet, and private subnet
- Deploy ECS instance in private subnet
- Use load balancer in public subnet
- Place managed database in private subnet
- Enable object storage for static assets/uploads
- Enable monitoring, logging, and backups
Great for early production and teams that want control without container complexity.
15.2 Pattern B: Containerized app with autoscaling
- Build container image via CI
- Deploy to container service with auto-scaling
- Use load balancer for traffic distribution
- Use managed database and object storage
- Set alerting for scaling and error rates
Great when you expect growth and want consistent runtime environments.
15.3 Pattern C: Data pipeline and batch workloads
- Use compute for scheduled jobs
- Store intermediate artifacts in object storage
- Write final results to database or data warehouse
- Monitor job runs and failure alerts
- Set retention policies for temporary data
Great for periodic processing and ETL-style workflows.
16. Frequently Asked Questions (With Answers That Don’t Just Shrug)
16.1 “How do I get started if I’m new?”
Start small. Choose one compute option (ECS or containers), deploy a simple web service, connect it to a managed database, and turn on monitoring and backups. Then iterate. Do not start by deploying ten microservices and three data lakes while wearing a blindfold. You’ll learn, but you’ll also pay.
16.2 “Should I use VPC from day one?”
In most cases, yes. Networking decisions affect security and deployment structure. Starting with VPC helps you avoid refactoring later.
16.3 “How do I prevent accidental exposure?”
Use security groups to restrict inbound access, place databases in private subnets, and avoid wide-open network rules. Also, use principle of least privilege in IAM. If you must expose something publicly, put it behind a load balancer and keep the attack surface minimal.
16.4 “What’s the fastest way to reduce costs?”
Right-size instances, stop unused resources, enable budgets and alerts, apply storage lifecycle policies, and watch network egress. Also, review logs and snapshots—those can grow quietly like indoor plants you forgot you watered.
17. The Ultimate “Don’t Regret It” Summary
The ultimate lesson of this manual is simple: Alibaba Cloud is not hard, but it is deep. You can succeed by making a few smart decisions early—security posture, networking structure, reliable deployment patterns, observability, and cost controls. When those fundamentals are in place, scaling and feature additions become manageable rather than chaotic.
Follow the checklist, use least privilege, keep your databases private, enable monitoring and alerting, and back up data with restore tests. Automation is your best friend. And when you feel tempted to “just quickly open the firewall,” pause and imagine future-you writing a calm postmortem titled “Why Was Everything on the Internet.”
Now go forth and deploy something useful. And if you accidentally create an extra resource, remember: the cloud is forgiving—until the bill arrives.

