What are cloud regions and availability zones?
A region is a geographic location containing cloud infrastructure; an availability zone is one or more physically separate data centres within that region, with independent power, cooling and networking. The distinction determines what kind of failure your architecture survives.
Why zones exist. A single data centre can fail entirely — power, cooling, fire, flood, or a fibre cut. Zones within a region are far enough apart that one failing should not take another with it, while being close enough for low-latency synchronous replication between them, usually a millisecond or two.
Why regions exist:
Latency. Physics sets a floor — light in fibre travels roughly 200 km per millisecond, so serving users from a nearby region is the only way to reduce round-trip time.
Data residency. Law frequently requires personal or regulated data to remain within a jurisdiction, which is now one of the main reasons regions are chosen.
Disaster recovery, against events affecting a whole region.
Availability and cost, both of which vary by region, sometimes substantially for the same service.
What this means architecturally:
Multi-zone is the normal baseline, and it is usually cheap — spreading instances across zones protects against the most common real failures.
Multi-region is expensive and complicated, because synchronous replication across regions is too slow, so you must choose between asynchronous replication with potential data loss and a design that tolerates eventual consistency. This is the trade-off people underestimate.
Cross-zone data transfer is usually charged, which surprises teams who spread services casually.
Not every service is regional. Some are global, some are zonal, and a zonal resource does not survive its zone.
What actually goes wrong in practice:
Control plane dependencies, where an outage prevents you launching replacement capacity even though your running instances are fine.
Correlated failures, where a software or configuration fault affects all zones simultaneously — zones protect against physical failure, not against a bad deploy.
Untested failover, which is the most common cause of a multi-zone design not helping.