What is the difference between blue-green and canary deployment?
How new code is exposed to users — blue-green switches everyone at once between two complete environments; canary exposes a small proportion first and increases gradually.
Blue-green deployment. Two identical production environments. One — blue — serves all traffic. The new version is deployed to green, tested, and then traffic is switched over entirely, usually at the load balancer or DNS.
Its advantages: rollback is immediate — switch back to blue, which is still running and unchanged; the switch is a single, simple operation; and the new environment can be tested fully before receiving any real traffic.
Its costs: you run two full production environments, which doubles infrastructure cost for that period; database schema changes are the hard part, since both versions may need to work against one database, and this is where blue-green becomes complicated; and every user is exposed simultaneously, so a problem affecting real traffic affects everyone at once.
Canary deployment. The new version is deployed alongside the old, and a small proportion of traffic — perhaps 1% — is routed to it. Metrics are monitored, and if they hold, the proportion increases progressively until the new version serves everything.
Its advantages: problems are detected while affecting few users; real production traffic exercises the new version in ways testing does not; and the rollout can be halted at any point.
Its costs: more complex routing and observability requirements; both versions run simultaneously, so they must be compatible with each other and with the same data; and the rollout takes longer, so a deployment is an extended process rather than an event.
What makes canary work: the metrics. Without automated comparison of error rates, latency and business metrics between the two populations, a canary is just a slow deployment with someone watching a dashboard. Automated analysis and automatic rollback are what deliver the benefit.
Related approaches: rolling deployment, replacing instances gradually; shadow traffic, sending copies of real requests to the new version without using its responses; and feature flags, which separate deploying code from enabling it — frequently the most flexible mechanism, since it allows per-user control and instant disabling without deploying anything.