Topics
DevOps & Delivery

Blue-Green and Canary Deployments

Blue-green switches all traffic at once, canary shifts it in steps judged by metrics. How each rolls back, and why schema changes must expand, then contract.

Intermediate·13 min read·Updated Oct 6, 2026

Most outages start with a change, so the way a new version reaches users decides how much a bad release can hurt. Blue-green runs the new version next to the old one and flips all traffic in one switch at the load balancer, so rollback is flipping it back. Canary sends a small, growing share of traffic to the new version and promotes it step by step only while its error rate and latency match the old one. Both depend on the database accepting the old and new code at the same time, which is why schema changes are split into expand and contract steps.

Context

For a long time a release meant stopping the application, copying the new build over the old one and starting it again, usually in a maintenance window at night. Blue-green deployment came out of ThoughtWorks projects in the mid-2000s and was popularised by Jez Humble and David Farley's book Continuous Delivery and Martin Fowler's 2010 write-up. Canary releases are named after the canaries miners carried to detect gas: a few users meet the new version first. Large companies automated them in the 2010s; Netflix and Google open-sourced Kayenta, an automated canary judge, in 2018, and Kubernetes controllers such as Flagger (2018) and Argo Rollouts (2019) made progressive delivery a configuration file.

You have used the simplest version already: a Kubernetes Deployment replaces pods gradually by default (a rolling update with up to 25% extra and 25% unavailable pods), and Vercel lets each production deploy roll out to a percentage of traffic with Rolling Releases (generally available since June 2025). The core of a canary is one weighted route:

canary-route.yaml
# Istio VirtualService: 95% stays on v1, 5% tries v2
http:
  - route:
      - destination: {host: checkout, subset: v1}
        weight: 95
      - destination: {host: checkout, subset: v2}
        weight: 5
Deploy vs release
Deploy puts new code on servers; release exposes it to users. Blue-green, canary and feature flags all separate the two.
Blue / green
Two complete production environments. One (say blue) serves traffic, the other (green) receives the new version, then they swap roles.
Canary
A small share of traffic routed to the new version while the rest stays on the stable one.
Baseline
In canary analysis, a fresh deployment of the old version of the same size, so the comparison is new vs old under equal conditions.
Rollback
Returning traffic to the previous version. How fast it is depends on whether that version is still running.
Expand / contract
Changing a schema in backward-compatible steps: add the new structure, migrate, and only later remove the old one.

Why it matters

Google's SRE (site reliability engineering) book estimates that roughly 70% of outages are caused by changes to a live system. A strategy that limits how many users see a bad change, and how quickly it can be undone, turns most of those outages into a blip on a dashboard. The hard parts are not the traffic switch, which any load balancer can do, but everything that is shared between versions: the database schema, cached data, queued messages and user sessions. A rollback that the database cannot follow is not a rollback.

The deployment strategies side by side

Every strategy answers the same three questions: how many users can a bad version reach, how long does rollback take, and how many servers do you pay for during the release. Recreate and rolling updates replace the old version, so rolling back means deploying the old version again. Blue-green and canary keep the old version running, so rolling back is a routing change.

StrategyHow it shipsRollbackCost / risk
RecreateStop v1, start v2Redeploy v1, with downtimeDowntime on every release
RollingReplace instances in batchesRoll forward or redeploy v1, minutesv1 and v2 mixed for the whole rollout
Blue-greenFull v2 next to v1, switch 100%Switch back, secondsDouble capacity; every user hits v2 at once
CanaryShift 1% → 5% → 25% → 100%Set weight to 0, secondsNeeds traffic splitting and good metrics
Feature flagCode shipped dark, enabled per userTurn the flag offFlag debt; both code paths in production

Blue-green: one switch, both ways

The new version is deployed to the idle environment, warmed up and tested there with real infrastructure but no users. The release is a single change at the routing layer: the load balancer target group, the Kubernetes Service selector, or the ingress points at green instead of blue. Blue keeps running untouched, so if green misbehaves, the same switch in the other direction restores the old version within seconds.

before switch
blue v1 · 100%green v2 · 0%smoke tests on greenv2 deployed, no users yet
after switch
green v2 · 100%blue v1 · idlekeep blue for the rollback window
rollback
blue v1 · 100%green v2 · 0%one routing change
Blue-green release and rollback. The new version is fully deployed and tested before it gets traffic; the old one stays warm until you are confident, so rollback is a routing change, not a redeploy.

Where the switch happens matters

Switching at a load balancer or service mesh takes effect for the next request. Switching by changing a DNS (Domain Name System) record does not: resolvers and clients cache the old address for at least the record's TTL (time to live) and sometimes much longer, so traffic drains over minutes or hours and a rollback is just as slow. Existing keep-alive and WebSocket connections also stay on the old environment until they close, which is why the old side must drain connections rather than be stopped the moment the switch happens.

Canary: shift traffic, measure, promote

A canary rollout raises the new version's share of traffic in steps and pauses after each one. During the pause, an analysis compares the canary's error rate, latency and key business metrics against a baseline running the old version. If the canary is within limits, the next step starts; if not, the weight goes back to 0 and the rollout is aborted, usually without anyone being paged.

5%25%50%100%promoted10 min10 min10 mindoneeach pause: compare canary vs baselinecanary shareafter each step:errors ≤ baseline?p99 ≤ baseline?any check fails:weight → 0%
A canary rollout over about 40 minutes. Each step holds the weight while metrics are compared with the stable version; any failed check sends the weight straight back to 0%.

Three details separate a real canary from a slow rolling update. Assignment is sticky: the router hashes a user ID or cookie, so a user does not bounce between versions on every request. The comparison is against a baseline of the old version started at the same time, not the long-running stable fleet, whose warm caches and older pods skew the numbers. And the decision is automated, with thresholds written down before the rollout, because a person staring at graphs promotes whatever looks roughly fine.

rollout.yaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: checkout
spec:
  replicas: 10
  strategy:
    canary:
      stableService: checkout-stable
      canaryService: checkout-canary
      trafficRouting:
        istio:
          virtualService: {name: checkout}
      analysis:                  # runs in the background from step 1
        templates: [{templateName: error-rate}]
        startingStep: 1
      steps:
        - setWeight: 5
        - pause: {duration: 10m}
        - setWeight: 25
        - pause: {duration: 10m}
        - setWeight: 50
        - pause: {duration: 10m}
  # selector and pod template omitted
analysis-template.yaml
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: error-rate
spec:
  metrics:
    - name: error-rate
      interval: 1m
      failureLimit: 2            # 3rd failed check aborts the rollout
      successCondition: result[0] < 0.01
      provider:
        prometheus:
          address: http://prometheus.monitoring:9090
          query: |
            sum(rate(http_requests_total{app="checkout",
              track="canary", code=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{app="checkout",
              track="canary"}[5m]))

The same pattern ships with other tools: Flagger for Kubernetes with Istio, Linkerd or NGINX, AWS CodeDeploy's canary and linear configurations for Lambda and ECS (Elastic Container Service), and Vercel Rolling Releases for frontends. At 1% of traffic a low-volume service may see a handful of requests per minute, which proves nothing, so small services need longer pauses or larger first steps.

The database and other shared state

During any of these rollouts, and again during a rollback, two versions of the code run against the same database at the same time. A migration that renames a column breaks whichever version expects the other name. The fix is to make every schema change backward compatible by splitting it across releases: first expand (add the new structure without removing anything), then move code and data over, then contract (remove the old structure) only when no running version uses it.

  1. 1
    Release N, expand: add the nullable full_name column. Old code ignores it, so this ships safely on its own.
  2. 2
    Release N+1: write both name and full_name, read full_name with a fallback to name. Rolling back to N still works because name is kept current.
  3. 3
    Backfill: copy old rows in small batches, so no single statement holds locks for minutes.
  4. 4
    Release N+2: read only full_name, keep writing both until the rollback window for N+1 has passed.
  5. 5
    Release N+3, contract: stop writing name, then drop the column. Only now is the old shape gone.
rename-column.sql
-- release N, expand: additive only, the running code keeps working
ALTER TABLE users ADD COLUMN full_name text;

-- backfill in batches; repeat until 0 rows are updated
UPDATE users SET full_name = name
WHERE id IN (
  SELECT id FROM users
  WHERE full_name IS NULL
  LIMIT 5000
);

-- release N+3, contract: no deployed version reads "name" any more
ALTER TABLE users DROP COLUMN name;

Sessions, caches and queues

The same rule applies to everything else both versions touch. Sessions kept in process memory vanish when traffic moves to the other environment, so they belong in a shared store or a signed token. A cache entry written by v2 in a new shape will be read by v1 after a rollback, so include a version in cache keys or keep the format compatible. Messages that v2 publishes may be consumed by v1 workers, so message schemas must be readable in both directions, and background workers and cron jobs need their own rollout, because they do not sit behind the load balancer the canary controls.

Pitfalls

  • A migration that the old version cannot survive

    Renaming or dropping a column in the same release as the code that needs it works only if nothing ever runs the old code again. During a canary, a rolling update or a rollback, it does, and it fails on every query that touches the column. Split the change into expand and contract releases.

  • Switching blue-green through DNS

    DNS answers are cached by resolvers, operating systems and clients, often beyond the TTL, so traffic moves gradually and unpredictably and rollback is just as slow. Switch at a load balancer, ingress or mesh you control, and keep DNS pointing at it.

  • Judging a canary on too little traffic

    Ten requests with zero errors is not evidence. With low traffic, a 1% canary passes every check by luck and the real test happens at 100%. Size steps and pauses so each step sees enough requests to detect the error rate you care about, or add synthetic traffic.

  • Comparing against the stable fleet instead of a baseline

    Long-running pods have warm caches and JIT (just-in-time) compiled code, while new pods start cold, so a perfectly good canary can look slower than stable. Compare it with a baseline of the old version started at the same moment and at the same size.

  • Forgetting the parts outside the load balancer

    Queue consumers, scheduled jobs and webhook receivers that read from a broker are not covered by traffic weights. A bad worker version processes 100% of jobs from the first minute. Roll them out separately and watch their own error metrics.

Interview questions

Q1When would you choose blue-green over canary?

Blue-green when you need an all-at-once switch with instant rollback and can afford double capacity during the release, for example for a version that cannot run mixed with the old one, or for a service with too little traffic for canary statistics to mean anything. Canary when you have enough traffic and good metrics, because it limits how many users meet a bad build instead of exposing everyone and relying on a fast rollback.

Q2Walk me through implementing a canary release for a stateless API on Kubernetes.

I would use a progressive delivery controller such as Argo Rollouts or Flagger with a mesh or ingress that supports weighted routing. The rollout defines steps, for example 5%, 25%, 50% with 10-minute pauses, and an analysis template that queries Prometheus for the canary's 5xx rate and p99 latency compared with a baseline. Routing is sticky by user hash, the thresholds are set before the release, and a failed check sets the weight to zero automatically. Schema changes ship separately using expand and contract.

Q3What happens when a release includes a database migration?

During the rollout and any rollback, old and new code share the database, so the migration must work with both. I make it additive first: add columns or tables, deploy code that writes both shapes, backfill, switch reads, and only remove the old structure in a later release once nothing running needs it. A destructive migration bundled with the code change makes rollback impossible.

Q4How do you decide automatically whether a canary is healthy?

By comparing it with a baseline of the old version on a few metrics with thresholds agreed in advance: error rate, tail latency, saturation such as CPU or memory, and one or two business metrics like successful checkouts. Each check runs repeatedly during the pause, and a small number of failures aborts the rollout. The step has to carry enough traffic for the difference to be statistically meaningful.

Q5What happens to sessions and in-flight requests during a blue-green switch?

New requests go to green immediately, while requests already being served by blue should finish, so blue drains connections instead of being stopped at the switch. Long-lived connections such as WebSockets stay on blue until they close or are told to reconnect. Sessions survive only if they live in a shared store or a signed token; in-memory sessions are lost.

Q6Why does a rolling update not give you instant rollback?

Because it replaces the old instances as it goes, so by the end there is nothing left to route back to. Rolling back means rolling out the old version again, which takes as long as the release did, while the bad version keeps serving. It also leaves both versions mixed for the whole rollout without any control over which users see which.

Q7How are feature flags different from a canary?

A canary controls which build serves a request; a flag controls which code path runs inside the same build. Canaries catch bad builds, such as crashes, leaks or performance regressions, while flags let you release or hide a feature per user, independently of deploys, and switch it off without shipping anything. Most teams use both, and clean up flags once a feature is fully on.

Key takeaways
  • Most incidents start with a change; a deployment strategy limits how many users a bad one reaches and how fast you can undo it.
  • Blue-green flips 100% of traffic between two full environments: rollback in seconds, at the cost of double capacity and a 100% blast radius.
  • Canary raises traffic in steps, compares the new version with a baseline on agreed metrics, and drops to 0% automatically when a check fails.
  • Switch at a load balancer or mesh, not DNS, and drain old connections instead of killing them.
  • Old and new code always share the database during a rollout or rollback: expand first, contract releases later.
  • Caches, queues, sessions and background workers must tolerate both versions too; traffic weights do not cover workers.

Preparing for interviews? DevRecall turns a job description into a prep plan that points at topics like this one.

Start free