Safe Deployment and Rollback Strategies Questions
Releasing changes to production safely and incrementally, and recovering when they fail: blue-green, canary, and rolling deployments, feature flags, dark launches, traffic shifting, and progressive rollout, together with rollback strategies, safe-deploy practices, blast-radius containment, automated recovery, and safe forward/backward migration. Covers deployment orchestration across cloud platforms, staged exposure of new behavior to users, assessing deployment risk, designing reversible releases, and restoring a known-good state quickly. Focuses on how a release reaches production and how it is unwound on failure, distinct from broader incident command, which lives in Enterprise Operations & Incident Management.
You're on a canary rollout at 5% traffic when p95 latency rises 1.5x while the error rate stays flat. Walk through your diagnostic steps in order, and how you'd decide whether to continue, pause, or roll back.
Sample Answer
Direct answer
A 1.5x latency increase with a flat error rate during a canary is exactly the ambiguous case automated canary analysis is built for: it's not an obvious failure (nothing's erroring) but it's also not obviously fine, so the right move is a structured diagnostic pass, not an immediate gut call either direction.
Structured elaboration
Step-by-step, in order:
- Confirm it's real, not a sample-size artifact: check the canary's request count for this window; 5% traffic might be a small enough sample that a couple of slow requests skew the p95/p99 without it being a genuine, broad regression.
- Check WHICH percentile moved: did the mean move, or specifically the tail (p99)? A tail-only shift suggests a subset of requests hitting a slow path (a cold cache, a specific input shape), while a broad shift across all percentiles suggests something more systemic.
- Compare against infrastructure-level signals: CPU/memory on the canary pods specifically, are they resource-constrained relative to stable? A canary running on fewer instances than the stable fleet can look "slower" purely from having less capacity per request, not from a code regression.
- Trace a slow request: pull a distributed trace for one of the slow requests and see WHERE the extra time is going; is it in the new code path itself, or in a downstream dependency call that both versions share (which would point away from the deploy as the cause)?
- Check logs for the canary specifically: any new warnings, retries, or timeout patterns that correlate with the deploy?
- Decide: continue if the increase traces to a benign, expected cause (e.g., cold cache that's now warming) and the trend is improving; pause and gather more data if the cause isn't yet clear and you have time to wait; roll back if the trace points to a genuine regression in the new code or if the trend is worsening rather than stabilizing.
Worked example
Tracing a slow canary request shows the extra ~150ms is spent in a downstream inventory-service call that BOTH old and new code make identically, ruling out the new code as the cause; checking resource metrics shows the canary pods are running at higher CPU utilization simply because the canary slice has fewer replicas than its 5% traffic share would proportionally need. In this case: continue the rollout (the latency increase traces to an infrastructure sizing artifact of the canary itself, not a code regression), but flag the canary-sizing mismatch as something to fix before the NEXT canary run.
Trade-offs and pitfalls
The temptation under a flat error rate is to assume "no errors means it's fine," but latency regressions are real user-experience problems even without a single error logged, so treating error rate as the only signal that matters is a common and costly mistake. Equally, panicking and rolling back on the FIRST ambiguous signal without doing the diagnostic work wastes the whole point of running a canary, which is to gather enough information to make a confident call rather than a reflexive one.
You run a globally distributed service behind a global load balancer. Design a canary that limits blast radius to a single region while preserving user session affinity and supporting cross-region failover.
Sample Answer
Direct answer
Limiting a canary's blast radius to a single region behind a global load balancer means routing based on BOTH region AND canary assignment together, so users in the target region get split between canary and stable while every other region stays entirely on stable, with session affinity handled so a user doesn't flip between versions mid-session, and a cross-region failover path that doesn't accidentally expose the canary to a region it was never meant to reach.
Structured elaboration
- Region-scoped canary: configure the global load balancer's routing so only requests already destined for the target region are further split by the canary weighting; requests to every other region bypass the canary logic entirely and go straight to stable, keeping the blast radius genuinely contained to one region's traffic.
- Session affinity: within the target region, use a stable hash of the user's identity (not a random per-request choice) to decide canary-vs-stable, so once a user lands on the canary, they consistently stay there for the DURATION of their session rather than flip-flopping between versions on each request, which would both confuse metrics and give users an inconsistent experience.
- Cross-region failover: if the target region fails over to another region (a genuine regional outage, unrelated to the canary itself), the failover target region needs to know NOT to apply the canary split, since the canary was only meant to affect that one specific region's traffic; failing over should route everyone, including the canary cohort, to STABLE in the failover-target region, rather than accidentally expanding canary exposure to a region it was never validated in.
- Metrics scoped to the region: canary-vs-stable comparison metrics need to be filtered to the target region specifically, since aggregating in metrics from unaffected regions (which are 100% on stable) would dilute or distort the comparison.
Worked example
flowchart TB
GLB[Global Load Balancer] -->|region=US-target| Split[Canary/Stable split, 10/90]
GLB -->|region=EU| Stable_EU[100% stable]
GLB -->|region=APAC| Stable_APAC[100% stable]
Split --> Canary_US[Canary, US only]
Split --> Stable_US[Stable, US]
Canary_US -->|failover| Stable_EU
A user in the target region hashed into the canary cohort stays on canary consistently across their session (via the stable-hash session affinity); if that region experiences an unrelated outage and traffic fails over to the EU region, the failover path routes explicitly to EU's STABLE tier, not attempting to preserve the canary assignment across a region boundary it was never validated for.
Trade-offs and pitfalls
The specific risk this design guards against is a REGIONAL FAILOVER accidentally becoming a canary-exposure EXPANSION, silently putting canary-cohort users onto a fresh region where the canary was never tested against that region's specific infrastructure, traffic patterns, or configuration; the common mistake is a failover mechanism built independently of the canary logic that doesn't know to override the canary assignment during a cross-region failover event.
Design a progressive-delivery ramp for a payment service: an initial 1% canary, ramp to 50% over two hours if clean, then 100% after 24 hours. What automation and metric checks run at each stage, and how do you handle a partial rollback if problems appear at the 50% stage?
Sample Answer
Direct answer
A progressive-delivery ramp for a payment service needs the automation to actively gate each stage's advance on real metric checks, not just wait out a timer, and the partial-rollback plan for the 50% stage needs to distinguish cleanly between requests that already went through the new code (which may have real side effects, like a payment already processed) and requests still ahead of the rollback taking effect.
Structured elaboration
- 1% canary: the smallest, most cautious stage, watched closely with a shorter observation window since the blast radius is tiny; metric checks focus on error rate and latency deltas against the stable baseline, plus a payment-specific correctness signal (successful-transaction rate, any reconciliation mismatch) since a payment service's most dangerous bugs may not show up as a raw HTTP error at all.
- Ramp to 50% over two hours if clean: this isn't a single jump, it's itself a staged ramp (say 1% -> 10% -> 25% -> 50%, each requiring its own clean metric window before advancing), automated so a human doesn't have to manually approve every micro-step, but with metric checks gating EVERY step, not just the final 50% checkpoint.
- 100% after 24 hours: a long hold at 50% specifically to accumulate enough transaction volume and TIME (payment issues can be slow-building, like a subtle reconciliation drift that only shows up after a batch settlement process runs) before committing to full exposure.
- Partial rollback if problems appear at 50%: reduce the new version's traffic share back down (not necessarily to zero immediately, potentially stepping back to a smaller, still-nonzero percentage to keep gathering diagnostic data on a contained population while you investigate), while the ALREADY-PROCESSED transactions on the new code path need their own review: were any payments processed incorrectly, and do they need a compensating action (a reversal, a manual reconciliation) distinct from the traffic-routing rollback itself?
Worked example
At the 50% stage, an automated check flags a reconciliation discrepancy in a batch of transactions processed by the new code. The traffic-routing rollback (scaling the new version's share back to 5%, not necessarily zero, to preserve some live diagnostic signal) happens within minutes via the automated pipeline. Separately and on a different timeline, a manual reconciliation process reviews every transaction that went through the new code path during its exposure window to determine whether any need a compensating correction, since simply routing future traffic away doesn't undo whatever the already-processed transactions did.
Trade-offs and pitfalls
Payment services are the canonical example of where "roll back the traffic" and "the problem is fixed" are NOT the same thing, since money may have already moved; the automation needs to be scoped clearly to what it CAN fix (stop MORE transactions from hitting the bad path) while explicitly flagging what it can't (undo transactions that already happened), which needs a human-driven reconciliation process rather than being folded into the automated rollback itself.
Define deployment frequency, mean time to recovery, and change-failure-rate, the DORA-style metrics used to gauge deployment health and velocity. How would a team measure each, and what's a reasonable target for a high-performing team?
Sample Answer
Direct answer
Deployment frequency (how often you ship to production), mean time to recovery (how fast you restore service after an incident), and change-failure-rate (what fraction of deployments cause a production problem) are three of the four DORA metrics that together describe how fast AND how safely a team ships. High-performing teams deploy often, recover fast, and fail rarely; the point of tracking all three together is that any one alone can be gamed or misleading.
Structured elaboration
- Deployment frequency: count of production deploys per day/week per service. Elite teams deploy on-demand, often multiple times a day; measuring it is usually just counting CI/CD pipeline "deploy to prod" events.
- Mean time to recovery (MTTR): from the moment an incident starts degrading users to the moment service is restored. Measuring it accurately requires two reliable timestamps: incident START (usually from the first alert or the first bad metric) and RESOLVED (usually from the alert clearing or an explicit "resolved" marker), which is harder to instrument well than it sounds, since teams often only record when the fix was DEPLOYED, not when the SERVICE actually recovered.
- Change-failure-rate: percentage of deployments that require a rollback, hotfix, or cause an incident, out of total deployments. Requires tagging deployments with an outcome, which usually means linking your deploy log to your incident/rollback log.
- Reasonable targets (per DORA's own research bands): elite performers deploy on-demand (multiple times per day), recover in under an hour, and keep change-failure-rate under 15%. Teams earlier in their DevOps maturity might deploy weekly to monthly, take a day or more to recover, and see failure rates well above that.
Worked example
A team ships 12 times a week, has 2 of those deploys cause an incident requiring rollback (change-failure-rate ~17%), and those 2 incidents took 25 and 40 minutes respectively to resolve (MTTR ~33 minutes). That profile is roughly "high" performing on frequency and recovery speed but borderline on failure rate, suggesting the team's canary/testing discipline needs tightening before pushing frequency even higher.
Trade-offs and pitfalls
Optimizing deployment frequency alone, without watching change-failure-rate, just means shipping more bugs faster; the metrics are meant to be read together. The most common measurement pitfall is MTTR calculated from "code fix deployed" instead of "users stopped being affected," which systematically understates real recovery time whenever a rollback or mitigation restores service before the actual fix ships.
Write a script that polls a service's health endpoint for a few minutes after a deploy, aggregates success rate and average latency, and marks the deployment FAILED if success rate or latency crosses a threshold.
Sample Answer
Direct answer
A post-deploy health-poll script watches the new version for a fixed window right after it starts receiving traffic, aggregating success rate and latency, and explicitly marks the deployment failed (triggering whatever the pipeline's next step is, typically an automatic rollback) if either crosses a threshold, rather than assuming silence means success.
Structured elaboration and worked example (logic verified)
import time
import requests
def poll_health(url: str, duration_seconds: int = 300, interval_seconds: int = 5,
success_rate_threshold: float = 0.99, latency_threshold_ms: float = 500):
# Polls a health endpoint for `duration_seconds`, aggregating results.
# Returns (passed: bool, summary: dict).
start = time.monotonic()
successes = 0
total = 0
latencies = []
while time.monotonic() - start < duration_seconds:
req_start = time.monotonic()
try:
resp = requests.get(url, timeout=5)
elapsed_ms = (time.monotonic() - req_start) * 1000
latencies.append(elapsed_ms)
total += 1
if resp.status_code == 200:
successes += 1
except requests.RequestException:
total += 1
latencies.append(latency_threshold_ms * 10) # count a hard failure as very slow, not silently dropped
time.sleep(interval_seconds)
if total == 0:
return False, {"reason": "no requests completed, cannot assess health"}
success_rate = successes / total
avg_latency = sum(latencies) / len(latencies)
passed = success_rate >= success_rate_threshold and avg_latency <= latency_threshold_ms
return passed, {"success_rate": success_rate, "avg_latency_ms": avg_latency, "requests": total}
Key structural decisions: a network exception is counted as a failed, slow request rather than silently skipped, since a connection timeout is itself a strong negative signal that shouldn't be excluded from the aggregate just because it didn't return an HTTP status code at all. Zero completed requests is treated as an explicit failure with a clear reason, never silently passed, since "we don't have enough data" should never be mistaken for "it's healthy."
Trade-offs and pitfalls
A fixed polling interval (5 seconds here) trades responsiveness for load on the health endpoint; a service under real stress from a bad deploy doesn't need to be hammered by an aggressive health-check loop on top of everything else. The retry/timeout handling matters more than it looks: a script that raises an unhandled exception on the FIRST network hiccup, rather than counting it as a data point and continuing, produces a false "script crashed" result instead of the actually useful "service is unhealthy" signal the pipeline needs to act on.
Unlock Full Question Bank
Get access to all Safe Deployment and Rollback Strategies interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.