Safe Deployment and Rollback Strategies Questions
Releasing changes to production safely and incrementally, and recovering when they fail: blue-green, canary, and rolling deployments, feature flags, dark launches, traffic shifting, and progressive rollout, together with rollback strategies, safe-deploy practices, blast-radius containment, automated recovery, and safe forward/backward migration. Covers deployment orchestration across cloud platforms, staged exposure of new behavior to users, assessing deployment risk, designing reversible releases, and restoring a known-good state quickly. Focuses on how a release reaches production and how it is unwound on failure, distinct from broader incident command, which lives in Enterprise Operations & Incident Management.
A rollback just occurred and thousands of clients reconnect or replay requests simultaneously, causing a secondary overload (a 'rollback thundering herd'). How do you throttle reconnections and apply backpressure to avoid it?
Sample Answer
Direct answer
A rollback thundering herd happens when a rollback (or the incident that triggered it) causes many clients to disconnect and reconnect, or retry requests, all at once, and the fix is to make reconnection and retry behavior deliberately staggered and rate-limited on BOTH the client and server sides, rather than letting every client react to the same event at the exact same instant.
Structured elaboration
- Client-side jitter: reconnection and retry logic should never fire at a fixed, synchronized delay after a disconnect; adding a random jitter window (say, a random delay between 0 and several seconds before reconnecting, rather than "reconnect immediately") spreads what would otherwise be a single sharp spike of reconnection attempts across a wider time window.
- Exponential backoff on retries: a client whose request failed should retry with an increasing delay (and its own jitter) rather than immediately retrying, so a wave of failed requests doesn't immediately become a wave of retried requests at the same intensity.
- Server-side rate limiting and backpressure: even with well-behaved clients, the server-side needs its own protection: rate-limiting new connection attempts per unit time, and applying backpressure (explicitly rejecting or queueing excess load with a clear signal, like a
503with aRetry-Afterheader, rather than silently accepting more connections than it can actually serve and degrading everyone). - Coordinating clients and servers: a
Retry-Afterheader (or an equivalent explicit signal) on a rejected request tells well-behaved clients specifically how long to wait, giving the server some influence over the retry timing rather than purely hoping client-side jitter alone is sufficient. - Graceful degradation as a last resort: if load genuinely exceeds capacity despite the above, shedding load deliberately (serving a degraded but fast response, or explicitly rejecting the lowest-priority traffic first) is better than every request queuing and timing out simultaneously, which tends to make the overload WORSE by having clients hold connections open waiting rather than failing fast and retrying later.
Worked example
A rollback causes 50,000 clients to lose their persistent connection simultaneously. Without jitter, all 50,000 attempt to reconnect within the same one-second window, overwhelming the now-recovering service before it can stabilize. With client-side jitter spreading reconnection attempts across a 30-second window, combined with server-side rate limiting on new connection acceptance and a Retry-After-based backpressure signal for anything beyond that rate, the same 50,000 reconnections complete successfully over that 30-second window instead of causing a second outage on top of the first.
Trade-offs and pitfalls
Jitter and backoff trade a small amount of individual-client reconnection latency (a client might wait a few extra seconds versus reconnecting instantly) for overall system stability, a trade that's almost always worth it during exactly the moment (right after a rollback) when the system is least able to absorb a sudden spike. The common mistake is only implementing client-side jitter without server-side rate limiting/backpressure as a backstop, which works fine if every client is well-behaved but leaves no protection against a client that isn't (a bug, an unusually aggressive retry implementation, or simply enough clients that even well-jittered retries still exceed capacity in aggregate).
What makes a database migration backward-compatible? Give an example of a safe and an unsafe schema change, and explain why backward compatibility matters for rollback and phased deployment.
Sample Answer
Direct answer
A backward-compatible migration is one where the OLD version of your application code still works correctly against the NEW schema. That matters because during a rolling or canary deployment, old and new code run against the same database simultaneously, so if the new schema breaks the old code, you have an outage the moment the rollout starts, before you've even finished deploying.
Structured elaboration
- Safe (backward-compatible) changes: adding a new nullable column, adding a new table, adding a new index, widening a column's type (e.g. int to bigint in most databases). Old code that doesn't know about the new column simply ignores it; nothing it does breaks.
- Unsafe (non-backward-compatible) changes: renaming or dropping a column the old code still reads or writes, adding a NOT NULL column with no default (old code's INSERT statements, which don't set that column, start failing), changing a column's type in an incompatible direction (bigint to int, string to enum), or adding a foreign-key constraint on data the old code might still write in a way that violates it.
- Why it matters for rollback specifically: if you deploy a schema change together with new code and then need to roll the CODE back, the old code needs to keep working against the schema as it now stands, which is exactly the backward-compatibility property. If the migration wasn't backward-compatible, rolling back the code doesn't fully undo the outage, because the schema is still in its new, incompatible state.
Worked example
Safe: adding a nullable discount_code column to an orders table. Old code that doesn't reference it keeps working exactly as before. Unsafe: renaming orders.total to orders.total_amount in the same deploy as the code change that uses the new name; if you need to roll back the code while the rename has already run, the rolled-back old code tries to read orders.total, which no longer exists, and every read fails.
Trade-offs and pitfalls
Backward-compatible migrations usually take more steps and more calendar time (add, backfill, switch, THEN remove the old column in a later, separate deploy) than a direct rename, which is the trade-off teams are making: more process for a rollback safety net. The common mistake is treating "the migration ran successfully" as the same thing as "the migration is safe," when the real test is whether the PREVIOUS version of the application still functions correctly against the new schema.
Services A and B were updated together. A's update is backward-compatible, but B's new version introduced incompatible writes, and you must roll back B while keeping A on its new version. How do you handle in-flight and already-persisted inconsistent state so the system reaches eventual consistency?
Sample Answer
Direct answer
When B must roll back but A stays on its new (backward-compatible) version, the core problem is that B's incompatible writes may have already happened before you detect the issue, so the fix isn't just "redeploy old B," it's identifying and correcting whatever inconsistent state those writes left behind, while B's rollback itself is straightforward since A's compatibility means A doesn't need to change at all.
Structured elaboration
- Stop further damage: roll back B's code immediately (this part is simple exactly because A is backward-compatible with B's old version, no coordination needed there), which stops NEW incompatible writes from happening, but doesn't undo ones that already occurred.
- Identify what's actually inconsistent: determine which writes B made in its brief new-version window were the problematic, incompatible ones, versus which writes (even during that window) were fine; this typically requires either an audit log of what changed, or a way to distinguish "written by new B" from "written by old B" (a version marker on the data itself, if you have one, or inferring from timestamps against the deploy window).
- In-flight requests specifically: any request that started against new-B logic but hasn't yet completed when the rollback happens needs explicit handling, either let it finish against the code it started with (avoiding a mid-request logic switch) and then reconcile its result afterward if needed, or, if that's not safe, actively cancel/retry it against the now-rolled-back old B.
- Reconciliation toward eventual consistency: for the identified inconsistent writes, either a compensating action (a corrective write that brings the data back to what old-B's logic would have produced) or, if the incompatible writes are numerous and hard to individually correct, a broader reconciliation job that recomputes affected records from source data, run once B is confirmed stable on the old version again.
- Idempotency throughout: both the rollback itself and any compensating/reconciliation actions need to be safely re-runnable, since this kind of recovery work is exactly the scenario where a retry (from a nervous on-call engineer, or from an automated retry mechanism) is likely.
Worked example
Service B's new version wrote records in a new, incompatible format for roughly 8 minutes before the regression was caught and B rolled back. An audit log (or a version-tagged field on the written records) identifies exactly which records fall in that window; a reconciliation job re-derives the correct, old-format value for each of those specific records from upstream source data, run idempotently (safe to re-trigger if it's interrupted partway through) so a retry doesn't double-apply the correction.
Trade-offs and pitfalls
The scope of this problem is directly proportional to how LONG the incompatible version was live before detection, which is a strong argument for the fast, automated rollback-trigger discipline covered elsewhere in this topic; a regression caught in 30 seconds via automated detection leaves a much smaller reconciliation problem than one caught 30 minutes later via a human noticing something looked off. The common mistake is treating the code rollback as the end of the incident, when the DATA reconciliation is often the harder, longer-running part of actually resolving it.
What signals commonly trigger an automated rollback, and how would you group them into categories? How long should you wait before triggering, and what techniques prevent oscillation between versions?
Sample Answer
Direct answer
Automated rollback is usually triggered by four signal classes: health-check failures (readiness/liveness probes failing repeatedly), an error-rate spike above baseline, a latency regression on a key percentile, or a failed post-deploy smoke test. The trigger doesn't fire the instant a signal crosses a threshold; it waits out a short observation window and requires the breach to persist, otherwise a normal traffic blip rolls back a perfectly good release.
Structured elaboration
- Health checks: a readiness probe failing on a fraction of new pods for more than a couple of consecutive checks is the fastest, cheapest signal and should gate before traffic even ramps up.
- Error rate: compare the new version's error rate against the stable version's, not against a fixed historical number, since traffic mix shifts by time of day.
- Latency: p95 or p99, not the mean, since a mean can hide a fat tail that half the users are actually feeling.
- Failed smoke tests: a synthetic transaction (place a test order, hit a canary endpoint) that fails is often the highest-confidence signal because it directly exercises the new code path.
- How long to wait: a grace period of one to a few observation windows (for example three consecutive one-minute buckets) before triggering, so a single noisy scrape doesn't fire the gate. The window should scale with traffic volume: a low-traffic service needs a longer window to get a statistically meaningful sample.
- Preventing oscillation (flapping): once a rollback fires, hold at the previous version for a cooldown period before allowing another promotion attempt, and require the SAME signal to clear for a similar dwell time before re-promoting, not just a single clean reading.
Worked example
A service normally runs 0.2% errors. During a canary, three consecutive one-minute windows show 1.8%, 2.1%, and 1.6% errors on the new version while the stable version stays at 0.2% for the same windows. That's a sustained, comparative regression across three windows, not a single spike, so the gate fires. Contrast that with one window at 2% followed by two windows back at 0.2%: that's noise, and a gate that fires on the first window alone would have rolled back a healthy release.
Trade-offs and pitfalls
A grace period that's too short creates false rollbacks and erodes trust in the automation (teams start ignoring or disabling it); one that's too long lets a bad release run in production longer than necessary. The common mistake is comparing the new version against a fixed absolute threshold instead of against the stable version's live baseline, which makes the gate miss regressions during a genuinely bad traffic day and over-trigger during a genuinely good one.
What is a rolling update, and how does it differ from a recreate deployment? For a stateless Kubernetes service, what does the rollout process look like, and what commonly goes wrong during it?
Sample Answer
Direct answer
A rolling update replaces old-version instances with new-version ones gradually, a few at a time, so the service stays available with a mix of old and new versions running simultaneously during the transition. A recreate deployment, by contrast, terminates ALL old instances first and only then starts the new ones, which means a period of full downtime but avoids ever running mixed versions.
Structured elaboration
For a stateless microservice in Kubernetes, a rolling update:
- Kubernetes creates a batch of new-version pods (controlled by
maxSurge, how many extra pods above the target replica count are allowed). - Waits for those new pods to pass their readiness probe before routing traffic to them.
- Terminates an equivalent batch of old-version pods (controlled by
maxUnavailable, how many pods can be down at once). - Repeats until all pods are on the new version.
Common failure modes to watch for during a rollout:
- Readiness probe misconfigured too loosely: traffic gets routed to a pod that's technically "ready" but not actually able to serve correctly yet (cache not warmed, connection pool not established).
- Version skew during the mixed-version window: old and new pods running simultaneously both talk to the same downstream dependencies (shared database, shared cache), so if the new version isn't backward-compatible with what the old version expects, you get intermittent failures purely from which version happened to handle a given request.
- Resource exhaustion from surge: if
maxSurgeallows too many extra pods at once relative to available cluster capacity, new pods can fail to schedule, stalling the rollout partway. - A bad new version rolling out gradually still affects a growing fraction of traffic before anyone notices, unlike blue-green where the bad version is fully isolated until an explicit cutover.
Worked example
A 20-replica deployment with maxSurge: 25% and maxUnavailable: 25% creates up to 5 extra pods (25 total temporarily) while taking down up to 5 old pods at a time, cycling through until all 20 are on the new version. If the new version has a subtle bug that only manifests under a specific downstream response, roughly a quarter of traffic is exposed to it at any point mid-rollout, growing toward 100% as the rollout proceeds, unless something halts it.
Trade-offs and pitfalls
Rolling update avoids downtime and extra infrastructure cost (no duplicate fleet, unlike blue-green) but accepts a mixed-version window where compatibility between old and new has to hold, and it doesn't isolate a bad release the way canary or blue-green does; it just gradually replaces capacity regardless of whether the new version is actually healthy, UNLESS combined with a readiness-probe-based or metrics-based halt condition.
Unlock Full Question Bank
Get access to all Safe Deployment and Rollback Strategies interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.