Safe Deployment and Rollback Strategies Questions
Releasing changes to production safely and incrementally, and recovering when they fail: blue-green, canary, and rolling deployments, feature flags, dark launches, traffic shifting, and progressive rollout, together with rollback strategies, safe-deploy practices, blast-radius containment, automated recovery, and safe forward/backward migration. Covers deployment orchestration across cloud platforms, staged exposure of new behavior to users, assessing deployment risk, designing reversible releases, and restoring a known-good state quickly. Focuses on how a release reaches production and how it is unwound on failure, distinct from broader incident command, which lives in Enterprise Operations & Incident Management.
Compare blue-green, canary, and rolling deployments (and note where a plain recreate deployment still fits). For each, explain how traffic is shifted, the resulting rollback complexity, the infrastructure cost, and which kind of service (stateless vs. stateful) it suits best.
Sample Answer
Direct answer
Blue-green, canary, and rolling all reduce the risk of a bad release, but through different mechanisms: blue-green switches ALL traffic at once between two full environments, canary exposes a small SLICE of traffic to the new version before widening it, and rolling replaces instances gradually IN PLACE. A plain recreate deployment, by contrast, tears down the old version entirely before starting the new one, accepting downtime in exchange for simplicity.
Structured elaboration
| Strategy | Traffic shift | Rollback complexity | Infra cost | Best for |
|---|---|---|---|---|
| Blue-green | All-at-once, via LB/DNS switch | Low (switch back) | High (2x during overlap) | Stateless services needing near-instant rollback |
| Canary | Gradual, percentage-based | Low-medium (shrink canary slice) | Low-medium (small extra capacity) | High-traffic services where blast-radius control matters most |
| Rolling | Gradual, instance-by-instance in place | Medium (redeploy previous version, also gradual) | Low (no duplicate fleet) | Stateless services where some capacity reduction during rollout is acceptable |
| Recreate | All-at-once, old torn down first | Trivial (redeploy old version) but WITH downtime | Lowest | Low-traffic or maintenance-window-tolerant services |
Rollback complexity nuance: blue-green's rollback is fastest because the old environment never stopped running; canary and rolling both have to actively redeploy or re-route, which takes real time even if it's automated; recreate's "rollback" is simple mechanically but means accepting a second period of downtime.
Stateful services: all three of blue-green/canary/rolling get significantly harder with state (a database, in-memory session data, local disk), because you can't just duplicate or partially expose the data layer the way you can stateless compute; the deployment strategy for the STATELESS layer often decouples from a separate, more careful strategy for the DATA layer.
Worked example
A stateless API fronting a shared database: canary is a strong default, since it limits blast radius on the code change while the shared database (which doesn't get canaried the same way) stays constant underneath. Blue-green would be a better fit if the team's top priority is minimizing time-to-rollback over minimizing blast radius, since flipping back to the old environment is close to instant.
Trade-offs and pitfalls
There's no universally "best" strategy: the choice trades off blast radius, rollback speed, infrastructure cost, and operational complexity, and the right answer depends on which of those the specific service and change profile cares about most. A common mistake is picking a strategy based on what's trendy (everyone reaches for canary) rather than what the actual risk profile of the change calls for; a low-risk config change might not need any of this ceremony at all.
What are feature flags, and what different categories of flag exist based on who owns them and how long they're meant to live? What pitfalls tend to show up as flags accumulate over time?
Sample Answer
Direct answer
Feature flags are runtime switches that let you turn a piece of behavior on or off without a redeploy, which decouples DEPLOYING code from RELEASING a feature to users. The common types are a release flag (temporarily gates a feature during rollout, meant to be removed once it's fully live), an experiment flag (drives an A/B test, removed once the experiment concludes), an operational flag (a longer-lived control, like a rate limiter toggle or a maintenance-mode switch), and a kill-switch (an emergency, instant off-switch for something risky).
Structured elaboration
- Release flags: owned by the engineer shipping the feature; short-lived by design, meant to be deleted once the feature is fully rolled out and stable.
- Experiment flags: owned jointly by product/data science and engineering; lifetime tied to the experiment's duration, removed once a winner is chosen.
- Operational flags: owned by the team operating the service; can be long-lived (a genuinely permanent operational lever), which is the ONE type where "long-lived" isn't automatically a problem.
- Kill-switches: owned by whoever's on-call; meant to be exercised rarely but tested regularly so it's trustworthy when actually needed.
- Common pitfalls as flags accumulate: flag debt (release flags that were never cleaned up after the feature fully shipped, cluttering the codebase with dead branches), flag sprawl (so many flags that nobody can reason about which combinations of flag states are even possible, let alone tested), and stale defaults (a flag's fallback value drifts out of sync with what's actually safe as the surrounding code evolves).
Worked example
A team ships a new checkout flow behind a release flag, ramps it from 5% to 100% of users over two weeks while watching conversion metrics, then deletes the flag and the old code path once it's fully rolled out and stable for a monitoring period. If that deletion step gets skipped (a common failure under deadline pressure to move to the next feature), six months later the codebase has a flag nobody remembers the purpose of, defaulting to a value nobody's sure is still correct.
Trade-offs and pitfalls
Flags buy real safety (instant, redeploy-free control over risky behavior) at the cost of code complexity: every flag is effectively a branch that has to be reasoned about and eventually tested in both states. The single most common failure mode in practice isn't misusing a flag while it's active, it's forgetting to remove it once its job is done, which is why healthy flag programs bake in an explicit cleanup step (an expiration date, a dashboard of stale flags, or a policy that blocks new flags until old ones are retired) rather than relying on developers to remember.
Compare client-side and server-side feature-flag evaluation for a mobile app with intermittent connectivity. Which is safer for a rollout, and how do you define the default behavior when a flag can't be fetched?
Sample Answer
Direct answer
Server-side evaluation is generally safer for a rollout because the server can change a flag's value instantly and consistently for every client, while client-side evaluation depends on each device having up-to-date flag configuration, which is exactly what breaks down under intermittent connectivity.
Structured elaboration
- Client-side evaluation: the app itself decides whether a feature is on, usually based on a flag configuration it fetched and cached at some point. Faster to evaluate at request time (no network round-trip needed) and works fully offline, but a device on a stale cache keeps using an OLD flag value until it next successfully syncs, which for an intermittently-connected mobile app could be minutes to hours out of date.
- Server-side evaluation: the server decides on each request. Always reflects the current flag state instantly and consistently, but requires a live connection for every decision that depends on the flag, which is a problem for a mobile app trying to work offline or under a flaky connection.
- Default behavior when a flag can't be fetched: this is the crux of the safety question. The system needs an explicit, deliberately chosen default for "no data available," and that default should almost always be the SAFE, conservative behavior (feature OFF, or the old known-good code path), never "assume the last cached value is still correct" for anything risk-sensitive, since an unreachable device might be running a stale cache for an unknown, possibly long, period.
Worked example
A payment-related feature flag on a mobile app with client-side evaluation: the SDK is configured so that if it can't reach the flag service within a short timeout, it falls back to a hardcoded, safe default (feature OFF) rather than serving whatever was last cached, which might be hours or days stale if the device has been offline. This costs some feature-availability under poor connectivity (some users see the old behavior more often than strictly necessary) in exchange for never risking a stale, potentially-unsafe flag state driving payment logic.
Trade-offs and pitfalls
A hybrid is common in practice: client-side evaluation for speed and offline resilience, paired with a short cache TTL and a conservative, explicit fallback value baked into the SDK for when the cache is stale or fetch fails, rather than a pure binary choice between client-side and server-side. The common mistake is an SDK that silently serves an arbitrarily-old cached value with no TTL or fallback logic at all, which quietly turns "intermittent connectivity" into "unpredictable flag state," exactly the failure mode a well-designed flag system exists to prevent.
What is 'blast radius' in the context of a deployment, and what practical techniques reduce it: resource isolation, traffic controls, small-batch deploys?
Sample Answer
Direct answer
Blast radius is how much of your system, and how many users, are exposed to a bad deployment before you can stop it. Reducing it means never letting a single change reach 100% of traffic or 100% of your infrastructure in one step: you deploy to a small slice first, isolate that slice from the rest, and give yourself controls that can cut it off fast.
Structured elaboration
Techniques, roughly cheapest-to-hardest:
- Small-batch / percentage rollouts: canary a change to 1-5% of traffic or instances before going wider, so a bug affects a small fraction of users instead of everyone.
- Resource isolation: run the new version in separate compute (a distinct pod set, node pool, or availability zone) so a resource-exhaustion bug in the new version can't starve the old version's capacity too.
- Traffic controls: circuit breakers that stop routing to a demonstrably unhealthy instance, and rate limiters that cap how much load any single new component can absorb before it's proven stable.
- Region/cell isolation: for a global service, containing a rollout to one region or one "cell" of a sharded architecture means a bad release can't take down every region at once.
- Feature flags: decoupling "deployed" from "exposed" means you can turn a specific feature off instantly without a full redeploy, which is a much smaller and faster blast-radius-reduction lever than rolling back code.
For a monolith specifically, blast radius reduction is harder because there's no natural unit smaller than "the whole app": the levers become instance-level canarying (a subset of instances behind the load balancer run the new build) and feature flags around risky code paths, since you can't isolate one internal module's resource usage the way you can with a separate microservice.
Worked example
A change to a recommendation algorithm rolled out to 2% of traffic in one region first. A latency regression showed up only under that region's specific traffic mix (a caching quirk tied to timezone-driven request patterns); because it was contained to 2% of one region, the fix-and-redeploy cycle affected a small, recoverable slice of users instead of the whole global user base.
Trade-offs and pitfalls
More blast-radius controls mean more operational complexity and slower time-to-full-rollout, so teams calibrate the aggressiveness of containment to the risk of the change: a config tweak might skip straight to 100%, while a payment-logic change might go through five separate stages. The pitfall is applying the same heavy process to every change regardless of risk, which erodes the very safety discipline it's meant to protect by making people route around it under deadline pressure.
What's the difference between a rollback (redeploying the previous artifact) and a revert (a new forward commit that undoes the change)? Which would you reach for after discovering a production regression, and why?
Sample Answer
Direct answer
A rollback redeploys the previous, already-tested version of the artifact; a revert is a NEW forward commit that undoes the change in source control and then gets built and deployed like any other change. After discovering a production regression, rollback is almost always the faster, safer first move, since it restores a known-good state immediately, while a revert (even though it also "undoes" the change conceptually) still has to go through the normal build-and-deploy pipeline before it takes effect.
Structured elaboration
- Rollback: uses infrastructure/deployment tooling (redeploy the previous artifact,
kubectl rollout undo, switch a blue-green environment back) to restore the PREVIOUS RUNNING STATE directly, without rebuilding anything; it's fast precisely because the previous version is already built, tested, and known-good. - Revert: a source-control operation (
git revert) that creates a new commit undoing the change; this new commit then needs to go through CI, build, and deploy like any normal change, which takes real time even if every step passes cleanly, and it's not automatically faster just because it "undoes" something. - When you'd reach for each: rollback for the immediate, fast restoration of service; revert as the FOLLOW-UP action that keeps the source-control history clean and honest about what's actually running, and as the mechanism for making the "undo" permanent once you've confirmed the rollback fixed the problem (otherwise the next normal deploy, built from a source tree that still contains the bad change, would silently reintroduce the regression).
- Why both matter, not just one: rolling back WITHOUT eventually reverting means the next deploy from the current source tree reintroduces the bug, since the source code still contains the bad change even though the RUNNING version has been reverted; reverting without rolling back first means waiting through a full build-and-deploy cycle before service actually recovers, when a faster path was available.
Worked example
A regression discovered five minutes after a deploy: immediately kubectl rollout undo restores the previous, known-good version running in production within seconds. Separately, and not blocking that fast recovery, git revert <bad-commit> is pushed to keep the source tree consistent with what's actually running, so the next unrelated deploy (which will build from the current source tree) doesn't accidentally reintroduce the regression.
Trade-offs and pitfalls
The common mistake is treating these as interchangeable or doing only one: rolling back without ever reverting leaves a latent landmine in the source tree that resurfaces on the next deploy; reverting without rolling back first needlessly extends the outage while waiting for a full pipeline run when a faster path existed. The strongest practice is rollback FIRST for immediate recovery, revert SECOND (often within the same incident) to make the fix permanent in source control.
Unlock Full Question Bank
Get access to all 13 Safe Deployment and Rollback Strategies interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.