Safe Deployment and Rollback Strategies Questions
Releasing changes to production safely and incrementally, and recovering when they fail: blue-green, canary, and rolling deployments, feature flags, dark launches, traffic shifting, and progressive rollout, together with rollback strategies, safe-deploy practices, blast-radius containment, automated recovery, and safe forward/backward migration. Covers deployment orchestration across cloud platforms, staged exposure of new behavior to users, assessing deployment risk, designing reversible releases, and restoring a known-good state quickly. Focuses on how a release reaches production and how it is unwound on failure, distinct from broader incident command, which lives in Enterprise Operations & Incident Management.
Design an automated rollback approach for a stateful service whose release includes a database migration, using blue-green environments plus a read-only clone of the database for pre-migration verification. How do you minimize data loss and handle replication lag?
Sample Answer
Direct answer
Combining blue-green with a read-only database clone lets you validate a migration's effect on real, current data BEFORE committing to it on the live database: the clone gets the migration applied first, in isolation, so you catch a problem against production-representative data without any risk to the actual live system, and the blue-green switch itself still gives you fast rollback for the application layer once you do commit.
Structured elaboration
- Clone the production database (read-only) into an isolated environment and apply the migration to the CLONE first, validating both that the migration runs successfully and that the resulting data is correct, against real data characteristics (volume, distribution, edge cases) that synthetic test data might miss.
- If clone validation passes, apply the actual migration to production using the same discipline covered elsewhere (backward-compatible, expand-contract, batched for large tables), since the clone validated the LOGIC and DATA EFFECT, not the operational safety of running it against a live, concurrently-written system.
- Blue-green for the application layer: once the schema is safely migrated (backward-compatible, so both old and new app code can run against it), deploy the new application version to the green environment, validate it, and cut over traffic, keeping blue as an instant application-level rollback path.
- Minimizing data loss and handling replication lag: the production migration's write path needs an explicit boundary that prevents an in-flight write from landing in the gap between the old and new state. Concretely: take a fresh, final backup/snapshot immediately before the real migration begins (the earlier clone can be stale by the time the actual migration runs, so it isn't a substitute for this), run the migration as backward-compatible expand-contract in small, monitored batches so a failure partway through never forces discarding already-migrated data, and gate the blue-to-green traffic cutover on replication lag explicitly: define a maximum acceptable lag (for example, hold the cutover while green's replica lag exceeds a few seconds) and only cut traffic over once green has caught up to that threshold, rather than cutting over on a fixed timer regardless of lag. During the cutover moment itself, a brief write-quiesce or dual-write window (writes are accepted by blue and also applied to or replicated into green before green starts serving reads) closes the specific gap where a write landing in the last moments before cutover could otherwise be lost to whichever side ends up not serving traffic.
- Rollback scope: if a problem emerges post-cutover, blue-green gives fast APPLICATION rollback (switch back to blue), which is why replication should keep flowing from green back to blue for a defined window after cutover, so blue doesn't fall behind and a same-day switch-back doesn't lose whatever writes landed on green in the meantime; but if the issue traces to the migration itself rather than the application code, that's a data-layer rollback with its own considerations (covered by the backward-compatibility discipline that made the migration safe to begin with), not something the blue-green switch alone fixes.
Worked example
A migration converting a JSON blob column into normalized relational fields: applied first to a read-only production clone, revealing that roughly 2% of real production rows have malformed JSON that the migration's parsing logic doesn't handle, a data-shape problem synthetic test fixtures never surfaced. The migration logic is fixed to handle that edge case, re-validated against the clone, and only THEN applied to the actual production database with the same batched, lag-monitored discipline: a fresh pre-migration backup is taken, batches are throttled to keep replica lag under a defined threshold, and the blue-to-green traffic cutover waits until that threshold is met, with a brief dual-write window bridging the cutover moment itself so no write is lost in the gap; the application's blue-green cutover happens afterward, once the schema itself is confirmed safely migrated.
Trade-offs and pitfalls
Cloning a large production database is itself a real operational cost (storage, time to create the clone, and it can go stale relative to live production if there's a meaningful delay between cloning and actually running the real migration), so this technique earns its cost specifically for migrations complex or risky enough that catching a data-shape problem before it hits live data is worth the overhead; a simple, well-understood migration probably doesn't need it. The common mistake is treating clone validation as a substitute for the real migration's own operational safety discipline (batching, lag monitoring, a fresh pre-cutover backup, and an explicit lag threshold gating cutover) rather than as a complementary, earlier-stage check.
List at least three smoke tests you'd run immediately after a release. For each, state what it verifies and your response if it fails: alert, auto-rollback, or disable via feature flag.
Sample Answer
Direct answer
Three concrete smoke tests for a service, each targeted at a different class of catastrophic failure, and each paired with the response that actually fits its severity and blast radius: a health-endpoint check (auto-rollback on failure, since it signals the process itself may be broken), an end-to-end transaction through a specific new/flagged code path (disable via feature flag on failure, since that's the fastest way to shed just the risky new logic without discarding the rest of the release), and a critical-dependency reachability check (alert or auto-rollback depending on whether the dependency is critical-path, since not every degraded dependency justifies an immediate rollback).
Structured elaboration
- Health endpoint:
GET /healthzreturns 200. Verifies: the process started, is listening, and basic internal wiring didn't crash on boot. Success criteria: HTTP 200 within a short timeout (a few seconds). On failure: auto-rollback, immediately and without waiting for further evidence, since a process that isn't even up is the most severe and least ambiguous signal there is; there's no narrower fix available at this layer. - End-to-end core transaction through a flagged new code path: for a checkout service, place a test order through a sandboxed test account, specifically exercising a new pricing-calculation path that shipped behind a feature flag. Verifies: the ACTUAL new business logic works, not just that the process is alive. On failure: disable via feature flag first, not a full rollback, since the failure is isolated to logic that's already gated behind a flag; flipping the flag off falls back to the previous, known-good pricing path instantly across all instances without discarding the rest of the release. Escalate to a full auto-rollback only if disabling the flag doesn't resolve the failure (meaning the regression isn't actually confined to the flagged path).
- Critical-dependency reachability: confirm the service reports its database and payment-processor connections as healthy (via the health endpoint's detailed response or a separate dependency-check endpoint). Verifies: the new version can actually reach what it needs. On failure: the response depends on which dependency: if the payment-processor or primary database is unreachable, auto-rollback, since the service will degrade further as traffic increases; if a non-critical, degradable dependency (a recommendation service, a non-blocking analytics sink) is unreachable, alert-only and let a human decide, since the core service can keep functioning in a degraded mode and an automatic rollback would be an overreaction to a non-blocking issue.
Worked example
Immediately after deploy: /healthz returns 200 (pass). The flagged new pricing path's test order returns an incorrect total (fail); the pipeline flips the pricing feature flag off, and a re-run of the same test order now returns the expected total, confirming the flag disable resolved it without a full rollback. Separately, the dependency check shows the payment-processor connection as unreachable (fail); because this is a critical-path dependency, this triggers an automatic rollback of the whole release regardless of the pricing-flag outcome, since a service that can't reach its payment processor will fail broadly once real traffic hits it.
Trade-offs and pitfalls
The common mistake is treating every smoke-test failure the same way (blanket auto-rollback for everything), which discards a release's healthy majority just to fix a narrow, flag-isolated regression, and is slower in practice since the whole release then has to be re-shipped and re-verified from scratch. The opposite mistake, alerting on everything and waiting for a human, is too slow for unambiguous, severe failures like a dead health endpoint. Matching the response to the test (full rollback only when the failure isn't narrowly containable, feature-flag disable when it is, alert-only when the dependency isn't on the critical path) gets the fastest safe recovery in each case rather than one blunt instrument for every failure.
Evaluate progressive-delivery platforms such as Argo Rollouts, Flagger, and LaunchDarkly for a mid-size org. What criteria and architecture considerations would drive picking one over building an in-house solution, and how does each integrate with CI/CD and monitoring?
Sample Answer
Direct answer
Argo Rollouts and Flagger are both Kubernetes-native, open-source progressive-delivery controllers that integrate tightly with a service mesh or ingress and are effectively free beyond operational overhead, while LaunchDarkly is a commercial feature-flag platform focused on flag-based (not traffic-weight-based) progressive delivery with broader multi-platform SDK support; the right choice depends on whether your primary rollout mechanism is INFRASTRUCTURE-level traffic splitting or APPLICATION-level flag evaluation.
Structured elaboration
- Argo Rollouts: a Kubernetes CRD-based controller, deeply integrated with Kubernetes' own Deployment model, supporting canary and blue-green natively, with built-in metric-analysis steps that query Prometheus (or several other supported metrics providers) to gate promotion. Strongest fit for teams already deeply invested in Kubernetes and wanting traffic-weight-based canaries as a first-class, GitOps-friendly resource.
- Flagger: similar goal to Argo Rollouts, but built with a stronger service-mesh-first design (originally Istio-focused, now supporting several meshes), automating the mesh's traffic-splitting resources directly. Strongest fit for teams already running a service mesh and wanting canary automation that plugs directly into the mesh's existing traffic-management primitives.
- LaunchDarkly: not Kubernetes-specific at all, a general-purpose feature-flag platform with SDKs across many languages and platforms (mobile, web, backend), targeting rules, and audit/governance tooling; its "progressive delivery" is fundamentally APPLICATION-CODE flag evaluation, not infrastructure-level traffic splitting, so it's the natural fit when your rollout mechanism needs to work outside Kubernetes too (mobile apps, client-side web) or when the SPECIFIC risky behavior is better gated inside the code than at the network layer.
- Build-vs-buy criteria: how deeply Kubernetes-native your infrastructure already is (favors Argo Rollouts/Flagger), whether you need flag-based control across NON-Kubernetes surfaces too (favors LaunchDarkly or a similar commercial flag platform), and how much you value avoiding vendor lock-in and operational cost versus SDK maturity and support (open-source tools cost engineering time to operate; commercial platforms cost a subscription but reduce that operational burden).
- Integration with CI/CD and monitoring: all three integrate with a standard CI/CD pipeline as a deploy-time step (apply the Rollout/Canary resource, or call the flag-platform's API to update targeting rules) and all three can consume metrics from common monitoring backends (Prometheus for the Kubernetes-native tools; most commercial flag platforms integrate with common APM/analytics tools for outcome tracking, though the analysis itself is often less automated than Argo Rollouts'/Flagger's built-in metric-gated promotion).
Worked example
A team fully on Kubernetes with Istio already deployed would likely choose Flagger, since it plugs directly into infrastructure they already operate. A team with both a Kubernetes backend AND a native mobile app that both need coordinated, flag-based rollout of the same feature would more likely choose LaunchDarkly (or build a lighter custom flag layer), since Argo Rollouts and Flagger have no mechanism to control mobile-app behavior at all, being purely Kubernetes-traffic-layer tools.
Trade-offs and pitfalls
A common mistake is choosing based on which tool is more popular or well-known rather than which mechanism (traffic-weight-based infrastructure control vs. application-level flag control) actually matches the team's real rollout needs; a Kubernetes-native canary tool can't help you gate a risky behavior on a mobile app, and a flag platform doesn't give you infrastructure-level traffic-weight canarying without the application explicitly checking a flag on every request.
What makes a database migration backward-compatible? Give an example of a safe and an unsafe schema change, and explain why backward compatibility matters for rollback and phased deployment.
Sample Answer
Direct answer
A backward-compatible migration is one where the OLD version of your application code still works correctly against the NEW schema. That matters because during a rolling or canary deployment, old and new code run against the same database simultaneously, so if the new schema breaks the old code, you have an outage the moment the rollout starts, before you've even finished deploying.
Structured elaboration
- Safe (backward-compatible) changes: adding a new nullable column, adding a new table, adding a new index, widening a column's type (e.g. int to bigint in most databases). Old code that doesn't know about the new column simply ignores it; nothing it does breaks.
- Unsafe (non-backward-compatible) changes: renaming or dropping a column the old code still reads or writes, adding a NOT NULL column with no default (old code's INSERT statements, which don't set that column, start failing), changing a column's type in an incompatible direction (bigint to int, string to enum), or adding a foreign-key constraint on data the old code might still write in a way that violates it.
- Why it matters for rollback specifically: if you deploy a schema change together with new code and then need to roll the CODE back, the old code needs to keep working against the schema as it now stands, which is exactly the backward-compatibility property. If the migration wasn't backward-compatible, rolling back the code doesn't fully undo the outage, because the schema is still in its new, incompatible state.
Worked example
Safe: adding a nullable discount_code column to an orders table. Old code that doesn't reference it keeps working exactly as before. Unsafe: renaming orders.total to orders.total_amount in the same deploy as the code change that uses the new name; if you need to roll back the code while the rename has already run, the rolled-back old code tries to read orders.total, which no longer exists, and every read fails.
Trade-offs and pitfalls
Backward-compatible migrations usually take more steps and more calendar time (add, backfill, switch, THEN remove the old column in a later, separate deploy) than a direct rename, which is the trade-off teams are making: more process for a rollback safety net. The common mistake is treating "the migration ran successfully" as the same thing as "the migration is safe," when the real test is whether the PREVIOUS version of the application still functions correctly against the new schema.
Design a rollback runbook for a Kubernetes StatefulSet backed by persistent volumes. What's different about an in-place rollback here versus a stateless Deployment, and how do you verify data integrity afterward?
Sample Answer
Direct answer
Rolling back a StatefulSet is fundamentally different from a stateless Deployment because each pod's identity and persistent volume are tied together and preserved across the rollback: you're not just swapping which image runs, you're confirming the DATA on each pod's volume is actually compatible with the version you're rolling back to, which a stateless rollback never has to worry about.
Structured elaboration
- In-place rollback mechanics: like a Deployment,
kubectl rollout undo statefulset/<name>redeploys the previous pod template, but StatefulSet's ordered, one-at-a-time (by default) pod replacement means the rollback itself proceeds sequentially by ordinal (pod-2 rolls back before pod-1, in defaultOrderedReadypolicy), not in parallel batches the way a Deployment'smaxSurge/maxUnavailableallows. - Volume compatibility when downgrading: the previous application version needs to be able to correctly read whatever the CURRENT version wrote to the persistent volume; if the new version wrote data in a new format the old version can't read, rolling back the CODE doesn't roll back the DATA on disk, which is the exact same backward-compatibility problem seen with shared databases, just now per-pod instead of centralized.
- Snapshot/restore as a stronger fallback: for genuinely incompatible data, a code-level rollback alone isn't sufficient, and the plan needs to include restoring the persistent volume itself from a pre-upgrade snapshot, which is slower and loses any writes made since the snapshot, a real trade-off to weigh against the alternative of not being able to roll back cleanly at all.
- Verifying data integrity afterward: beyond confirming the pod is Running and Ready, this needs an application-level or storage-level integrity check specific to what the volume holds (a checksum, a consistency check the application itself can run, or for a database-backed StatefulSet, its own internal consistency-check tooling), since Kubernetes' own health signals say nothing about whether the DATA on the volume is actually intact and correct.
Worked example
A StatefulSet running a distributed cache with local persistent storage: rolling back the container image alone works cleanly IF the new version's on-disk cache-file format is backward-compatible with the old version, verified beforehand the same way a database schema's backward-compatibility would be. If the new version wrote an incompatible format, the rollback plan instead needs to fall back to restoring from a pre-upgrade volume snapshot per pod, accepting the loss of any cache writes since that snapshot, since there's no code-level fix that makes the old version able to read data in a format it was never built to understand.
Trade-offs and pitfalls
The most common mistake is treating StatefulSet rollback as mechanically identical to a Deployment's, assuming rollout undo alone is sufficient, when the real question, whether the on-disk data is compatible with the version being rolled back to, needs its own explicit answer before the rollback is trusted; a rollback that "succeeds" at the Kubernetes level (pod Running, Ready) can still be serving corrupted or misread data if that compatibility question was never actually checked.
Unlock Full Question Bank
Get access to all Safe Deployment and Rollback Strategies interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.