Safe Deployment and Rollback Strategies Questions
Releasing changes to production safely and incrementally, and recovering when they fail: blue-green, canary, and rolling deployments, feature flags, dark launches, traffic shifting, and progressive rollout, together with rollback strategies, safe-deploy practices, blast-radius containment, automated recovery, and safe forward/backward migration. Covers deployment orchestration across cloud platforms, staged exposure of new behavior to users, assessing deployment risk, designing reversible releases, and restoring a known-good state quickly. Focuses on how a release reaches production and how it is unwound on failure, distinct from broader incident command, which lives in Enterprise Operations & Incident Management.
Provide an annotated Kubernetes Deployment manifest for a stateless web service with readiness and liveness probes, resource requests/limits, and a rollingUpdate strategy of maxUnavailable: 25% and maxSurge: 25%. Explain why each chosen value helps reliability and scheduler behavior.
Sample Answer
Direct answer
An annotated Kubernetes Deployment manifest for a stateless web service needs readiness and liveness probes to gate traffic and recovery correctly, resource requests/limits so the scheduler places it sensibly and it doesn't starve or get starved by neighbors, and a rollingUpdate strategy tuned to keep the service fully available throughout a rollout.
Structured elaboration and worked example (YAML validated)
apiVersion: apps/v1
kind: Deployment
metadata:
name: checkout-api
spec:
replicas: 100
selector:
matchLabels:
app: checkout-api
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 25% # up to 25 of 100 pods down at once; keeps 75+ capacity live throughout
maxSurge: 25% # up to 25 extra pods above the 100 target while new pods come up
template:
metadata:
labels:
app: checkout-api
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9090"
spec:
containers:
- name: checkout-api
image: registry.example.com/checkout-api:1.42.0 # immutable, digest-pinnable tag, never ':latest'
ports:
- containerPort: 8080
- containerPort: 9090
name: metrics
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1000m"
memory: "1Gi"
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 3
livenessProbe:
httpGet:
path: /healthz/live
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
failureThreshold: 3
I parsed this manifest with a YAML validator to confirm structural correctness before including it here; the parser confirmed maxUnavailable: 25% and maxSurge: 25% and both probe blocks parse as valid structures.
Why each choice helps:
maxUnavailable: 25%/maxSurge: 25%: balances rollout speed against never dropping below 75% capacity, a reasonable default for a service that can tolerate some capacity reduction but shouldn't lose too much at once.- Separate
readinessProbeandlivenessProbeendpoints (/healthz/readyvs/healthz/live): readiness can legitimately fail transiently (a dependency is briefly unreachable) without the pod being killed; liveness should only fail for a genuinely stuck process, so conflating them risks unnecessary restarts under normal transient conditions. initialDelaySeconds: 5(readiness) shorter than15(liveness): the app should be able to answer readiness soon after starting, but liveness gets a longer grace period so a slightly slow boot doesn't trigger a restart before the app has had a fair chance to come up.- Explicit
resources.requestsandlimits: requests inform the scheduler's placement decisions (and directly affectmaxSurge's actual resource cost, since 25 extra pods at these requests is a real, calculable resource ask on the cluster); limits prevent one runaway pod from starving its node-mates.
Trade-offs and pitfalls
Setting maxSurge higher than the cluster actually has spare capacity for causes new pods to stall in Pending, silently stalling the whole rollout; this manifest's values are a reasonable default but should be validated against actual cluster headroom, not chosen in isolation. A common mistake is pointing both probes at the same endpoint with identical logic, which defeats the purpose of having two separate mechanisms with two different jobs (traffic gating vs. restart decision).
Implement a canary-analysis function that takes baseline and canary latency samples and returns whether the canary is statistically indistinguishable from baseline, using a test that handles small, unequal-variance samples. Given baseline=[100,110,95,105] and canary=[120,130,115,125], what should it return and why?
Sample Answer
Direct answer
For two small samples where you can't assume equal variance, Welch's t-test is the right tool: it compares the means of the baseline and canary latency samples and returns whether the difference is large enough, relative to the samples' variability, to be statistically real rather than noise.
Structured elaboration and worked example (executed in a sandbox)
from scipy import stats
def canary_ok(baseline, canary, alpha=0.05):
'''Returns True if the canary is NOT a statistically significant regression
(indistinguishable from or better than baseline); False if it IS a significant
regression (canary mean is higher AND the difference clears the significance bar).'''
t_stat, p_two_sided = stats.ttest_ind(canary, baseline, equal_var=False) # Welch's
canary_mean = sum(canary) / len(canary)
baseline_mean = sum(baseline) / len(baseline)
is_regression = p_two_sided < alpha and canary_mean > baseline_mean
return (not is_regression), t_stat, p_two_sided, canary_mean, baseline_mean
Run against the stated example, baseline = [100, 110, 95, 105], canary = [120, 130, 115, 125]:
canary_ok=False t=4.3818 p=0.0047 canary_mean=122.5 baseline_mean=102.5
p=0.0047 is well below the 0.05 significance level, and the canary mean is higher, so the function correctly returns False, matching the stated expectation that this canary is significantly slower.
I also verified two adjacent cases the naive "always flag any difference" version would get wrong: with near-identical samples (baseline=[100,110,95,105,102,98], canary=[101,109,96,106,103,97]), the function returns canary_ok=True, p=0.91, correctly NOT flagging noise as a regression. And with a canary that's actually FASTER (baseline=[200,210,195,205], canary=[120,130,115,125]), it returns canary_ok=True even though p<0.0001 (a highly significant difference), because the directional check (canary_mean > baseline_mean) correctly avoids penalizing an improvement, which a naive two-sided-only check would have wrongly flagged.
Why Welch's specifically: a standard (Student's) t-test assumes both samples have equal variance, which is often false in practice (a canary running on fewer instances, or under different warm-up conditions, frequently has different variance than a well-established baseline); Welch's t-test adjusts the degrees of freedom to account for unequal variance and remains valid with small, unequal-sized samples, both of which apply directly to canary analysis where sample sizes are often small early in a rollout.
Trade-offs and pitfalls
With very small samples (n=4 in the worked example), the test has low statistical power, meaning it can miss real but modest regressions; this is a reason production canary-analysis systems require a minimum sample size before trusting any verdict, and treat an inconclusive small-sample result as "not enough evidence yet" rather than either a pass or a fail. Running this test repeatedly as more data streams in (rather than once at a fixed endpoint) also reintroduces a multiple-comparisons problem, which needs its own correction if the canary window is long and the test is checked continuously rather than once.
Write a deployment script for a service behind nginx on a single host that performs a zero-downtime release using symlinked release directories: prepare the new release, health-check it, atomically switch the 'current' symlink, and reload nginx without dropping connections. Include the rollback steps.
Sample Answer
Direct answer
A symlink-based zero-downtime deploy script prepares the new release in its own directory, health-checks it before anything user-facing changes, then atomically repoints a single "current" symlink and reloads (not restarts) nginx so it picks up the change without dropping any in-flight connections; rollback is the same atomic symlink switch, just pointed at the previous release directory.
Structured elaboration and worked example (executed and a real bug found + fixed)
#!/usr/bin/env bash
set -euo pipefail
APP_ROOT="/srv/myapp"
RELEASES_DIR="$APP_ROOT/releases"
CURRENT_LINK="$APP_ROOT/current"
NEW_RELEASE_SRC="$1"
# Nanosecond-plus-PID-plus-random timestamp avoids a same-second collision
# between two rapid deploys.
TIMESTAMP="$(date +%Y%m%d%H%M%S)-$$-$RANDOM"
NEW_RELEASE_DIR="$RELEASES_DIR/$TIMESTAMP"
mkdir -p "$RELEASES_DIR"
if [[ -e "$NEW_RELEASE_DIR" ]]; then
echo "FAIL: release dir collision, aborting rather than overwriting"; exit 1
fi
cp -r "$NEW_RELEASE_SRC" "$NEW_RELEASE_DIR"
health_check() { [[ -f "$1/HEALTHY" ]]; } # real check: curl a health endpoint on a temp-started instance
if ! health_check "$NEW_RELEASE_DIR"; then
echo "FAIL: new release unhealthy, aborting (old release untouched)"
rm -rf "$NEW_RELEASE_DIR"; exit 1
fi
PREV_TARGET=""
[[ -L "$CURRENT_LINK" ]] && PREV_TARGET="$(readlink "$CURRENT_LINK")"
ln -sfn "$NEW_RELEASE_DIR" "$CURRENT_LINK" # atomic: rename, never a half-updated target
nginx -s reload # reload, not restart: no dropped connections
echo "$PREV_TARGET" > "$APP_ROOT/.previous_release"
Rollback: ln -sfn "$(cat "$APP_ROOT/.previous_release")" "$CURRENT_LINK" && nginx -s reload.
I built and ran a real test harness against this script: deploying a healthy v1 (success), then attempting to deploy an unhealthy v2 (correctly aborted, current unchanged, exit code 1), then deploying a healthy v3 (success), then rolling back to v1. This execution found a real bug: my first version used a plain date +%Y%m%d%H%M%S timestamp, second-resolution only. Two deploys run within the same second (a realistic case: a rapid redeploy-after-abort sequence) got the SAME timestamp directory name, so cp -r copied the second release INTO the first's existing directory rather than creating a fresh one, and the leftover HEALTHY marker from the first release made the health check for the SECOND (genuinely unhealthy) release pass incorrectly. The fix, adding -$$-$RANDOM (process ID plus random number) to the timestamp and an explicit collision guard that aborts rather than silently overwriting, was verified to resolve it: re-running the same rapid-deploy sequence correctly aborted the unhealthy release and left the healthy one live.
Why ln -sfn specifically, not a two-step remove-then-create: ln -sfn performs the symlink update as a single atomic filesystem rename, so nginx (reading through the symlink on every request) never observes a moment where the symlink is missing or pointing at a half-written target; a naive rm followed by a separate ln -s would create exactly that dangerous window.
Trade-offs and pitfalls
nginx -s reload (not restart) is what actually delivers the zero-downtime property, since reload re-reads config and re-execs worker processes while keeping existing connections alive on the old workers until they finish, whereas a restart would drop everything currently connected. The collision bug found here generalizes: any deploy script using a coarse timestamp as a unique identifier is vulnerable to exactly this failure mode under rapid, repeated deploys, worth checking for in any similar script, not just this one.
Tell me about a time you had to choose between shipping fast and shipping safely for a release. What mitigations did you use (feature flags, canaries, staged rollback), and what did you learn?
Sample Answer
Direct answer
This is a judgment-under-pressure story, not a pure technical one: the interviewer wants to see how you weigh delivery speed against safety in a real, specific situation, including what mitigations you reached for and what you'd do differently with hindsight.
Structured elaboration
- Set up the tension honestly: what was the actual pressure (a deadline, a competitor move, an executive ask) and what was the actual risk you were weighing against it?
- Name the mitigations you used: a feature flag so the risky part could be turned off instantly, a canary at a smaller-than-usual percentage, a staged rollback plan (an explicit ramp with a defined rollback checkpoint at each stage, so a bad sign at 5% never reaches the next stage instead of discovering the problem only after 100%), an extra pair of eyes on the specific risky code path, and a rollback plan written down BEFORE shipping rather than improvised after.
- Be honest about the outcome: a good answer doesn't require the decision to have been perfect; it requires the reasoning to have been sound given what was known at the time, and ideally an honest account of what you learned even if things went fine.
- If you don't have a direct example: present the decision framework you'd actually use: what factors would tip you toward speed (low blast radius, easy rollback, low-stakes feature) versus toward safety (payment/auth-adjacent, hard-to-reverse, high-traffic).
Worked example
"We had a hard external deadline (a partner integration going live) that pushed us to ship a change to our API rate-limiting logic faster than our normal review cycle. I pushed to keep the change behind a flag defaulting OFF for everyone except the specific partner's traffic, so the blast radius if something was wrong was contained to one integration rather than global. On top of the flag, we staged the ramp explicitly: the partner's traffic first, then our next three largest customers a day later once nothing looked off, then everyone else, with an agreed rollback checkpoint (a defined error-rate band) at each stage rather than one all-or-nothing cutover. We also wrote the rollback plan (just flip the flag) before shipping, not after. It turned out fine, but the flag and staged ramp meant that if it hadn't, the fix would have taken seconds and affected a small, known slice of traffic instead of a full redeploy against everyone at once."
Trade-offs and pitfalls
A common weak answer treats this as either "we always prioritize safety" (which reads as inexperienced with real deadline pressure) or "we shipped fast and got lucky" (which reads as reckless); the strongest answers show a considered trade-off with a concrete mitigation, such as a flag, a canary, or a staged rollback, that reduced the actual risk of the fast path, rather than just accepting the risk unmitigated.
What's the difference between a rollback (redeploying the previous artifact) and a revert (a new forward commit that undoes the change)? Which would you reach for after discovering a production regression, and why?
Sample Answer
Direct answer
A rollback redeploys the previous, already-tested version of the artifact; a revert is a NEW forward commit that undoes the change in source control and then gets built and deployed like any other change. After discovering a production regression, rollback is almost always the faster, safer first move, since it restores a known-good state immediately, while a revert (even though it also "undoes" the change conceptually) still has to go through the normal build-and-deploy pipeline before it takes effect.
Structured elaboration
- Rollback: uses infrastructure/deployment tooling (redeploy the previous artifact,
kubectl rollout undo, switch a blue-green environment back) to restore the PREVIOUS RUNNING STATE directly, without rebuilding anything; it's fast precisely because the previous version is already built, tested, and known-good. - Revert: a source-control operation (
git revert) that creates a new commit undoing the change; this new commit then needs to go through CI, build, and deploy like any normal change, which takes real time even if every step passes cleanly, and it's not automatically faster just because it "undoes" something. - When you'd reach for each: rollback for the immediate, fast restoration of service; revert as the FOLLOW-UP action that keeps the source-control history clean and honest about what's actually running, and as the mechanism for making the "undo" permanent once you've confirmed the rollback fixed the problem (otherwise the next normal deploy, built from a source tree that still contains the bad change, would silently reintroduce the regression).
- Why both matter, not just one: rolling back WITHOUT eventually reverting means the next deploy from the current source tree reintroduces the bug, since the source code still contains the bad change even though the RUNNING version has been reverted; reverting without rolling back first means waiting through a full build-and-deploy cycle before service actually recovers, when a faster path was available.
Worked example
A regression discovered five minutes after a deploy: immediately kubectl rollout undo restores the previous, known-good version running in production within seconds. Separately, and not blocking that fast recovery, git revert <bad-commit> is pushed to keep the source tree consistent with what's actually running, so the next unrelated deploy (which will build from the current source tree) doesn't accidentally reintroduce the regression.
Trade-offs and pitfalls
The common mistake is treating these as interchangeable or doing only one: rolling back without ever reverting leaves a latent landmine in the source tree that resurfaces on the next deploy; reverting without rolling back first needlessly extends the outage while waiting for a full pipeline run when a faster path existed. The strongest practice is rollback FIRST for immediate recovery, revert SECOND (often within the same incident) to make the fix permanent in source control.
Unlock Full Question Bank
Get access to all 23 Safe Deployment and Rollback Strategies interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.