Fault Tolerance, High Availability, and Disaster Recovery Questions
Keeping a system serving despite failure, from code-level resilience to infrastructure-level recovery: circuit breakers, retries with backoff and jitter, timeouts, bulkheads, graceful degradation, and preventing cascading failures, alongside redundancy, failover (active-active versus active-passive), RPO and RTO objectives, backup and restore, and multi-region failover. Covers dependency-failure isolation, chaos engineering to validate resilience, failure-mode analysis, designing to nines of availability, cost-versus-availability tradeoffs, and recovery runbooks. Spans both the patterns that isolate partial failure and the disaster-recovery planning that restores a business-critical system after a major outage.
What does 'blast radius' mean when you're talking about a production failure? Name a few concrete engineering practices that reduce it, and what that costs you.
Sample Answer
Direct answer
Blast radius is the scope of impact when a component fails: how many users, tenants, or dependent services are affected, and how severely, not just whether the failure happened at all. Reducing blast radius means designing so a single failure touches the smallest possible slice of the system, which makes outages smaller, easier to detect, and faster to recover from, even if it doesn't reduce how often failures happen at all.
Practices that reduce it, and what they cost
| Practice | How it shrinks blast radius | What it costs |
|---|---|---|
| Circuit breakers | Stop repeated calls to a failing dependency, isolating the failure to the caller instead of letting it spread | Added latency and complexity in the failure path; a poorly tuned breaker can trip on transient blips |
| Finer-grained service decomposition | A failure or overload in one bounded service only affects its own consumers, not unrelated functionality | More services to deploy, monitor, and operate; cross-service calls add their own new failure modes |
| Bulkheads (per-tenant or per-dependency resource pools) | One tenant's or one dependency's exhaustion doesn't consume capacity meant for everyone else | More total resources provisioned (dedicated pools cost more than one shared pool sized for the average case) |
| Traffic shaping and rate limits | Caps how much load a single misbehaving client or spike can push into downstream systems | Legitimate bursty clients can get throttled unless limits are tuned carefully |
Worked example
Consider a service with 1,000 tenants sharing a single connection pool. If that pool exhausts, every tenant is affected. Now split that same total capacity into 10 isolated pools of 100 tenants each, so each pool serves 100 of the 1,000 tenants and only that pool's own tenants are affected if it exhausts:
1,000100=10% of tenants affected (isolated pools)vs.100% (shared pool)Splitting the same total capacity into 10 pools of 100 tenants each means a single pool's exhaustion now affects only 100 of the 1,000 tenants, 10% of the blast radius of the shared-pool design, for the same total resources. The cost is operational: 10 pools to monitor and size instead of one, and if traffic isn't evenly distributed across tenants, some pools may be under-utilized while others are tight, which the shared pool didn't have to worry about.
Trade-offs & pitfalls
Reducing blast radius is generally a trade of operational complexity and some resource inefficiency for smaller, more contained failures; it doesn't reduce the underlying failure rate of any individual component. The common mistake is treating blast-radius reduction as free: partitioning by tenant, region, or dependency multiplies the number of things to monitor and can hide a systemic bug (one that affects every partition equally) behind what looks like ten separate, unrelated small incidents instead of one clearly systemic one.
Design a retry strategy with exponential backoff and jitter for calls to a downstream dependency that's struggling. Walk through why jitter matters, and how you'd make sure your retries don't make the dependency's problem worse when it starts recovering.
Sample Answer
Plain exponential backoff (double the delay after each failed attempt) reduces load on a struggling dependency over time, but it has a hidden flaw: if many clients failed at roughly the same moment (which is exactly what happens when the dependency itself goes down), they all compute the same delay sequence and retry in lockstep, so the "backoff" just delays the same synchronized spike instead of spreading it out. Jitter fixes that by randomizing the delay so clients that failed together don't retry together.
Jitter strategies compared
| Strategy | Delay formula | Behavior |
|---|---|---|
| No jitter | delay=base×2attempt | Deterministic; every client that failed together retries together, recreating the spike at each step |
| Full jitter | delay=random(0, base×2attempt) | Maximum spread; delay can be anywhere from 0 up to the cap, so retries are smeared thinly across the whole window |
| Equal jitter | delay=2cap+random(0, 2cap) | Keeps a guaranteed minimum delay (never retries immediately) while still spreading the upper half randomly |
| Decorrelated jitter | delay=random(base, previous delay×3) | Grows the delay based on the client's own previous delay rather than a fixed exponential schedule, avoiding a hard cap while still spreading load |
Worked example: how much jitter actually reduces the spike
Pin a concrete scenario: 1000 clients failed at the same moment, base delay = 1 second, and this is their 3rd retry attempt (attempt = 3), so the backoff cap is:
cap=1s×23=8 secondsWithout jitter: every one of the 1000 clients computes the identical 8-second delay and retries at exactly the same instant, a spike of 1000 concurrent requests hitting the dependency in one moment, right as it may just be starting to recover.
With full jitter, each client independently draws a delay uniformly from [0,8] seconds. Dividing that 8-second window into 100 ms buckets gives 8000/100=80 buckets, and under a uniform distribution the expected number of clients landing in any single bucket is:
801000=12.5 requests per 100ms bucketThat's a peak-to-average reduction factor of 1000/12.5=80× under this modeling assumption (uniform, independent draws), turning one instantaneous spike of 1000 into a smooth trickle of roughly 12-13 requests every 100 ms across the full 8-second window, which a recovering dependency can absorb where a single 1000-request spike would knock it back down.
Why retries shouldn't make recovery worse
Jitter alone doesn't prevent the retry storm from getting worse over time if attempts aren't capped: a client that keeps failing and keeps retrying at base×2attempt forever will eventually be sending requests at a cap so large it's functionally giving up, or, worse, if the cap is bounded, converges back to a steady drumbeat of load that never lets the dependency fully recover. The fix is a hard cap on both the maximum delay and the maximum number of attempts, plus honoring any explicit signal the server provides (a Retry-After header or a 429/503 status) as authoritative over the client's own backoff schedule, since the server is in the best position to know its own recovery state.
Trade-offs and pitfalls
Full jitter maximizes spread but means some unlucky clients draw a near-zero delay and retry almost immediately, which is fine in aggregate (that's still only ~12-13 requests per 100ms bucket in the example above) but means full jitter alone doesn't guarantee a minimum backoff for any individual client; equal jitter trades some of that spread for a guaranteed floor, useful when even a small number of near-instant retries is unacceptable. A pitfall specific to mobile or otherwise unreliable-network clients: retries are only safe to jitter and reattempt if the underlying operation is idempotent (repeating it produces the same end result as doing it once, so a duplicate attempt is harmless), a non-idempotent submit (a payment, an order) retried after a client-side timeout can double-execute if the server had actually processed the first attempt and just failed to deliver the response, so the fix belongs on the server (idempotency keys deduping identical requests) not just in the client's backoff logic, jitter reduces load, it does not make an unsafe retry safe.
Design a multi-tenant platform so that a noisy tenant can't degrade availability for everyone else. Walk through your isolation and enforcement mechanisms, and how autoscaling and billing interact with the quotas you set.
Sample Answer
Direct answer
Isolate at three layers so no single control point being bypassed breaks the guarantee: admission control at the edge (reject or queue over-quota requests before they consume shared resources), hard resource enforcement at the OS/scheduler level (so a tenant that gets past admission still can't starve others), and fairness-aware autoscaling and billing (scale for real aggregate demand, allocate new capacity by weighted share rather than first-come-first-served, and meter overage so legitimate extra usage is paid for instead of stolen from neighbors).
Architecture and enforcement
flowchart LR
T[Tenant request] --> GW[Admission gateway: token bucket]
GW -->|within quota| ENQ[Per-tenant queue]
GW -->|over quota| REJ[429 / shed]
ENQ --> SCHED[Scheduler: cgroup CPU/IO caps]
SCHED --> POOL[Shared compute pool]
POOL --> AS[Autoscaler: per-tenant saturation signal]
AS --> POOL
POOL --> MET[Usage metering]
MET --> BILL[Billing: quota + overage]
Base quota per tenant is a weighted fair share of total capacity C:
Qi=C×∑jwjwi- Admission layer: per-tenant token bucket enforces request-rate and concurrency caps before work is queued; this stops a noisy tenant from ever occupying shared queue depth.
- Enforcement layer: the OS and scheduler cap what a tenant's workload can actually consume, CPU time, memory, and network bandwidth, even if a request slips past admission control, defense in depth rather than trusting one gate alone. (Advanced implementation detail: on Linux this is typically cgroups for CPU bandwidth limits, OOM-score tuning to control which processes get killed first under memory pressure, and network policers such as tc/ingress qdiscs to cap per-tenant bandwidth.)
- Autoscaling: the trigger must look at per-tenant saturation, not just cluster-wide CPU. If a single tenant is the one hammering the cluster, scaling out on aggregate cluster metrics effectively rewards that tenant with new capacity paid for by everyone, which defeats the isolation goal economically even if it holds technically.
- Billing: usage beyond the base quota is metered as overage rather than throttled outright for tenants on a plan that allows bursting; quota tier is the plan's contract, burst is a paid escape valve, not a loophole.
Worked example
Cluster capacity C=1000 vCPU across four tenants with weights A=3, B=1, C=1, D=1 (their plan tiers):
C=1000, ∑w=6:QAQB=QC=QD=1000×63=500=1000×61=166.67If tenant D is idle, its unused 166.67 vCPU can be reclaimed and redistributed proportionally among the active tenants (weights 3, 1, 1 summing to 5), rather than sitting wasted:
RA′B′C′=QD=166.67=QA+R×53=500+100=600=QB+R×51=166.67+33.33=200=QC+R×51=166.67+33.33=200A' + B' + C' + 0 = 1000, the full capacity, allocated fairly by weight instead of by who asked first. If D comes back online, the reclaim has to be preemptible (graceful, with a grace period) so D isn't hard-killed the moment it wakes up; it gets its guaranteed floor back and the borrowers scale down.
Trade-offs and pitfalls
Hard caps with no reclaim protect worst-case isolation perfectly but waste real capacity whenever a tenant is idle, which is most of the time for most tenants. Soft caps with reclaim use capacity efficiently but add real scheduler complexity, and the failure mode to avoid is yanking borrowed capacity abruptly instead of draining it, which turns "fairness" into a second source of noisy-neighbor-style disruption for the tenant who briefly borrowed it. The single most common wrong turn in this design is coupling the autoscaler's scale-out trigger directly to raw cluster-wide metrics: it looks like a capacity problem and "solves" it by adding nodes, but if the actual cause is one tenant blowing past its quota, autoscaling just subsidizes the noisy tenant at everyone else's expense until billing catches up, if it ever does.
What does graceful degradation mean for a resilient system, and why does it matter? Pick a user-facing service, like search or checkout, and walk through which features you'd disable first under partial failure, and which you'd protect at all costs.
Sample Answer
Direct answer: Graceful degradation means a system keeps serving its core value under partial failure by deliberately shedding non-essential features, instead of failing completely because one dependency is unhealthy. It matters because most real outages are partial, not total, and a system that can't distinguish "checkout is down" from "product recommendations are down" ends up treating both the same way: total outage, when only one of them actually deserved it.
Structured elaboration
The core discipline is ranking features by how essential they are to the user's actual goal, then deciding in advance what happens to each tier when its supporting dependency fails:
| Priority | Category | What happens under partial failure |
|---|---|---|
| Protect at all costs | The core transaction (e.g., add to cart, checkout, payment) | Never disabled; if its own dependency fails, fail the request loudly rather than silently corrupt it |
| Degrade first | Personalization and enrichment (recommendations, "customers also bought," rich previews) | Hide the widget or fall back to a generic/cached version; the page still loads and functions |
| Degrade next | Non-critical background work (analytics events, telemetry sampling, async inventory sync) | Drop or buffer, since losing this doesn't affect the current user's experience |
How you decide what's "core": ask whether the feature is on the path the user came for. For a checkout service, that's the cart-to-payment path; product recommendations, reviews, and "recently viewed" are enrichment around that path, valuable but not why the user is there. For a search service, returning some relevant results is core; typo-correction, personalized re-ranking, and query autocomplete are enrichment that can be dropped without breaking the user's ability to search.
Detecting when to degrade: this has to be automatic, not something a human decides mid-incident. Health checks and latency/error-rate thresholds on each dependency feed a circuit breaker; when the breaker for the recommendations service opens, the front end (or an API gateway) simply omits that section rather than waiting on a call that's failing. The degraded state should be visible in monitoring (a "degraded mode" flag, not silence) so the team knows it's active and can address root cause.
Trade-offs & pitfalls
- Degrading too aggressively removes revenue-generating features (recommendations often drive real conversion) for failures that didn't actually require it; the tiering has to be based on actual dependency health, not a blanket "anything non-core gets cut."
- Degrading too conservatively (waiting too long, or requiring a human to flip a switch) means the cascading failure the degradation was supposed to prevent happens anyway, because by the time a human reacts, the core path is already backed up.
- Static thresholds don't generalize across traffic levels; a latency threshold tuned for average traffic can either never trigger during a real incident at peak load, or trigger too eagerly during a routine traffic spike that isn't actually a failure.
- Testing degraded paths is easy to skip because they're rarely exercised in normal operation; without deliberately forcing dependencies to fail in staging (or via chaos testing in production), the first real test of the degraded path is during an actual incident, which is the worst time to discover it's broken.
- The same tiering logic applies outside typical web services: an ML-serving system facing a slow or unavailable model can fall back to a cached prior response, swap to a smaller/cheaper model that's faster but less accurate, or return a safe default decision, the exact same "protect the core interaction, shed the enrichment" reasoning, just with "model quality" instead of "page richness" as the thing being traded off.
How would you plan and run a game day to validate your team's DR readiness? Walk through how you'd scope it, who you'd involve, how you'd measure impact against your SLIs, and what you'd do with the findings afterward.
Sample Answer
Direct answer
A good game day has a tightly bounded scope, a named set of stakeholders who signed off before the experiment starts, a real-time comparison of the system's behavior against its SLIs (service level indicators: the specific numbers you track, like latency and error rate, that tell you whether the system is healthy) during the run, and a retrospective that turns findings into tracked action items, not just a summary email. The hard part is not running the experiment; it's building the recurring program and organizational trust that lets you run harder ones over time.
Scoping the experiment
Pick a single, realistic failure mode against a bounded slice of traffic: a specific dependency (cache, database replica, a downstream API), a specific service, and ideally a canary or staging slice of load rather than 100% of production on the first run. Define upfront what "done" looks like: which SLIs you'll watch, what the abort condition is, and who has authority to hit the kill switch.
Who to involve
- Service owners and on-call engineers for the system under test, since they know the failure modes and own the runbook being validated.
- A designated incident commander for the exercise itself, separate from whoever is executing the fault injection, so there's a clear decision-maker if things go sideways.
- Product or support stakeholders when the blast radius could touch real users, so they understand what "the recommendations service is intentionally broken for 20 minutes" means for anyone who notices.
- Observability or SRE tooling owners to make sure dashboards and alerting are actually wired up to catch what you're about to do, not just to catch organic incidents.
Measuring impact against SLIs
Capture a baseline of your SLIs (latency percentiles, error rate, saturation) before injecting the fault, then watch the same SLIs in real time during the run and compare against the SLOs (service level objectives: the target values you've committed to for those same indicators, e.g. 99.9% success rate). The goal is not "did it break" (you know it will) but "did it break within the bounds you predicted, and did the defenses (timeouts, circuit breakers, autoscaling) behave the way the runbook assumes they do."
Turning findings into a recurring program
A single successful game day proves one thing worked once. Standing up a recurring practice requires a roadmap: start with low-risk, staging-only experiments to build muscle memory and trust, then progressively widen scope (larger blast radius, real production traffic, less-scripted scenarios) as the team demonstrates it can run these safely. Getting buy-in usually means showing leadership a concrete finding from an early, low-risk drill (a specific gap the exercise surfaced) rather than asking for blanket permission to break production up front. Once a cadence is established (for example, monthly), track a maturity metric across runs, such as the fraction of prior findings that were actually remediated before the next drill, so the program itself is accountable.
Worked example
Consider a payments API game day: inject 200ms of added latency into its database replica for a scoped window, on a canary slice of traffic, with a monthly error budget of 43.2 minutes at a 99.9% SLO:
43.2=30×24×60×(1−0.999) minutes, the monthly error budget at a 99.9% SLODuring the drill, the induced latency causes synchronous retries to queue up, and the service is measurably degraded (error rate above SLO) for 12 minutes before the circuit breaker trips and the fallback path kicks in. That single test consumed:
43.212≈27.8% of the monthly error budget consumed by one testThat is a legitimate, alarming finding on its own: a single scoped drill burning over a quarter of the monthly error budget means either the blast radius needs to be tightened further (smaller canary percentage) or the circuit breaker's failure threshold needs to trip faster. Either way it's a concrete, numeric input for the retrospective and the case for continued investment in the program, rather than a vague "went well."
Trade-offs & pitfalls
Widening scope too fast is the single biggest risk to the program's survival: one game day that causes a real customer-visible incident before the team has built confidence can kill the practice for a year. The opposite failure is scoping every drill so conservatively that it never surfaces anything new, which also erodes stakeholder buy-in because the exercise starts to look like theater. The retrospective is where most of the value is either captured or lost; findings that don't get a tracked owner and a re-test in the next cycle tend to silently repeat.
Unlock Full Question Bank
Get access to all Fault Tolerance, High Availability, and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.