Fault Tolerance, High Availability, and Disaster Recovery Questions
Keeping a system serving despite failure, from code-level resilience to infrastructure-level recovery: circuit breakers, retries with backoff and jitter, timeouts, bulkheads, graceful degradation, and preventing cascading failures, alongside redundancy, failover (active-active versus active-passive), RPO and RTO objectives, backup and restore, and multi-region failover. Covers dependency-failure isolation, chaos engineering to validate resilience, failure-mode analysis, designing to nines of availability, cost-versus-availability tradeoffs, and recovery runbooks. Spans both the patterns that isolate partial failure and the disaster-recovery planning that restores a business-critical system after a major outage.
What does graceful degradation mean for a resilient system, and why does it matter? Pick a user-facing service, like search or checkout, and walk through which features you'd disable first under partial failure, and which you'd protect at all costs.
Sample Answer
Direct answer: Graceful degradation means a system keeps serving its core value under partial failure by deliberately shedding non-essential features, instead of failing completely because one dependency is unhealthy. It matters because most real outages are partial, not total, and a system that can't distinguish "checkout is down" from "product recommendations are down" ends up treating both the same way: total outage, when only one of them actually deserved it.
Structured elaboration
The core discipline is ranking features by how essential they are to the user's actual goal, then deciding in advance what happens to each tier when its supporting dependency fails:
| Priority | Category | What happens under partial failure |
|---|---|---|
| Protect at all costs | The core transaction (e.g., add to cart, checkout, payment) | Never disabled; if its own dependency fails, fail the request loudly rather than silently corrupt it |
| Degrade first | Personalization and enrichment (recommendations, "customers also bought," rich previews) | Hide the widget or fall back to a generic/cached version; the page still loads and functions |
| Degrade next | Non-critical background work (analytics events, telemetry sampling, async inventory sync) | Drop or buffer, since losing this doesn't affect the current user's experience |
How you decide what's "core": ask whether the feature is on the path the user came for. For a checkout service, that's the cart-to-payment path; product recommendations, reviews, and "recently viewed" are enrichment around that path, valuable but not why the user is there. For a search service, returning some relevant results is core; typo-correction, personalized re-ranking, and query autocomplete are enrichment that can be dropped without breaking the user's ability to search.
Detecting when to degrade: this has to be automatic, not something a human decides mid-incident. Health checks and latency/error-rate thresholds on each dependency feed a circuit breaker; when the breaker for the recommendations service opens, the front end (or an API gateway) simply omits that section rather than waiting on a call that's failing. The degraded state should be visible in monitoring (a "degraded mode" flag, not silence) so the team knows it's active and can address root cause.
Trade-offs & pitfalls
- Degrading too aggressively removes revenue-generating features (recommendations often drive real conversion) for failures that didn't actually require it; the tiering has to be based on actual dependency health, not a blanket "anything non-core gets cut."
- Degrading too conservatively (waiting too long, or requiring a human to flip a switch) means the cascading failure the degradation was supposed to prevent happens anyway, because by the time a human reacts, the core path is already backed up.
- Static thresholds don't generalize across traffic levels; a latency threshold tuned for average traffic can either never trigger during a real incident at peak load, or trigger too eagerly during a routine traffic spike that isn't actually a failure.
- Testing degraded paths is easy to skip because they're rarely exercised in normal operation; without deliberately forcing dependencies to fail in staging (or via chaos testing in production), the first real test of the degraded path is during an actual incident, which is the worst time to discover it's broken.
- The same tiering logic applies outside typical web services: an ML-serving system facing a slow or unavailable model can fall back to a cached prior response, swap to a smaller/cheaper model that's faster but less accurate, or return a safe default decision, the exact same "protect the core interaction, shed the enrichment" reasoning, just with "model quality" instead of "page richness" as the thing being traded off.
A downstream service you depend on starts responding slowly, and requests to it start backing up on your side, growing queues and increasing latency. Walk through your immediate mitigations and your longer-term architectural fix, and explain the trade-off each one introduces.
Sample Answer
Direct answer: The immediate priority is to stop the slowdown from consuming your own resources: set aggressive timeouts, open a circuit breaker so you stop calling the failing dependency, and isolate the connection/thread pool used for that call so it can't starve everything else. The longer-term fix is architectural: decouple the caller from the dependency's latency entirely, usually via an async queue or by making the call non-blocking, so a slow downstream degrades throughput instead of taking the whole service down with it.
Structured elaboration
Why this happens (the mechanism): by Little's Law, the number of requests in flight L equals arrival rate λ times the time each request spends in the system W: L=λW. If a downstream call's latency goes from 50ms to 500ms while your request rate stays at, say, 200 requests/second, the in-flight count grows from L=200×0.05=10 to L=200×0.5=100, a 10x increase, purely from the latency change with no change in incoming traffic. If your thread or connection pool was sized for ~10-20 concurrent in-flight requests to that dependency, it's now exhausted, and requests start queueing on your side, which is exactly the symptom described.
Immediate mitigations (minutes, not a redesign):
| Mitigation | What it does | Trade-off it introduces |
|---|---|---|
| Tight timeouts | Caps how long you'll wait, preventing unbounded queue growth | Cuts off requests that might have succeeded a moment later; needs to be shorter than your own SLA to the caller |
| Circuit breaker | Stops calling the dependency once error/latency crosses a threshold, failing fast instead of queueing | Can trip on transient blips if thresholds are too sensitive; denies service even to calls that might succeed |
| Bulkhead (isolated pool) | Gives this dependency its own thread/connection pool so its slowdown can't exhaust pools shared by healthy dependencies | Reduces pooled efficiency (can't borrow capacity across dependencies); requires knowing sizing up front |
| Load shedding / fast 503 | Rejects excess requests immediately when queue depth crosses a threshold, protecting the instances still healthy | Directly reduces availability for shed requests; needs to shed selectively, not randomly, if some requests matter more |
Longer-term architectural fix:
- Decouple via an async queue: put a durable queue between the caller and the slow dependency so the caller can return quickly (accept-and-acknowledge) and the dependency is drained at its own sustainable pace, rather than the caller blocking on it synchronously. Trade-off: the caller can no longer return a synchronous success/failure for that operation; the interaction model has to change to something the client and product can tolerate (a "pending" state, a webhook, a poll).
- Idempotent retries with backoff and jitter: if retries are needed, they must be capped, exponential, and jittered so a fleet of callers doesn't retry in lockstep and re-create the exact overload it's recovering from. Trade-off: added complexity, and retries must be provably idempotent on the downstream side or they risk duplicate side effects.
- Capacity planning against the tail, not the average: provision the dependency (or the pool sized to call it) based on observed p99 latency, not p50, since it's the tail that determines when queues start building. Trade-off: costs more standing capacity for headroom that's idle most of the time.
Applying this to concrete variants of the same pattern: the reasoning above is the same whether the slow dependency is a payment-validation service (immediate: circuit breaker + fast-fail with a clear "try again" to the user rather than a silent hang; long-term: async payment confirmation via webhook), a message-queue consumer falling behind (immediate: shed or dead-letter the oldest low-priority messages, bulkhead the consumer pool by message type; long-term: scale consumers horizontally and partition by priority), a retry storm from a flood of client-side 503s (immediate: the client-side backoff-with-jitter above is the direct fix; long-term: make the shedding threshold adaptive so it doesn't itself become the trigger for a thundering herd), or a synchronous order-processing pipeline backing up (immediate: bulkhead the slow stage's pool; long-term: convert that stage to the async-queue pattern above).
Trade-offs & pitfalls
- Every immediate mitigation above trades some availability or correctness for stability: timeouts drop requests that might have succeeded, circuit breakers deny service during their open window, load shedding sacrifices some requests to save the rest. The point isn't to avoid the trade-off, it's to make it deliberately and visibly rather than let an unbounded queue make it for you via an eventual crash.
- A common wrong turn: adding retries as the first response to a slowdown. Naive retries without backoff amplify load on an already-struggling dependency and can turn a partial slowdown into a full outage (a retry storm).
- Circuit breakers and bulkheads need to be tuned against real traffic and latency distributions; thresholds copied from a different service's runbook are a common source of either false trips (unnecessary unavailability) or no protection at all (thresholds too loose to matter).
When should failover be fully automated versus require a human to approve it? Walk through the factors that push you toward one or the other.
Sample Answer
Direct answer
Automate failover when the detector is high-precision, the failover action is reversible and idempotent, and the cost of a wrong automatic trigger is bounded and recoverable. Require a human when any of those breaks down, especially when a wrong trigger risks unrecoverable data divergence or an irreversible action. Expected-value math on detection accuracy alone favors automation more than intuition suggests, but it is reversibility and blast radius, not raw precision, that should gate the decision.
Structured elaboration
| Factor | Pushes toward automation | Pushes toward manual approval |
|---|---|---|
| Detection precision | High, multi-signal, correlated | Single noisy signal, history of false positives |
| Reversibility of the action | Fully reversible, idempotent | One-way (data promotion, DNS cutover with no clean undo) |
| Blast radius of a wrong trigger | Isolated to one service or region | Cross-service, cross-customer, or financial |
| Data consistency risk | Stateless or conflict-free (CRDT, idempotent) | Risk of split-brain (two nodes each independently believing they are the current leader, and both accepting writes at the same time, so the data silently diverges) or a double write |
| Regulatory or audit requirement | None, or satisfied by an audit log | Explicit approval-before-action mandate |
| Operational maturity | Tested runbooks, regular chaos drills | First time this failover path has been exercised |
A hybrid middle ground. Mature systems rarely pick one point on the automate-versus-manual spectrum. They tier it: automated detection and containment (circuit breakers, traffic throttling) run automatically because those actions are cheap to reverse, while the highest-blast-radius action (full regional failover, promoting a new primary) goes through an automated-detect, human-approve gate with an escalation timeout if nobody responds.
flowchart TD
A[Alert fires] --> B{Multi-signal, high-precision detector?}
B -->|No| M[Manual: page human, human confirms before failover]
B -->|Yes| C{Action reversible and idempotent?}
C -->|No| H[Hybrid: auto-detect and auto-contain, human approves full failover]
C -->|Yes| D{Wrong trigger risks split-brain or data loss?}
D -->|High risk| H
D -->|Low risk| E[Automate: auto-detect and auto-failover with fencing token and audit log]
A fencing token here is a number that increases with every failover action; if a stale, already-superseded actor (an old primary that thinks it's still in charge, for example) tries to act after a newer one has taken over, its writes carry an outdated token and get rejected, so a late-arriving action from a process that no longer should be acting can't silently corrupt state.
Framing it as expected value. For a given alert, the expected value of automatic failover is:
EVauto=p×value saved by faster RTO−(1−p)×cost of a false triggerwhere p is the detector's precision, the probability an alert reflects a real failure.
Worked example
Assume correct auto-failover cuts RTO from a 15-minute human-paged response to a 2-minute automatic one, a 13-minute improvement, against a downtime cost of $50k/hour:
value saved per true incident=6013×50,000=10,833A false trigger causes roughly 3 minutes of avoidable disruption (connection draining and reconnect storms) at the same rate:
cost per false trigger=603×50,000=2,500At a detector precision of p=0.9:
EV=0.9×10,833−0.1×2,500=9,750−250=9,500Solve for the breakeven precision where EV=0:
p×10,833=(1−p)×2,500 p=10,833+2,5002,500≈0.19Pure expected value favors automation down to a detector that is right only 19% of the time, far noisier than any detector actually deployed. That is the point: raw EV almost always says automate. The equation treats every false trigger as a bounded $2,500 cost, which is only true if the action is reversible. If a wrong trigger can cause split-brain or an irreversible data promotion, the real cost of that tail case is not in the equation at all, which is why reversibility, not precision, is the dominant factor in practice.
Trade-offs & pitfalls
- The most common wrong turn is optimizing for detector precision and stopping there; a 99%-precision detector triggering an irreversible action is still a bad automation candidate if the 1% case is catastrophic.
- Automating containment (throttle, circuit-break) before automating the full failover captures most of the RTO benefit with much lower blast radius; teams often skip straight to automating the whole failover and take on risk they did not need.
- An approval gate with no timeout just becomes a slower manual failover with extra steps; if a human stays in the loop, define an explicit escalation timeout.
- Chaos-testing the automated path before trusting it in production is not optional. An automation that has never been exercised against a real failure is a new, untested failure mode, not a safety net.
What's the difference between redundancy and replication when it comes to service reliability? Walk through an example for a stateless service and a stateful service, and name a failure mode that redundancy alone doesn't protect against for the stateful one.
Sample Answer
Direct answer: Redundancy is having extra, interchangeable components standing by so one can take over if another fails; replication is actively keeping copies of state (data) synchronized across multiple nodes so the state itself survives a failure, not just the compute that serves it. They're often used together, but they solve different problems: redundancy alone is enough for a stateless service, because any interchangeable instance can serve any request; a stateful service needs replication too, because a fresh redundant instance with no data isn't actually a working replacement.
Structured elaboration
| Aspect | Redundancy (stateless) | Replication (stateful) |
|---|---|---|
| What's duplicated | Compute/serving capacity | Data/state itself |
| Failover requirement | Route traffic to a healthy instance; done | Promote a replica that has the data, and ensure it's sufficiently up to date |
| Consistency concern | None, any instance is interchangeable | Central concern: how in-sync are the replicas at failover time |
| Typical mechanism | Load balancer + auto-healing instance group | Leader-follower or multi-leader data replication |
Stateless example: a set of identical app server instances behind a load balancer, handling API requests with no local state. If one instance dies, the load balancer routes around it and an autoscaler replaces it; the new instance needs no data transfer because there was never any instance-local state to lose. This is pure redundancy: extra interchangeable copies of the same stateless computation.
Stateful example: a primary-replica database. The primary accepts writes; replicas continuously receive a copy of the write stream (replication). If the primary fails, a replica is promoted to take over. Unlike the stateless case, simply having an extra database instance running (redundancy alone, no replication) would give you an empty database, not a working replacement, because there's no mechanism copying the actual data into it.
A failure mode redundancy alone doesn't protect against, for the stateful case: data loss or corruption on the primary itself. If the primary's disk corrupts a row, or a bad write silently corrupts application-level data, having a redundant (but not yet caught-up, or synchronously replicating that same bad write) standby doesn't help, because either the standby doesn't have the data yet (async lag) or it faithfully replicated the corruption along with everything else (synchronous replication of a logically bad write). Redundancy protects against a node dying; it does not protect against the data itself being wrong, that requires backups (a separate, point-in-time copy decoupled from live replication) and, for silent corruption specifically, checksums or application-level validation.
How this generalizes: a useful mental checklist for fault tolerance covers five distinct techniques, and redundancy and replication are only two of them: retries (recover from a transient failure by trying again), bulkheads (isolate one failure from spreading to unrelated resources), failover (the mechanism that switches traffic to a healthy replacement), redundancy (having that replacement exist at all), and replication (making sure the replacement actually has the state it needs). A strong answer names which of these a given design decision is actually addressing, since "redundancy" gets used loosely to mean all five in casual conversation.
Trade-offs & pitfalls
- Replication has a cost redundancy alone doesn't: network bandwidth, storage for extra copies, and a consistency model to reason about (synchronous replication costs write latency; asynchronous replication risks data loss on failover, the classic RPO trade-off).
- A common wrong turn: assuming "we have 3 replicas" automatically means "we're protected," without checking replication lag. A replica that's minutes behind at failover time silently loses however much data arrived in that window, unless the promotion logic explicitly accounts for lag and refuses to promote a too-far-behind replica.
- Redundancy for stateless services is comparatively cheap and low-risk to over-provision; replication for stateful services is not, since more replicas means more write-path coordination overhead (for synchronous replication) or more divergence risk (for asynchronous/multi-leader), so it isn't a "just add more" lever in the same way.
Design the failure detection that decides when to trigger an automated failover for a critical service. What health signals would you check, how would you set thresholds and windows to avoid mistaking a blip for a real failure, and when would you still want a human in the loop instead of a fully automatic failover?
Sample Answer
Direct answer: I'd layer health signals from cheap-and-fast (process liveness) to expensive-and-meaningful (dependency and business-metric checks), require multiple consecutive failures before declaring a node unhealthy to avoid reacting to a single blip, and keep a human in the loop specifically for the step that's hardest to reverse: promoting a new primary for stateful services, where an incorrect automated failover can cause data loss or a split-brain (two nodes each believing they're the one true primary and accepting conflicting writes at the same time), versus something like removing an unhealthy instance from a load-balancer pool, which is cheap to reverse and safe to fully automate.
Structured elaboration
| Signal | What it catches | Suggested check interval | Failures needed before acting |
|---|---|---|---|
| Process liveness | Process crashed or hung | Every 5s | 3 consecutive (15s) |
| Readiness / dependency connectivity | Process is up but can't reach its DB, cache, or queue | Every 10s | 2 consecutive (20s) |
| Latency / error-rate threshold | Process is up and connected, but degraded (slow, erroring) | Every 10-15s, rolling window | Sustained breach over a window (e.g., p95 > threshold for 3 samples), not a single sample |
| Business-metric sanity | Everything upstream looks healthy but the service is doing something wrong (e.g., checkout success rate collapsed) | Every 30-60s | Requires a real threshold breach, not a spike; slower-moving signal used as a final gate |
Why "N consecutive failures" instead of one: a single failed check can be a genuine blip (a GC pause, a brief network hiccup) rather than a real failure. If each check independently has some baseline flakiness probability p (a transient failure unrelated to a real outage), then requiring N consecutive failures before declaring unhealthy makes the false-trigger probability:
P(false trigger)=pNWith, say, p=0.05 (5% chance any single check fails transiently) and N=1 (react on the first failure), the false-trigger probability is just p=5%, meaning roughly 1 in 20 blips would incorrectly trigger action. Requiring N=3 consecutive failures drops that to:
P(false trigger)=0.053=0.000125=0.0125%a 400x reduction, at the cost of a real failure now taking 3 check intervals (here, up to 15 seconds at a 5s interval) longer to detect instead of one. That's the actual dial being turned: detection speed versus false-positive rate, and it should be set from an observed flakiness rate for your specific checks, not copied from another team's runbook.
Decision flow, including where automation stops and a human is required:
flowchart TD
A[Health check runs] --> B{N consecutive<br/>failures?}
B -->|No| A
B -->|Yes| C{What kind of<br/>action?}
C -->|Remove from LB pool| D[Fully automated:<br/>cheap, instantly reversible]
C -->|Open circuit breaker| D
C -->|Promote new primary<br/>for stateful service| E{Safety checks pass?<br/>replica caught up,<br/>quorum reachable}
E -->|No| F[Page human,<br/>do not auto-promote]
E -->|Yes, and blast radius<br/>is well-understood| G[Auto-promote,<br/>but page for review]
E -->|Yes, but ambiguous<br/>e.g. partition, not clear failure| F
"Quorum reachable" in that safety check means enough replicas are online and able to vote that a new primary can be safely elected, a majority agreeing on who's in charge, without risking two nodes each believing they're the primary at once.
Trade-offs & pitfalls
- Fully automated failover is right when the action is cheap and reversible (removing an unhealthy node from rotation); it's risky when the action is expensive or irreversible (promoting a database replica, since promoting the wrong one, or promoting during a network partition rather than an actual failure, can cause split-brain or data loss). The line isn't "how critical is the service," it's "how reversible is this specific action."
- Requiring consecutive failures trades detection speed for false-positive protection; too aggressive a requirement (e.g., N=10) means a genuine failure runs uncaught far longer than the blast radius justifies.
- Checks that themselves depend on a shared resource (e.g., every health check queries the same central database) can produce correlated, simultaneous "failures" across an entire fleet when that shared resource degrades, which looks like a mass outage but is really a single point of failure in the monitoring path itself.
- A common wrong turn: only checking process liveness and assuming that's sufficient. A process can be alive, passing liveness checks, and still be completely unable to serve real traffic because its only database connection pool is exhausted; readiness and dependency checks catch what liveness checks structurally cannot.
Unlock Full Question Bank
Get access to all Fault Tolerance, High Availability, and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.