Fault Tolerance, High Availability, and Disaster Recovery Questions
Keeping a system serving despite failure, from code-level resilience to infrastructure-level recovery: circuit breakers, retries with backoff and jitter, timeouts, bulkheads, graceful degradation, and preventing cascading failures, alongside redundancy, failover (active-active versus active-passive), RPO and RTO objectives, backup and restore, and multi-region failover. Covers dependency-failure isolation, chaos engineering to validate resilience, failure-mode analysis, designing to nines of availability, cost-versus-availability tradeoffs, and recovery runbooks. Spans both the patterns that isolate partial failure and the disaster-recovery planning that restores a business-critical system after a major outage.
What does a disaster recovery runbook actually need to contain to be useful during a real region failure? Walk through the essential sections: owner, RTO/RPO, step-by-step actions, and verification.
Sample Answer
Direct answer
A disaster recovery runbook is only useful if a stressed engineer can execute it top to bottom without needing to look anything else up. That means naming an owner, stating the RTO and RPO it is designed to meet, listing ordered and specific actions (exact commands or console steps, not descriptions), and ending with concrete verification steps that prove the system is actually back, not just that the steps were followed.
Structured elaboration
| Section | What it contains | Why it's required |
|---|---|---|
| Title and scope | Which system or service, and which failure modes it covers | A runbook that does not say what it's for gets grabbed for the wrong incident |
| Owner and escalation contacts | Primary owner, backup owner, and how to reach them, not just a name | Someone must be accountable for keeping it accurate and reachable during the incident |
| RTO / RPO | The target time to recover and the acceptable data loss this runbook is designed to hit | Without a target, "did the runbook work" has no answer |
| Prerequisites | Required access, credentials, tickets, and any upstream dependency that must already be healthy | Discovering you're locked out mid-incident is the worst time to find out |
| Step-by-step actions | Numbered, specific commands or console actions, including a rollback for each risky step | Vague steps like "promote the standby" force the responder to improvise under pressure |
| Verification steps | Health checks, smoke tests, and the specific metrics that confirm recovery | "The steps finished" is not the same as "the system works" |
| Post-incident tasks | Root cause capture, stakeholder communication, runbook update | A runbook that isn't updated after every real use rots |
Versioning and access. Store the runbook in source control with required review on changes, so every edit has an author and a diff. Test it in a real drill, not a tabletop discussion only, on a cadence tied to how critical the service is, quarterly for anything customer-facing. Keep it reachable when the primary systems it recovers are down: a runbook that lives only on an internal wiki hosted in the region that just failed is not a disaster recovery runbook.
Worked example
For a service with RTO = 30 minutes, a well-built runbook's step timings should sum to that budget, and the sum should be checked, not assumed:
| Step | Budget |
|---|---|
| Detection and paging | 5 minutes |
| Triage and decision to fail over | 5 minutes |
| Execution (promote standby, update routing) | 15 minutes |
| Verification (smoke tests, dashboards green) | 5 minutes |
That sum matching the stated RTO exactly is what makes the RTO a testable claim rather than a number pasted at the top of the document. If a quarterly drill shows execution consistently takes 20 minutes instead of 15, the runbook's RTO is wrong and needs to be corrected, not explained away.
Trade-offs & pitfalls
- A runbook with no owner drifts out of date the first time the architecture changes; ownership is not optional metadata.
- Testing via tabletop discussion only, never an actual drill, hides the gap between the steps sounding right and the steps working; most drift is caught only by execution.
- Over-specifying every command for a fast-moving system creates a maintenance burden that causes the runbook to be abandoned; balance specificity against how often the underlying commands change.
- Storing the only copy behind the same authentication system that depends on the region that just failed is a common, self-defeating mistake.
A downstream service you depend on starts responding slowly, and requests to it start backing up on your side, growing queues and increasing latency. Walk through your immediate mitigations and your longer-term architectural fix, and explain the trade-off each one introduces.
Sample Answer
Direct answer: The immediate priority is to stop the slowdown from consuming your own resources: set aggressive timeouts, open a circuit breaker so you stop calling the failing dependency, and isolate the connection/thread pool used for that call so it can't starve everything else. The longer-term fix is architectural: decouple the caller from the dependency's latency entirely, usually via an async queue or by making the call non-blocking, so a slow downstream degrades throughput instead of taking the whole service down with it.
Structured elaboration
Why this happens (the mechanism): by Little's Law, the number of requests in flight L equals arrival rate λ times the time each request spends in the system W: L=λW. If a downstream call's latency goes from 50ms to 500ms while your request rate stays at, say, 200 requests/second, the in-flight count grows from L=200×0.05=10 to L=200×0.5=100, a 10x increase, purely from the latency change with no change in incoming traffic. If your thread or connection pool was sized for ~10-20 concurrent in-flight requests to that dependency, it's now exhausted, and requests start queueing on your side, which is exactly the symptom described.
Immediate mitigations (minutes, not a redesign):
| Mitigation | What it does | Trade-off it introduces |
|---|---|---|
| Tight timeouts | Caps how long you'll wait, preventing unbounded queue growth | Cuts off requests that might have succeeded a moment later; needs to be shorter than your own SLA to the caller |
| Circuit breaker | Stops calling the dependency once error/latency crosses a threshold, failing fast instead of queueing | Can trip on transient blips if thresholds are too sensitive; denies service even to calls that might succeed |
| Bulkhead (isolated pool) | Gives this dependency its own thread/connection pool so its slowdown can't exhaust pools shared by healthy dependencies | Reduces pooled efficiency (can't borrow capacity across dependencies); requires knowing sizing up front |
| Load shedding / fast 503 | Rejects excess requests immediately when queue depth crosses a threshold, protecting the instances still healthy | Directly reduces availability for shed requests; needs to shed selectively, not randomly, if some requests matter more |
Longer-term architectural fix:
- Decouple via an async queue: put a durable queue between the caller and the slow dependency so the caller can return quickly (accept-and-acknowledge) and the dependency is drained at its own sustainable pace, rather than the caller blocking on it synchronously. Trade-off: the caller can no longer return a synchronous success/failure for that operation; the interaction model has to change to something the client and product can tolerate (a "pending" state, a webhook, a poll).
- Idempotent retries with backoff and jitter: if retries are needed, they must be capped, exponential, and jittered so a fleet of callers doesn't retry in lockstep and re-create the exact overload it's recovering from. Trade-off: added complexity, and retries must be provably idempotent on the downstream side or they risk duplicate side effects.
- Capacity planning against the tail, not the average: provision the dependency (or the pool sized to call it) based on observed p99 latency, not p50, since it's the tail that determines when queues start building. Trade-off: costs more standing capacity for headroom that's idle most of the time.
Applying this to concrete variants of the same pattern: the reasoning above is the same whether the slow dependency is a payment-validation service (immediate: circuit breaker + fast-fail with a clear "try again" to the user rather than a silent hang; long-term: async payment confirmation via webhook), a message-queue consumer falling behind (immediate: shed or dead-letter the oldest low-priority messages, bulkhead the consumer pool by message type; long-term: scale consumers horizontally and partition by priority), a retry storm from a flood of client-side 503s (immediate: the client-side backoff-with-jitter above is the direct fix; long-term: make the shedding threshold adaptive so it doesn't itself become the trigger for a thundering herd), or a synchronous order-processing pipeline backing up (immediate: bulkhead the slow stage's pool; long-term: convert that stage to the async-queue pattern above).
Trade-offs & pitfalls
- Every immediate mitigation above trades some availability or correctness for stability: timeouts drop requests that might have succeeded, circuit breakers deny service during their open window, load shedding sacrifices some requests to save the rest. The point isn't to avoid the trade-off, it's to make it deliberately and visibly rather than let an unbounded queue make it for you via an eventual crash.
- A common wrong turn: adding retries as the first response to a slowdown. Naive retries without backoff amplify load on an already-struggling dependency and can turn a partial slowdown into a full outage (a retry storm).
- Circuit breakers and bulkheads need to be tuned against real traffic and latency distributions; thresholds copied from a different service's runbook are a common source of either false trips (unnecessary unavailability) or no protection at all (thresholds too loose to matter).
You're using DNS failover with a 5-minute TTL, but in practice you're seeing a 3-minute real-world failover window, and it's too slow. How would you redesign this to get failover under 30 seconds for most clients, and what do you give up to get there?
Sample Answer
Direct answer
The 5-minute TTL isn't the actual bottleneck: DNS caching in the real world (ISP resolvers with minimum-TTL floors, browser caches, persistent keep-alive connections that never re-resolve at all) means real failover time doesn't track the advertised TTL cleanly, which is exactly why you're seeing 3 minutes instead of something close to 5. Getting under 30 seconds for most clients means building an explicit time budget (detect, update, propagate) that sums under 30s for compliant clients, and accepting that DNS alone can't guarantee it for the tail of clients whose resolvers or connections don't re-check in time.
Building the time budget
Tfailover≤(k×interval)+Tpush+TTLwhere k is the number of consecutive failed health checks required before failing over (the detection threshold) and interval is the health-check period.
sequenceDiagram
participant C as Client
participant R as Resolver
participant D as Authoritative DNS
participant H as Health Monitor
participant A as Origin A
participant B as Origin B
H->>A: probe every 5s
A--xH: 2 consecutive failures (10s)
H->>D: update record to B
C->>R: resolve hostname
R->>D: query (TTL expired)
D-->>R: return B, TTL 10s
R-->>C: B
C->>B: connect
Worked example
Redesign inputs: health-check interval 5s, failure threshold k=2 (avoids single-blip flaps), API-driven record push under 1s, and TTL lowered from 300s to 10s.
TdetectTpushTttlTfailover≤2×5s=10s≈1s=10s≤10+1+10=21s<30sThat covers clients and resolvers that honor the lowered TTL, with about 9 seconds of margin. It does not cover the two categories that caused the original 3-minute number: resolvers that enforce a minimum TTL floor above what you set, and clients holding a persistent connection that has no reason to re-resolve DNS at all until it errors. For those, add a client-side backstop that's independent of TTL: short keep-alive and idle timeouts so connections periodically re-establish (and therefore re-resolve), and connect-level retry to a secondary IP on failure (a Happy-Eyeballs-style fallback) rather than trusting DNS to be the only failover signal.
Trade-offs and pitfalls
What you give up: a 10s TTL multiplies authoritative DNS query volume roughly 30x versus the 300s baseline, which is a real cost and load increase on your DNS infrastructure, and a false-positive failover (from setting k too low) now flips production traffic in as little as 5 to 10 seconds, so your health check needs to be more conservative about what counts as "down," not less. The pitfall that caused the original bug is assuming all clients and resolvers honor your TTL uniformly; they don't, and any redesign that only lowers the TTL without a client-side or network-level backstop will hit the same wall for the same tail of misbehaving resolvers, just with a lower number attached to it.
Leadership wants to cut cloud costs by 30% without dropping below 99.99% uptime for critical services. Walk through how you'd find the savings and what you'd protect no matter what.
Sample Answer
Direct answer
Frame the cut through the 99.99% error budget, not a blanket percentage. 99.99% allows about 52.6 minutes of downtime a year; anything that doesn't touch the paths that consume that budget is safe to cut aggressively, and anything that does must be modeled against it explicitly before you touch it. In practice that means chasing waste and inefficiency hard (usually most of the 30%), and treating redundancy, failover paths, and DR test cadence as protected unless you can show the specific budget impact is acceptable.
A three-bucket framework for the savings
D99.99%=(1−0.9999)×525600 min=52.56 min/yrThat number is the currency every cut gets priced in.
| Bucket | Examples | Availability impact |
|---|---|---|
| 1. Waste elimination | Idle/overprovisioned instances, orphaned disks/snapshots/IPs, unscheduled non-prod environments | None, doesn't touch the serving or failover path |
| 2. Efficiency gains | Rightsizing with headroom preserved, reserved/spot capacity for stateless, replaceable workers, caching hot reads | Neutral to positive if done with canary rollout and autoscaling guardrails |
| 3. Structural changes | Fewer replicas, longer failover windows, single cloud/provider consolidation, reduced DR test cadence | Directly spends the 52.6 min/yr error budget, must be modeled before approval |
Chase bucket 1 and 2 first and aggressively (they rarely conflict with availability); only reach into bucket 3 if 1 and 2 don't get you to 30%, and price every bucket-3 cut against the budget above.
Worked example
Suppose an audit of monthly compute spend of $300,000 finds 15% idle or overprovisioned capacity, plus another 10% recoverable through rightsizing and reserved commitments on stateless, replaceable capacity:
wasterightsizingsubtotal=$300,000×0.15=$45,000/mo=$300,000×0.10=$30,000/mo=$45,000+$30,000=$75,000/mo=25% of spendThat's 25 of the 30 points from buckets 1 and 2 alone, with essentially zero availability risk. The remaining 5 points has to come from bucket 3, for example dropping a redundant standby from 3 independent paths to 2 on some service. Whether that's safe depends entirely on whether that service is in the 99.99% critical scope. Reusing the same series/parallel redundancy model above (an N-way redundant path's unavailability is the product of each independent unit's own unavailability, since all N units have to fail at the same time for the whole path to be down: UN=(1−a)N), for a building block at a=0.99:
N=3:N=2:U3=(0.01)3=0.000001⇒D3=0.000001×525600=0.5256 min/yrU2=(0.01)2=0.0001⇒D2=0.0001×525600=52.56 min/yrDropping that one path from N=3 to N=2 raises its expected downtime contribution from about half a minute a year (0.5256 min/yr) to 52.56 minutes a year, a 100x jump that would consume the entire annual budget for the 99.99% target on that one path alone. That's the concrete argument for why redundancy counts on in-scope critical services are protected regardless of the cost target: the math shows the cut doesn't save what it looks like it saves once you price the risk.
Trade-offs and pitfalls
The classic failure is applying a flat 30% cut across the board instead of segmenting critical from non-critical, which either misses easy wins in non-critical systems or, worse, quietly erodes redundancy on a critical path because nobody explicitly modeled the budget impact. Watch for Goodhart's-law style gaming too: cutting observability or alerting spend looks free on the invoice but raises mean time to detect, which inflates the effective downtime against the same budget without showing up as an "availability" line item until an incident hits. What to protect no matter what: replica or quorum counts below the tested minimum on in-scope services, cross-region failover paths, backup and DR test cadence, and on-call staffing, because all four either directly hold the redundancy math above or determine how fast you can react when it fails.
Design a multi-tenant platform so that a noisy tenant can't degrade availability for everyone else. Walk through your isolation and enforcement mechanisms, and how autoscaling and billing interact with the quotas you set.
Sample Answer
Direct answer
Isolate at three layers so no single control point being bypassed breaks the guarantee: admission control at the edge (reject or queue over-quota requests before they consume shared resources), hard resource enforcement at the OS/scheduler level (so a tenant that gets past admission still can't starve others), and fairness-aware autoscaling and billing (scale for real aggregate demand, allocate new capacity by weighted share rather than first-come-first-served, and meter overage so legitimate extra usage is paid for instead of stolen from neighbors).
Architecture and enforcement
flowchart LR
T[Tenant request] --> GW[Admission gateway: token bucket]
GW -->|within quota| ENQ[Per-tenant queue]
GW -->|over quota| REJ[429 / shed]
ENQ --> SCHED[Scheduler: cgroup CPU/IO caps]
SCHED --> POOL[Shared compute pool]
POOL --> AS[Autoscaler: per-tenant saturation signal]
AS --> POOL
POOL --> MET[Usage metering]
MET --> BILL[Billing: quota + overage]
Base quota per tenant is a weighted fair share of total capacity C:
Qi=C×∑jwjwi- Admission layer: per-tenant token bucket enforces request-rate and concurrency caps before work is queued; this stops a noisy tenant from ever occupying shared queue depth.
- Enforcement layer: the OS and scheduler cap what a tenant's workload can actually consume, CPU time, memory, and network bandwidth, even if a request slips past admission control, defense in depth rather than trusting one gate alone. (Advanced implementation detail: on Linux this is typically cgroups for CPU bandwidth limits, OOM-score tuning to control which processes get killed first under memory pressure, and network policers such as tc/ingress qdiscs to cap per-tenant bandwidth.)
- Autoscaling: the trigger must look at per-tenant saturation, not just cluster-wide CPU. If a single tenant is the one hammering the cluster, scaling out on aggregate cluster metrics effectively rewards that tenant with new capacity paid for by everyone, which defeats the isolation goal economically even if it holds technically.
- Billing: usage beyond the base quota is metered as overage rather than throttled outright for tenants on a plan that allows bursting; quota tier is the plan's contract, burst is a paid escape valve, not a loophole.
Worked example
Cluster capacity C=1000 vCPU across four tenants with weights A=3, B=1, C=1, D=1 (their plan tiers):
C=1000, ∑w=6:QAQB=QC=QD=1000×63=500=1000×61=166.67If tenant D is idle, its unused 166.67 vCPU can be reclaimed and redistributed proportionally among the active tenants (weights 3, 1, 1 summing to 5), rather than sitting wasted:
RA′B′C′=QD=166.67=QA+R×53=500+100=600=QB+R×51=166.67+33.33=200=QC+R×51=166.67+33.33=200A' + B' + C' + 0 = 1000, the full capacity, allocated fairly by weight instead of by who asked first. If D comes back online, the reclaim has to be preemptible (graceful, with a grace period) so D isn't hard-killed the moment it wakes up; it gets its guaranteed floor back and the borrowers scale down.
Trade-offs and pitfalls
Hard caps with no reclaim protect worst-case isolation perfectly but waste real capacity whenever a tenant is idle, which is most of the time for most tenants. Soft caps with reclaim use capacity efficiently but add real scheduler complexity, and the failure mode to avoid is yanking borrowed capacity abruptly instead of draining it, which turns "fairness" into a second source of noisy-neighbor-style disruption for the tenant who briefly borrowed it. The single most common wrong turn in this design is coupling the autoscaler's scale-out trigger directly to raw cluster-wide metrics: it looks like a capacity problem and "solves" it by adding nodes, but if the actual cause is one tenant blowing past its quota, autoscaling just subsidizes the noisy tenant at everyone else's expense until billing catches up, if it ever does.
Unlock Full Question Bank
Get access to all Fault Tolerance, High Availability, and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.