Fault Tolerance, High Availability, and Disaster Recovery Questions
Keeping a system serving despite failure, from code-level resilience to infrastructure-level recovery: circuit breakers, retries with backoff and jitter, timeouts, bulkheads, graceful degradation, and preventing cascading failures, alongside redundancy, failover (active-active versus active-passive), RPO and RTO objectives, backup and restore, and multi-region failover. Covers dependency-failure isolation, chaos engineering to validate resilience, failure-mode analysis, designing to nines of availability, cost-versus-availability tradeoffs, and recovery runbooks. Spans both the patterns that isolate partial failure and the disaster-recovery planning that restores a business-critical system after a major outage.
Explain active-active versus active-passive architecture. For each, walk through the typical failover behavior, what it takes to detect a failure, and when you'd choose one over the other.
Sample Answer
Active-active runs all nodes or regions serving live traffic at the same time, so failover is mostly a routing problem: stop sending traffic to the unhealthy node. Active-passive keeps one side idle (or partially warmed) as a standby, so failover is a promotion problem: detect the primary is down, make the standby the new primary, then redirect traffic to it. That difference in what has to happen during failover is what drives everything else: recovery speed (RTO, recovery time objective: how long it takes to restore service after a failure), data-consistency risk (RPO, recovery point objective: how much data, measured in time, you could lose in a failure), and cost.
Comparing the two
| Dimension | Active-active | Active-passive |
|---|---|---|
| What's serving traffic | All nodes/regions, concurrently | Only the primary; standby is idle or warm |
| Failure detection | Health checks per node feed a load balancer or GSLB (Global Server Load Balancer: a DNS-based load balancer that routes traffic across regions, not just across servers in one place), which simply stops routing to the failed one | Health checks must trigger an explicit promotion decision, usually with a consensus/quorum step to avoid promoting during a false alarm |
| What "failover" does | Reroute traffic; no state transition needed | Promote replica to primary, update routing (DNS or LB config), then reroute traffic |
| Typical RPO | Near-zero if writes are synchronously replicated or conflict-resolved; otherwise bounded by replication lag | Near-zero with a synchronous standby, up to minutes with an async one |
| Consistency risk | Needs conflict resolution or partitioned ownership if writes happen on both sides (dual-writer problem) | Simpler: single writer at any point in time, no conflict resolution needed |
| Cost/complexity | Higher: full capacity running everywhere, plus distributed-write tooling | Lower: standby can run at reduced capacity (or be provisioned only at failover time) |
| Best fit | Latency-sensitive, globally distributed traffic; teams that can invest in multi-writer data patterns | Systems needing a single source of truth for writes (most transactional relational databases); cost-sensitive setups |
Worked example: deriving RTO for each
Pin the same detection policy to both: a health check runs every 5 seconds, and 3 consecutive failures are required before the system acts (15 seconds to declare a node down; this avoids reacting to a single dropped probe).
Active-active RTO: once the node is declared down, the load balancer or GSLB removes it from rotation immediately (no promotion step). If we assume that rotation update takes about 5 seconds to propagate to all edge/LB nodes:
RTOactive-active≈15s (detect)+5s (reroute)=20 secondsActive-passive RTO: the same 15 seconds to detect, plus a promotion step (electing the standby, replaying any un-applied log entries, opening it for writes, say 20 seconds for a warm standby with low replication lag) plus DNS or routing propagation (assume a low TTL of 30 seconds, and worst case the full TTL has to expire before every client picks up the change):
RTOactive-passive≈15s (detect)+20s (promote)+30s (routing propagation)=65 secondsUnder these pinned assumptions, active-active recovers about 3x faster, entirely because it skips the promotion step, and that gap only grows if the standby is cold rather than warm (add provisioning time) or if the routing layer is DNS with a high TTL instead of a fast health-checked LB.
Trade-offs and pitfalls
Active-active's speed advantage is real but not free: the moment two regions can both accept writes to the same record, you have a distributed-write problem, and skipping it (assuming replication will "just" reconcile) is the most common wrong turn. It needs either a conflict-resolution strategy (last-write-wins with vector clocks, CRDTs) or partitioned ownership (each region owns a disjoint key range) to avoid silently losing or corrupting data during a partition. Active-passive's failure mode is different: over-eager health checks or a flapping network link can trigger a premature promotion while the old primary is still technically reachable by some clients, producing two nodes that both believe they're primary (split-brain), which is why production active-passive systems add fencing (forcibly cutting off the old primary's access to shared storage) on top of the detection logic above, not just a timer. The same trade-off shows up outside a classic web/DB stack too: failing over an ML-serving fleet is closer to active-active in spirit (multiple replicas of the same model serving concurrently, so losing one just drops capacity) unless the model itself is being hot-swapped, in which case the promotion-style risk (serving from a half-loaded or stale model version) reappears.
Here's a simple architecture: a single load balancer, three identical application servers behind it, and one primary database instance handling all writes. Walk through it and identify the single points of failure. For each one, what would you do about it, and what does that cost you?
Sample Answer
This architecture has three single points of failure once you look past the app tier: the load balancer, and the primary database, are both singletons that the whole request path depends on; the app-server tier looks redundant on paper (three instances) but is only actually redundant if those three instances sit in different fault domains, so it's worth confirming rather than assuming.
Walking the diagram
flowchart LR
U[Users] --> LB[Load Balancer\nSINGLE instance]
LB --> A1[App Server 1]
LB --> A2[App Server 2]
LB --> A3[App Server 3]
A1 --> DB[(Primary DB\nSINGLE writer)]
A2 --> DB
A3 --> DB
Load balancer (single instance). Every request passes through it, so its failure is a total outage regardless of how healthy the three app servers are behind it. Mitigation: run an active-active pair of LB nodes behind a floating IP or DNS-based failover, or use a managed cloud load balancer where the provider owns that redundancy. Cost: a small amount of extra infrastructure and configuration; the bigger cost is usually operational (health-check tuning, avoiding split traffic during LB failover), not dollars.
App servers (three instances, conditionally redundant). If all three run in the same rack, same availability zone, or share an underlying host, they aren't actually independent, a single power or network event takes out all three at once. Mitigation: spread them across at least two, ideally three, availability zones and confirm the LB health-checks each independently and routes around a dead one automatically. Cost: cross-AZ data transfer costs and slightly higher latency on some requests; this is usually the cheapest SPOF to fix since it's mostly a placement decision, not new infrastructure.
Primary database (single writer, no replica). This is the highest-blast-radius SPOF: if it fails, every write path is down and, depending on the failure mode, recent unreplicated data can be at risk. Mitigation: add at least one synchronous or semi-synchronous replica in another AZ with automated failover (promote-on-failure), plus continuous backups for protection against logical corruption that replication alone wouldn't catch. Cost: this is the most expensive fix of the three, both in infrastructure (a standing replica) and in write latency if replication is synchronous.
Worked example: quantifying the SPOFs
Assume, for illustration, per-component annual availability of 99.95% for the load balancer, 99.9% for each app-server instance, and 99.9% for the database, all figures pinned as inputs for this calculation, not measured facts about any real vendor.
The three-server app tier, if truly independent across fault domains, only fails when all three fail simultaneously, so its unavailability multiplies:
1−Aapp tier=(1−0.999)3=(0.001)3=10−9That's an app-tier availability of essentially 99.9999999%, negligible. But the LB and DB are each in series with the whole request path (either one being down takes the whole system down), so their unavailabilities add through multiplication of the availabilities:
Aoverall=ALB×Aapp tier×ADB≈0.9995×(1−10−9)×0.999≈0.998500(99.850%)That's about 525,600×(1−0.998500)≈788.4 minutes of downtime per year, roughly the sum of the LB's own downtime (about 262.8 min/yr at 99.95%) and the DB's own downtime (about 525.6 min/yr at 99.9%), because the redundant app tier contributes essentially nothing to the failure budget while the two singletons dominate it completely. This is the concrete version of "fix the SPOFs first": no amount of extra app-server redundancy moves that 788-minute number until the LB and DB are addressed.
Trade-offs and pitfalls
The most common wrong turn is stopping at "add more app servers," which is the SPOF that's already effectively solved in this diagram and contributes the least to the real number above; teams do this because it's the cheapest, least disruptive change, not because it's the highest-leverage one. A second pitfall is fixing the database with synchronous cross-region replication by default: it does reduce RPO to near zero, but the added write latency (and reduced availability during a partition, since a strict quorum, requiring a majority of replicas to agree before a write is accepted, can block writes when too few replicas are reachable) is often the wrong trade for a service that would have been fine with an in-region synchronous replica and async cross-region for disaster recovery only. The same "look for the singleton" walk generalizes past this exact diagram: in a streaming ingestion pipeline the SPOF is usually a single partition leader or a schema registry with no standby; in an ML-serving stack it's a lone model server or a feature store with no fallback; and at a more abstract level, any shared control plane (service discovery, config store, secrets manager) or shared cache that every downstream service depends on is a SPOF even when nobody draws it on the diagram.
What is chaos engineering, and why would a company deliberately break its own production systems on purpose? Walk through the basic methodology: how you'd define steady state, form a hypothesis, and run a safe first experiment.
Sample Answer
Chaos engineering is the practice of deliberately injecting failure into a system, in a controlled way, to find weaknesses before they find you during a real incident. The reasoning behind doing it on purpose: most production failures aren't hypothetical, dependencies do time out, nodes do crash, networks do partition, and the choice isn't between "failures happen" and "failures don't happen," it's between discovering how your system responds to them during a planned, low-stakes experiment or during an unplanned, high-stakes 3 a.m. page.
Methodology
1. Define steady state. Pick measurable indicators of normal health, request success rate, latency percentiles, throughput, that represent "the system is working" in terms an on-call engineer would actually check on a dashboard, not an abstract notion of "healthy."
2. Form a hypothesis. State, before running anything, what you expect to happen and why: "if we kill one instance of the recommendation service, overall page error rate will stay flat because the client has a fallback path." A real hypothesis is falsifiable; "let's see what happens" isn't chaos engineering, it's just causing an outage without a way to learn from it.
3. Design a safe first experiment. Choose the smallest fault that could test the hypothesis (kill one non-critical replica, not the whole fleet) and decide the blast radius up front: what fraction of traffic or users can be affected, and for how long.
4. Run it with an abort condition already defined. Before starting, decide the exact metric threshold that ends the experiment immediately (for example, page error rate exceeding a set ceiling), so the decision to stop isn't made under pressure in the moment.
5. Observe against the steady-state baseline. Watch the same metrics defined in step 1, not new ones invented mid-experiment, so the comparison is apples-to-apples.
6. Learn and iterate. If the hypothesis held, expand the blast radius gradually on future runs. If it didn't, that's the actual finding, fix the missing fallback or retry logic, and re-run the same experiment to confirm the fix works before calling it done.
Worked example: a first, safe experiment
Target: a non-critical "related items" widget on a product page, deliberately chosen because a broken hypothesis here degrades a widget, not checkout. Steady state: page load success rate and p95 latency, whatever their current normal values are for that page. Hypothesis: "terminating one replica of the related-items service will not change page load success rate or p95 latency, because the front end treats that service as optional with a client-side timeout and empty-state fallback." Experiment: kill one replica (not all of them) during a low-traffic window, with an abort condition of "page success rate drops below its normal range" defined before starting. Outcome either confirms the fallback works as designed, or reveals it doesn't, which is the actual value of running it: finding that out on a Tuesday afternoon experiment instead of during a real node failure at peak traffic.
Trade-offs and pitfalls
The most common misunderstanding is that chaos engineering means "randomly break things in production," when the entire method is built around the opposite instinct: a stated hypothesis, a bounded blast radius, and a predefined abort condition are what separate a chaos experiment from just causing an outage. A related pitfall is skipping the hypothesis step and injecting a fault "to see what happens": without a stated expectation, there's no way to say afterward whether the result was surprising or how bad it was relative to what should have happened. Teams also sometimes skip straight to production chaos before validating the tooling and abort mechanism in staging first, running the injection and rollback machinery against a stage environment is itself a smaller, safer experiment worth doing before trusting it against real traffic.
What's the bulkhead pattern, and how does it stop one failing dependency or noisy tenant from taking down the whole system? Give a concrete example of where you'd draw the isolation boundary.
Sample Answer
Direct answer
The bulkhead pattern partitions a system's resources (thread pools, connection pools, CPU, or entire nodes) into isolated compartments, named after a ship's watertight bulkheads, so that one failing dependency or one noisy tenant can only exhaust the resources in its own compartment, not the resources every other caller depends on. Without bulkheads, a single slow or misbehaving dependency can consume every available thread or connection in a shared pool, and a completely healthy code path fails simply because it couldn't get a thread to run on.
Where to draw the isolation boundary
A concrete example: an API gateway calls three downstream services, an inventory service, a recommendations service, and a payments service, all through one shared thread pool. If recommendations starts responding slowly, every thread in the shared pool eventually ends up blocked waiting on recommendations calls, and inventory and payment requests start timing out too, even though nothing is wrong with either of them. The fix is a dedicated, bounded thread pool (or connection pool) per downstream dependency: recommendations gets its own pool of, say, 10 threads, so a recommendations outage can stall at most those 10 threads and its own queue, while inventory and payments keep running normally on their own separate pools.
The boundary should sit wherever one caller's failure or slowness shouldn't be able to spill onto another caller's request. Common places to draw it:
- Per-downstream-dependency, as in the example above: each external service or database gets its own pool so a slow one can't starve calls to a fast one.
- Per-tenant, in a multi-tenant system: each tenant (or tenant tier) gets a capped share of connections or CPU so one noisy or abusive tenant can't degrade service for everyone else on shared infrastructure.
- Per-criticality-tier: payment and auth paths get reserved capacity separate from lower-priority paths like analytics or notifications, so a spike in low-priority traffic can't crowd out the paths that actually matter.
Trade-offs & pitfalls
Bulkheads trade utilization for isolation: reserved capacity that a compartment isn't currently using sits idle rather than being available to a busier compartment, so a poorly sized bulkhead can cause localized throttling even while the system as a whole has spare capacity. Sizing is the actual hard part in practice, not the pattern itself: too small and a legitimate burst of normal traffic gets rejected by its own bulkhead; too large and the isolation becomes theoretical, because if every pool is sized close to the shared pool's original total, a single compartment can still consume enough of the machine's real resources (CPU, memory, file descriptors) to degrade its neighbors even though the pool counters look fine. Bulkheads are also a different tool from a circuit breaker and the two are frequently confused: a bulkhead limits how much of a shared resource one dependency can consume (a capacity boundary), while a circuit breaker stops sending requests to a dependency once it's clearly failing (a decision to stop calling at all); they're complementary, since the bulkhead caps the damage while the circuit breaker is deciding whether to keep trying, and production systems typically use both on the same dependency together. The same reasoning extends beyond web request threads: an ML-serving platform running GPU inference for multiple models on shared hardware applies the identical idea by pinning each model (or tenant) to a dedicated slice of GPU memory and compute, so one model that starts issuing runaway-batch-size requests can't starve GPU capacity away from every other model sharing that hardware.
For availability targets of 99.9%, 99.99%, and 99.999%, calculate the allowed downtime per year and per month for each. Then walk through what architectural changes actually get you from one tier to the next.
Sample Answer
Direct answer: Availability is the fraction of time a system is usable, and "N nines" is shorthand for how close that fraction is to 100%. Going from 99.9% to 99.99% to 99.999% shrinks allowed downtime by roughly 10x at each step, and each step also costs roughly an order of magnitude more in engineering and infrastructure, because you're eliminating an entire category of failure (single-host, then single-zone, then single-region) rather than just adding more of the same redundancy.
Structured elaboration
Allowed downtime per year is derived from the availability target directly:
downtimeyear=(1−A)×8760 hourswhere 8760 is the number of hours in a 365-day year (24 x 365), and per-month downtime uses 730 hours (8760 / 12):
downtimemonth=(1−A)×730×60 minutesPlugging in each target:
| Availability | Allowed downtime / year | Allowed downtime / month |
|---|---|---|
| 99.9% ("three nines") | (1−0.999)×8760=8.76 hours -> 8h 45m 36s | (1−0.999)×730×60=43.8 min |
| 99.99% ("four nines") | (1−0.9999)×8760=0.876 hours -> 52m 34s | (1−0.9999)×730×60=4.38 min |
| 99.999% ("five nines") | (1−0.99999)×8760=0.0876 hours -> 5m 15s | (1−0.99999)×730×60=0.438 min -> 26s |
Each jump divides allowed downtime by exactly 10, because each availability target divides (1−A) by 10.
What actually changes architecturally between tiers
- 99.9% -> 99.99%: eliminate single points of failure inside one facility. Multi-AZ deployment, N+1 redundancy (one extra standby unit beyond what's strictly needed to handle normal load, so a single failure doesn't drop capacity below what's required) on stateful components (load balancers, databases with a standby replica), automated health-check-driven failover, and a real on-call rotation with paging. Most of the gain here comes from removing manual recovery steps: a human restarting a service takes minutes and that alone can burn the entire four-nines monthly budget.
- 99.99% -> 99.999%: eliminate the facility (zone or region) itself as a single point of failure. Multi-region active-active or hot standby (a fully-running backup kept ready to take over instantly, unlike a cold standby that would first need to be started up and warmed), automated cross-region failover (not human-triggered), synchronous or tightly-bounded-lag replication for the data that must survive a region loss, and rigorous testing of the failover path itself (chaos drills: deliberately triggering the failover in a controlled test so a broken failover path is discovered on a Tuesday afternoon, not during a real outage), because at this tier the failover mechanism is now a bigger risk to availability than the failures it's protecting against.
- Beyond 99.999%, the limiting factor usually isn't infrastructure, it's deployment risk (bad releases) and dependency risk (a vendor or DNS provider you don't control), so the remaining budget goes to progressive rollouts, fast automated rollback, and reducing the number of hard external dependencies on the critical path.
Worked example: composing a dependency chain
A request that serially depends on a load balancer (99.99%), an app tier (99.95%), and a database (99.99%) has a combined availability equal to the product of the individual availabilities, because all three must be up simultaneously:
Aserial=0.9999×0.9995×0.9999=0.9993That's roughly 99.93%, worse than any single component, which is why a system built entirely from 99.99%-rated pieces chained together does not automatically deliver 99.99% end to end. Adding a redundant standby database (parallel, either one being up is sufficient) with independent 99.99% availability changes only that term:
Adb,pair=1−(1−0.9999)2=1−0.00012=0.99999999so the pair is effectively always up, and the chain's availability is then bounded by the weakest remaining serial link (the app tier at 99.95%), not the database.
Trade-offs & pitfalls
- Availability composes multiplicatively across a serial chain and the weakest link dominates: chasing five nines on your database while your app tier sits at three nines is wasted spend.
- Each nine costs disproportionately more: 99.9% to 99.99% is mostly process and automation (cheap-ish); 99.99% to 99.999% usually means paying for a second region and the operational overhead of keeping it truly independent (expensive, and dangerous if the failover path itself is untested).
- Downtime budgets don't distinguish planned from unplanned; a team that spends its whole error budget on deploy-related outages hasn't actually built a more resilient system, just a riskier release process.
- A very common interview trap: treating "99.99% uptime" as a promise about any single request rather than a time-integrated average. A system can meet 99.99% for the year while having a full 50-minute outage in one bad afternoon.
Unlock Full Question Bank
Get access to all 16 Fault Tolerance, High Availability, and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.