Fault Tolerance, High Availability, and Disaster Recovery Questions

Keeping a system serving despite failure, from code-level resilience to infrastructure-level recovery: circuit breakers, retries with backoff and jitter, timeouts, bulkheads, graceful degradation, and preventing cascading failures, alongside redundancy, failover (active-active versus active-passive), RPO and RTO objectives, backup and restore, and multi-region failover. Covers dependency-failure isolation, chaos engineering to validate resilience, failure-mode analysis, designing to nines of availability, cost-versus-availability tradeoffs, and recovery runbooks. Spans both the patterns that isolate partial failure and the disaster-recovery planning that restores a business-critical system after a major outage.

MediumTechnical
86 practiced

Explain active-active versus active-passive architecture. For each, walk through the typical failover behavior, what it takes to detect a failure, and when you'd choose one over the other.

MediumTechnical
86 practiced

Here's a simple architecture: a single load balancer, three identical application servers behind it, and one primary database instance handling all writes. Walk through it and identify the single points of failure. For each one, what would you do about it, and what does that cost you?

EasyTechnical
76 practiced

What is chaos engineering, and why would a company deliberately break its own production systems on purpose? Walk through the basic methodology: how you'd define steady state, form a hypothesis, and run a safe first experiment.

EasyTechnical
81 practiced

What's the bulkhead pattern, and how does it stop one failing dependency or noisy tenant from taking down the whole system? Give a concrete example of where you'd draw the isolation boundary.

EasyTechnical
70 practiced

For availability targets of 99.9%, 99.99%, and 99.999%, calculate the allowed downtime per year and per month for each. Then walk through what architectural changes actually get you from one tier to the next.

Unlock Full Question Bank

Get access to all 16 Fault Tolerance, High Availability, and Disaster Recovery interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.