Disaster Recovery and Business Continuity Questions
Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.
What does a disaster recovery runbook actually need to contain to be usable by someone under pressure at 3am, and who should own keeping each section current? Walk through the essential sections and why each one earns its place.
Sample Answer
Direct answer
A usable 3 a.m. runbook has to work for someone who is tired, stressed, and not the person who wrote it, so every section exists to answer a question that person will actually have in the moment: am I authorized to act, exactly what do I do, how do I know it worked, and who do I call if it doesn't. Ownership has to be distributed across the people who actually know each piece, not centralized in whoever happened to draft the document, because a runbook only stays accurate if its owner is the one who'd notice when their piece goes stale.
Essential sections and why each earns its place
| Section | What it answers | Typical owner |
|---|---|---|
| Purpose, scope & triggers | What event this runbook is for, and the specific conditions that mean "use this one" | Service or function lead |
| Preconditions & prerequisites | What has to already be true or confirmed before starting (e.g., declaration made, backup integrity confirmed where relevant) | Continuity or incident commander |
| Roles & ownership for this runbook | Who executes, who verifies, who's the backup if the primary is unreachable | Engineering manager / function lead |
| Step-by-step procedure | Ordered, numbered actions specific enough for someone unfamiliar with the system to follow | The person closest to the recovery work, reviewed by their manager |
| Verification checks | Concrete, observable confirmation each major step actually worked, not just that it executed | Whoever owns that step |
| Rollback criteria & procedure | The specific conditions that mean "this isn't working, reverse it," and how | Release / change owner |
| Estimated per-step timelines | A rough expectation for how long each major step should take | Step owner, refined after each real use or exercise |
| Contact lists & escalation matrix | Who to reach, and at what elapsed time or procedural point each escalation level triggers | On-call / escalation manager |
| Communication templates | Pre-written internal and external status language for this scenario | Communications lead |
| Post-recovery actions | Root-cause capture, runbook update, scheduling the next exercise of this runbook | Incident commander |
A precondition section without a trigger stops someone from starting a recovery sequence that assumes a decision hasn't actually been made. A verification section separate from the step it checks exists because executing a step isn't the same as confirming it succeeded, and conflating the two is how partial recoveries get reported as complete. A per-step timeline exists so someone can tell "this is taking the expected amount of time" from "this step has stalled and needs escalation," a distinction that's very hard to judge from scratch under pressure.
Multiple runbook types, not one document
A mature program doesn't maintain a single generic "disaster runbook." It maintains several distinct runbooks, each using this same skeleton but with different triggers, sequencing, and named owners: one for a full data-center-loss scenario, with its own escalation and contact list because the responders differ; one specifically for a ransomware or security-incident scenario, where the preconditions section is doing much more work because it has to gate on backup-integrity confirmation before any restore step; one for a database primary-region outage specifically, where the contact-list and timeline sections matter even more, since escalation routes to the database team and expected step durations look different from other scenarios; and one for a critical third-party or vendor outage. Treating these as separate documents, rather than branches inside one bloated runbook, keeps each one short enough to actually follow under pressure. NIST's SP 800-34 contingency-planning guidance, written for U.S. federal information systems but widely used as a general template, reflects this same idea: multiple purpose-specific plans rather than one document trying to cover every scenario.
Maintenance ownership
Every section above has a named owner specifically because "the runbook" as a whole has no natural owner. Distributing ownership by section is what keeps individual pieces current as systems and people change, paired with a fixed review cadence tied to the service's criticality tier, the same logic as the exercise cadence, so staleness gets caught on a schedule rather than discovered during a real event.
Worked example: critical authentication service outage
- Purpose & scope: use this runbook when the primary authentication service is confirmed unreachable or failing above a defined threshold of login attempts.
- Preconditions: confirm this isn't a downstream dependency issue (check the identity provider's own status) before proceeding; confirm continuity declaration has occurred if the outage has crossed the Tier 1 trigger duration.
- Step-by-step (abbreviated): (1) engage the secondary authentication path per the platform team's documented failover procedure; (2) notify downstream service owners that authentication is degraded, using the pre-written template; (3) confirm with each downstream owner that their service functions against the backup path.
- Verification: confirm login success rate has returned to the expected baseline range, and confirm at least the top dependent services report normal operation, not just that the failover step completed.
- Rollback criteria: if the backup path itself shows elevated failures beyond a defined threshold, escalate to the platform on-call lead rather than continuing to push traffic through it.
- Escalation: unresolved after 30 minutes, escalate to the platform engineering manager; unresolved after 60 minutes, escalate to the incident commander for continuity-declaration review.
This mirrors the same skeleton the data-center-loss, ransomware, and database-outage runbooks use, with the specific triggers, steps, and named escalation owners swapped for what's actually true of authentication as a service.
Trade-offs and pitfalls
- A runbook centralized under one owner tends to rot, because no single person notices every section going stale. Section-level ownership is more work to coordinate but far more likely to stay accurate.
- Steps that are too vague ("restore the service") fail exactly when they're needed most; steps scripted too rigidly for every possible variation become unmaintainable and get skipped. The right level of detail is specific enough to follow without deep system expertise, general enough to survive minor environment changes.
- Skipping verification checks in favor of "the step ran without erroring" produces false-confidence recoveries that unravel later, often in front of a customer.
- A single sprawling runbook trying to cover every disaster type becomes too long to use under pressure; splitting into scenario-specific runbooks with a shared skeleton keeps each one usable, at the cost of a bit more coordination to keep the shared skeleton consistent across all of them.
- Estimated timelines never revisited after real exercises stay theoretical and stop being useful for spotting a stalled step; they need the same review cadence as the rest of the document.
You're given a service dependency map: several services depend on a shared database, one depends on an external payment API, and a couple depend on each other. Build a recovery order and explain your reasoning. Call out at least one opportunity to recover things in parallel rather than strictly sequentially, and what you'd do if one step in the sequence fails.
Sample Answer
Direct answer. Build a directed dependency graph from the map, topologically order it so nothing starts before what it depends on is verified healthy, and treat truly independent branches as parallel recovery tracks. Validate each step before advancing, and when a step fails, hold everything downstream in a known state and escalate through a defined decision point rather than retrying indefinitely.
1. Build and order the graph
- List every service and every "depends on" edge from the map. Anything with no unresolved incoming edges (here, the shared database) starts first.
- Order the rest by dependency depth: services that depend only on the database, then services that depend on those, and so on. A service that depends on the external payment API doesn't block anything upstream, since nothing needs it in order to start; it just needs its own dependency confirmed before it accepts payment-related traffic.
- Two services that depend on each other (a cycle) can't be strictly ordered against one another. Treat them as one recovery unit: bring both up together, then verify the pair as a whole before opening it to traffic.
2. Find the parallel opportunity
Any two services with no dependency edge between them, direct or transitive, can recover on separate tracks. The common case here is two services that both depend only on the shared database but not on each other: once the database is verified healthy, both can start at the same time instead of one waiting on the other. Parallelizing shortens wall-clock recovery time without changing the correctness of the order, since neither branch depends on the other's state.
3. Checkpoint and verify data integrity at each step, not just liveness
"Started" and "verified healthy" are not the same thing. At each step, confirm the service isn't just accepting connections but is serving data consistent with the last known-good state (for example: row counts within an expected range, no partial-write artifacts, a plausible latest-committed-transaction time). Skipping this and checking only process liveness is how a corrupted or half-restored dependency gets silently built on top of.
Log a checkpoint (what was verified, when, by what evidence) at each step. That makes the sequence auditable afterward and lets a later step's failure be traced back to a specific, verified starting point instead of an assumption.
4. Which pieces can stay in a degraded state while others recover
Not everything needs to be fully healthy before it's usable. A service that only reads reference data from the shared database can often be brought up in a degraded, read-only mode as soon as the database is verified, well before every other service has finished its own recovery. The service that depends on the external payment API is a good candidate to deliberately hold in a degraded state (accept and queue requests, decline new payment attempts) until its own dependency check passes, rather than blocking the rest of the sequence on an external system nobody on the recovery team controls.
What counts as an acceptable degraded state is a business call as much as a technical one: it depends on what each function actually needs to deliver during the outage, not on what's technically possible to bring online.
5. Handling a failed step
Define a bounded number of retries with a real timeout, not an indefinite retry loop; retrying against a dependency that isn't going to recover in that window just burns time. When the bound is hit, hold everything downstream of the failed step in its current state and escalate to a named decision point (the recovery coordinator, or the owner of the affected business function) who chooses between: wait longer, fail over to an alternate path if one exists, or proceed with dependents running in a degraded mode that doesn't require the failed step. Whatever is decided gets logged with the reasoning, same as the checkpoints, so a later review can tell what was planned and what was improvised under pressure.
Worked example
Map: shared-database (no dependencies); service-A and service-B (both depend only on shared-database, not on each other); service-C (depends on service-A); service-D (depends only on the external payment API).
Order: shared-database first. Then service-A and service-B in parallel, since neither depends on the other and both only need the database. service-C waits specifically for service-A's checkpoint to pass, and can start while service-B is still on its own track. service-D can be started any time after the database step but is deliberately held in a degraded (read-only or queued) mode until the external payment API's own health check clears, since that dependency is outside the team's control.
If service-A's checkpoint fails (health check passes but the row-count integrity check does not), service-C does not start. service-C's traffic is held, service-A is retried up to the defined bound, and if it doesn't clear, the recovery coordinator escalates: either fail service-A over to a secondary if one exists, or accept that service-C runs in a reduced-functionality mode that doesn't need service-A's data, whichever gets the organization to an acceptable state fastest.
Trade-offs & pitfalls. The most common wrong turn is treating "process is up" as "safe to build on top of": a liveness check is not a correctness check, and skipping the integrity checkpoint is how a downstream service inherits a silently corrupted starting point. Being too conservative (fully sequential, no parallel tracks) wastes recovery time when the graph has real independence in it; being too aggressive (parallelizing across an actual dependency edge) risks starting a service against data that isn't there yet. The graph should decide what's parallelizable, not intuition. And a dependency map that looks authoritative but isn't kept current is worse than no map at all, because it gives false confidence about what's actually safe to run in parallel.
Once you have BIA findings for a set of business services, how do you translate that into recovery-priority tiers, and what actually determines whether a service lands in the top tier versus the bottom one? Walk through how those tiers then drive budget and staffing decisions.
Sample Answer
Direct answer
Tiering translates BIA (business impact analysis) findings into a small number of recovery-priority bands, using the business's tolerance for downtime and data loss as the primary axis, not technical difficulty. A service lands in the top tier because sustained disruption threatens revenue, safety, legal or regulatory standing, or a large share of customers within a very short window, not because it happens to be technically easiest to protect. The tiers then become the lever finance and staffing use to decide where headroom, retainer contracts, and dedicated on-call coverage get funded.
From BIA findings to tier criteria
- MTPD / MAO (maximum tolerable period of disruption, also called maximum acceptable outage): the longest a service can be down before consequences become unacceptable to the business. This is the anchor input from the BIA.
- Combine MTPD with: financial loss per hour of disruption, regulatory or contractual exposure, safety impact, breadth of customers or users affected, and dependency fan-out (how many other services or business functions break if this one is down).
- Score, don't guess: weight the BIA findings for each service across these dimensions. Example weighting for illustration only (the actual weights are a business decision the BIA sponsor and finance sign off on, not something one team sets unilaterally): financial impact 35%, regulatory/legal exposure 25%, customer scope 25%, dependency fan-out 15%.
Translating scores into tiers
| Tier | Recovery expectation (MTPD-driven) | What lands here | Staffing / budget posture |
|---|---|---|---|
| 1 | Minutes to a few hours | Revenue-critical or safety/regulatory-critical, broad customer exposure | Dedicated on-call rotation, pre-funded standby resources, exercised quarterly |
| 2 | Several hours to one business day | Degraded operation tolerable briefly, moderate exposure | Named on-call owner, exercised semi-annually |
| 3 | One to several business days | Internal or narrow-scope, workaround exists | Best-effort recovery, reviewed annually |
| 4 | No fixed near-term deadline | Low-impact, batch, or back-office | Recovery scheduled opportunistically |
What actually pushes a service to the top isn't "sounds important," it's the combination of speed of consequence onset (does damage start accruing in minutes) and irreversibility (can the loss be made up later, like a delayed nightly batch, or is it gone for good, like a breached SLA credit or a safety incident). A service with high total impact but slow-accruing consequence, a monthly reporting job that only starts costing money after a week of delay, is often Tier 2, not Tier 1, even if its total dollar impact looks large on paper.
How tiers drive budget and staffing
- Rotas: Tier 1 justifies funded, cross-trained on-call rotations with named backups; Tier 3/4 rely on the general on-call pool at best effort.
- Standby spend: Tier 1 gets pre-approved budget for whatever standby resources the business decided it needs; this is where the tiering output hands off into the technical recovery-strategy conversation, but the tier itself is a business-impact call, not an architecture decision.
- Exercise cadence: higher tiers are drilled more often, because the cost of an untested plan scales with the cost of the outage it exists to prevent.
- Supplier spend: Tier 1 is where organizations pay for expedited third-party support contracts and priority escalation paths.
- Governance: tiers get reviewed at a fixed cadence (at least annually, or whenever the BIA is refreshed) and signed off by the business owner, because a service's tier is a statement about acceptable business risk, not just an IT classification. This is the same discipline external frameworks like the ISO 22301 business continuity standard expect: BIA and risk assessment as a documented, periodically reviewed input to recovery prioritization, without the standard itself prescribing specific weights or tier counts; those stay organization-specific.
Worked example
A mid-size B2B SaaS company scores three services 0-10 on each dimension using the weighting above (financial 0.35, regulatory 0.25, customer scope 0.25, dependency fan-out 0.15), with band cutoffs of Tier 1 at 7.5+, Tier 2 at 5.0-7.49, Tier 3 at 2.5-4.99, Tier 4 below 2.5:
- Payment processing: financial 9, regulatory 9, customer scope 8, dependency 7.
9(0.35)+9(0.25)+8(0.25)+7(0.15)=3.15+2.25+2.00+1.05=8.45
Composite 8.45 lands in Tier 1. - Customer support ticketing: financial 5, regulatory 4, customer scope 6, dependency 4.
5(0.35)+4(0.25)+6(0.25)+4(0.15)=1.75+1.00+1.50+0.60=4.85
Composite 4.85 lands in Tier 3. - Internal expense-reporting tool: financial 2, regulatory 1, customer scope 1, dependency 2.
2(0.35)+1(0.25)+1(0.25)+2(0.15)=0.70+0.25+0.25+0.30=1.50
Composite 1.50 lands in Tier 4.
Consequence for staffing and budget: payment processing (Tier 1) gets a funded dedicated on-call rotation and a pre-approved contract with a backup payment partner; the expense tool (Tier 4) is recovered on a best-effort basis by the general helpdesk queue with no dedicated budget line.
Trade-offs and pitfalls
- Letting engineering effort quietly redefine tiers ("it's already highly available so it must be Tier 1") inverts the logic. Tiering has to stay anchored to business impact, not to what's already been built.
- Too many tiers dilutes the prioritization signal; two to four is the usual practical range. Beyond that, tiering stops driving clear resourcing decisions.
- Scoring once and never revisiting is a common failure: business models shift (an internal tool becomes customer-facing) and tiering needs the same refresh cadence as the BIA itself, not a one-time exercise.
- Weighting choices are political. Finance, legal, and the business owner need to agree the weights before scores are trusted, or every re-score becomes a renegotiation.
- A score sitting near a boundary, like the ticketing example above, deserves judgment, not automatic sorting; treat cutoffs as a starting point, not a verdict.
Before a disaster recovery event ever happens, what stakeholder roles and communication structure should already be defined? Walk through who owns which responsibility and how the update cadence should work.
Sample Answer
Direct answer
Before any event happens, the organization needs a named, rehearsed structure for who acts, who decides, and who talks to whom, because working this out for the first time mid-crisis is how communication chaos happens. At minimum: one clearly designated decision-maker per event, a communications owner who is a different person from whoever is doing the operational recovery work, and a pre-agreed update cadence for each audience.
Roles that need to exist before an event
- Continuity commander: single point of decision authority for the event's duration. Who this is, and how backup authority is documented if they're unreachable, is its own governance question, declaration authority (who has the standing to formally declare and activate the plan); what matters here is that the role exists and is named in advance.
- Communications lead: separate from the people fixing the problem, so the team doing operational recovery isn't also drafting customer-facing language under pressure.
- Function or service leads: one named owner per affected business function, reporting status and progress into the commander.
- Executive sponsor: kept informed, and consulted on higher-stakes trade-offs depending on severity (for example, accepting a longer outage versus authorizing emergency mitigation spend).
- Legal or compliance liaison: engaged whenever an event may carry regulatory notification obligations, contractual exposure, or litigation risk, so anything going external gets reviewed first.
- Scribe: keeps a timestamped decision log, separate from the commander's job, so no one is both making decisions and writing the record at the same time.
Communication structure by audience
| Audience | What they need | Typical cadence |
|---|---|---|
| Responders / internal technical teams | Granular, real-time status | Continuous, in the working channel |
| Executives | Business impact, ETA, decisions needed from them | Fixed interval (e.g., every 30-60 minutes) or on material change |
| Customers / external | Plain-language status, no internal detail | At defined checkpoints: acknowledgment, updates, resolution |
| Regulators / legal (when applicable) | Factual, reviewed language, notification-timeline aware | Per the applicable compliance obligation, not ad hoc |
Cadence gets committed to in advance, not decided live: pick a default interval per audience, and let the communications lead depart from the default when conditions genuinely warrant it, rather than negotiating frequency mid-event. What makes that cadence achievable under pressure is pre-built material: templated messages for each audience and status ("investigating," "mitigating," "resolved," "follow-up to come"), a pre-identified channel per audience, and a standing decision-log format the scribe isn't inventing on the fly.
Worked example
A mid-size retailer's checkout service fails during a promotional weekend. Because roles were pre-assigned, the on-call engineering lead isn't the one fielding calls from the VP of Sales; the communications lead handles that, posting a pre-written "investigating" update within minutes and following the pre-agreed cadence rather than waiting for engineering to have spare attention. The legal liaison is looped in immediately because payment processing is affected, even though no breach has occurred, purely because that trigger condition was defined in advance rather than judged in the moment.
Trade-offs and pitfalls
- If the same person both fixes the problem and writes the customer update, both jobs suffer. The split has to be a standing structure, not an improvisation invented under pressure.
- Over-specifying cadence for minor events creates alert fatigue and unnecessary overhead; scale the structure to the declared severity level instead of applying one rhythm to everything.
- A structure that only lives in a document nobody has read is no better than none. It has to be exercised so people know their role before they're asked to play it for real.
- Looping in legal or compliance "when needed," without a pre-defined trigger, tends to mean they're looped in too late. Define the trigger condition in advance, not in the moment.
Design a communications plan for major incidents that affect availability. Cover who needs updates (executives, customers, engineering, legal), how often, what a status-page update should say, and when you escalate from an engineering update to an executive one. Sketch out what the first sixty minutes of communication would look like for a critical outage.
Sample Answer
Direct answer
I'd build the plan around three things: a stakeholder map that separates internal audiences (engineering, executives, legal) from external ones (customers, partners, press) since they need different content and tone, a RACI matrix (who is Responsible, Accountable, Consulted and Informed) so it's unambiguous who is allowed to say what to whom, and a defined escalation threshold for when a routine engineering update becomes an executive one. The first sixty minutes follow a fixed cadence: detect, declare, and get an initial acknowledgment out fast, even before the cause is known, then update on a regular clock rather than only when there's new information.
Structured elaboration
Stakeholder map: internal vs. external.
| Audience | Content | Tone |
|---|---|---|
| Engineering (internal) | Technical symptoms, current hypothesis, what's being tried, what's needed from them | Direct, technical, fast-moving |
| Executives (internal) | Business impact, decision points that need their authority, reputational or legal exposure | Concise, framed around impact and decisions, not implementation detail |
| Legal and compliance (internal) | Anything with regulatory or contractual exposure, review of external language before it ships | Precise, focused on obligations and risk |
| Customers and partners (external) | What's affected, current status, workaround if any, next update time | Plain language, no internal jargon, no speculation about root cause until confirmed |
| Press (external, if triggered) | A single approved statement, routed only through the designated spokesperson | Careful, legally reviewed, minimal |
RACI for the response.
| Role | RACI |
|---|---|
| Incident Commander | Responsible for coordinating the technical response and deciding when thresholds are crossed |
| Communications Lead | Responsible for drafting and publishing every external update; the single pen holding the message |
| Executive sponsor | Accountable for approving customer-facing language once the escalation threshold is crossed |
| Legal | Consulted before any public statement that touches liability, data exposure, or contractual terms |
| Customer support | Informed, and executes pre-approved response templates for inbound customer questions |
Escalation thresholds, engineering to executive. A routine update stays at the engineering level until it crosses a defined line, for example greater than 25% of customers affected, any customer-visible data integrity risk, a duration past 30 minutes with no root cause identified, or any event that spans multiple regions or products at once. Crossing that line triggers the executive sponsor and legal into the loop immediately, not at the next scheduled update.
Cross-timezone coordination for multi-region events. When an incident spans regions, assign a single Communications Lead of record per shift with a documented, timestamped handoff (what's been said publicly, what's pending, what's the next scheduled update time) so a shift change never produces two conflicting messages. All approved language lives in one shared document that both shifts write to, not in separate regional threads.
Status page content. Header with incident ID and affected services, a plain-language impact summary, current status (investigating, identified, mitigating, resolved), a short list of actions taken, and a concrete next-update time rather than a vague "soon."
Escalation flow.
flowchart TD
A[Incident detected] --> B{Meets major-incident threshold?}
B -- No --> C[Engineering tracks it, no exec notify]
B -- Yes --> D[Incident Commander declares major incident]
D --> E[Internal update: scope, IC, ETA]
D --> F[Status page: Investigating]
E --> G{Unresolved at 30 min or revenue-affecting?}
G -- No --> H[Engineering-led updates every 30 min]
G -- Yes --> I[Escalate to executives and legal]
I --> J[Exec liaison approves customer-facing language]
J --> K[Status page and customer comms updated]
H --> L[Resolution and postmortem scheduled]
K --> L
Worked example
First sixty minutes of a critical, multi-region outage at a mid-size payments company:
- Minute 0: Incident Commander declares the major incident based on the on-call alert.
- Minute 5: Communications Lead publishes an initial status page entry: "We're investigating reports of failed transactions. Next update in 15 minutes." No cause is stated yet.
- Minute 15: Internal update to the executive sponsor and legal: current hypothesis, estimated affected percentage, and whether the threshold has been crossed. In this case it has, since more than 25% of customers in two regions are affected.
- Minute 20: Status page updated with the confirmed impact scope, still no cause speculation externally.
- Minute 40: A mitigation is applied; the Communications Lead updates internally and externally with early signs of recovery, careful to say "error rates are declining" rather than "resolved," since a partial rollback can regress.
- Minute 60: If recovery is confirmed, a resolution update is published with a commitment to a postmortem within 72 hours; if not, the next update time is restated and the executive sponsor is briefed on the extended timeline.
Trade-offs and pitfalls
Too frequent updates without new information create alarm fatigue and can look like the team doesn't have control; too infrequent updates erode customer trust faster than the outage itself does. Publishing a root cause before it's confirmed is a common and expensive mistake, since retracting it externally is more damaging than staying vague a little longer. A single point of failure in the plan itself is common: if the Communications Lead is unreachable, there must be a named backup, otherwise the whole cadence stalls. Finally, this is a crisis-communications discipline, not an incident-tooling configuration exercise: the plan should specify who says what to whom and when, not which paging tool routes the alert.
Unlock Full Question Bank
Get access to all 20 Disaster Recovery and Business Continuity interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.