Disaster Recovery and Business Continuity Questions
Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.
Design a communications plan for major incidents that affect availability. Cover who needs updates (executives, customers, engineering, legal), how often, what a status-page update should say, and when you escalate from an engineering update to an executive one. Sketch out what the first sixty minutes of communication would look like for a critical outage.
Sample Answer
Direct answer
I'd build the plan around three things: a stakeholder map that separates internal audiences (engineering, executives, legal) from external ones (customers, partners, press) since they need different content and tone, a RACI matrix (who is Responsible, Accountable, Consulted and Informed) so it's unambiguous who is allowed to say what to whom, and a defined escalation threshold for when a routine engineering update becomes an executive one. The first sixty minutes follow a fixed cadence: detect, declare, and get an initial acknowledgment out fast, even before the cause is known, then update on a regular clock rather than only when there's new information.
Structured elaboration
Stakeholder map: internal vs. external.
| Audience | Content | Tone |
|---|---|---|
| Engineering (internal) | Technical symptoms, current hypothesis, what's being tried, what's needed from them | Direct, technical, fast-moving |
| Executives (internal) | Business impact, decision points that need their authority, reputational or legal exposure | Concise, framed around impact and decisions, not implementation detail |
| Legal and compliance (internal) | Anything with regulatory or contractual exposure, review of external language before it ships | Precise, focused on obligations and risk |
| Customers and partners (external) | What's affected, current status, workaround if any, next update time | Plain language, no internal jargon, no speculation about root cause until confirmed |
| Press (external, if triggered) | A single approved statement, routed only through the designated spokesperson | Careful, legally reviewed, minimal |
RACI for the response.
| Role | RACI |
|---|---|
| Incident Commander | Responsible for coordinating the technical response and deciding when thresholds are crossed |
| Communications Lead | Responsible for drafting and publishing every external update; the single pen holding the message |
| Executive sponsor | Accountable for approving customer-facing language once the escalation threshold is crossed |
| Legal | Consulted before any public statement that touches liability, data exposure, or contractual terms |
| Customer support | Informed, and executes pre-approved response templates for inbound customer questions |
Escalation thresholds, engineering to executive. A routine update stays at the engineering level until it crosses a defined line, for example greater than 25% of customers affected, any customer-visible data integrity risk, a duration past 30 minutes with no root cause identified, or any event that spans multiple regions or products at once. Crossing that line triggers the executive sponsor and legal into the loop immediately, not at the next scheduled update.
Cross-timezone coordination for multi-region events. When an incident spans regions, assign a single Communications Lead of record per shift with a documented, timestamped handoff (what's been said publicly, what's pending, what's the next scheduled update time) so a shift change never produces two conflicting messages. All approved language lives in one shared document that both shifts write to, not in separate regional threads.
Status page content. Header with incident ID and affected services, a plain-language impact summary, current status (investigating, identified, mitigating, resolved), a short list of actions taken, and a concrete next-update time rather than a vague "soon."
Escalation flow.
flowchart TD
A[Incident detected] --> B{Meets major-incident threshold?}
B -- No --> C[Engineering tracks it, no exec notify]
B -- Yes --> D[Incident Commander declares major incident]
D --> E[Internal update: scope, IC, ETA]
D --> F[Status page: Investigating]
E --> G{Unresolved at 30 min or revenue-affecting?}
G -- No --> H[Engineering-led updates every 30 min]
G -- Yes --> I[Escalate to executives and legal]
I --> J[Exec liaison approves customer-facing language]
J --> K[Status page and customer comms updated]
H --> L[Resolution and postmortem scheduled]
K --> L
Worked example
First sixty minutes of a critical, multi-region outage at a mid-size payments company:
- Minute 0: Incident Commander declares the major incident based on the on-call alert.
- Minute 5: Communications Lead publishes an initial status page entry: "We're investigating reports of failed transactions. Next update in 15 minutes." No cause is stated yet.
- Minute 15: Internal update to the executive sponsor and legal: current hypothesis, estimated affected percentage, and whether the threshold has been crossed. In this case it has, since more than 25% of customers in two regions are affected.
- Minute 20: Status page updated with the confirmed impact scope, still no cause speculation externally.
- Minute 40: A mitigation is applied; the Communications Lead updates internally and externally with early signs of recovery, careful to say "error rates are declining" rather than "resolved," since a partial rollback can regress.
- Minute 60: If recovery is confirmed, a resolution update is published with a commitment to a postmortem within 72 hours; if not, the next update time is restated and the executive sponsor is briefed on the extended timeline.
Trade-offs and pitfalls
Too frequent updates without new information create alarm fatigue and can look like the team doesn't have control; too infrequent updates erode customer trust faster than the outage itself does. Publishing a root cause before it's confirmed is a common and expensive mistake, since retracting it externally is more damaging than staying vague a little longer. A single point of failure in the plan itself is common: if the Communications Lead is unreachable, there must be a named backup, otherwise the whole cadence stalls. Finally, this is a crisis-communications discipline, not an incident-tooling configuration exercise: the plan should specify who says what to whom and when, not which paging tool routes the alert.
You're given a service dependency map: several services depend on a shared database, one depends on an external payment API, and a couple depend on each other. Build a recovery order and explain your reasoning. Call out at least one opportunity to recover things in parallel rather than strictly sequentially, and what you'd do if one step in the sequence fails.
Sample Answer
Direct answer. Build a directed dependency graph from the map, topologically order it so nothing starts before what it depends on is verified healthy, and treat truly independent branches as parallel recovery tracks. Validate each step before advancing, and when a step fails, hold everything downstream in a known state and escalate through a defined decision point rather than retrying indefinitely.
1. Build and order the graph
- List every service and every "depends on" edge from the map. Anything with no unresolved incoming edges (here, the shared database) starts first.
- Order the rest by dependency depth: services that depend only on the database, then services that depend on those, and so on. A service that depends on the external payment API doesn't block anything upstream, since nothing needs it in order to start; it just needs its own dependency confirmed before it accepts payment-related traffic.
- Two services that depend on each other (a cycle) can't be strictly ordered against one another. Treat them as one recovery unit: bring both up together, then verify the pair as a whole before opening it to traffic.
2. Find the parallel opportunity
Any two services with no dependency edge between them, direct or transitive, can recover on separate tracks. The common case here is two services that both depend only on the shared database but not on each other: once the database is verified healthy, both can start at the same time instead of one waiting on the other. Parallelizing shortens wall-clock recovery time without changing the correctness of the order, since neither branch depends on the other's state.
3. Checkpoint and verify data integrity at each step, not just liveness
"Started" and "verified healthy" are not the same thing. At each step, confirm the service isn't just accepting connections but is serving data consistent with the last known-good state (for example: row counts within an expected range, no partial-write artifacts, a plausible latest-committed-transaction time). Skipping this and checking only process liveness is how a corrupted or half-restored dependency gets silently built on top of.
Log a checkpoint (what was verified, when, by what evidence) at each step. That makes the sequence auditable afterward and lets a later step's failure be traced back to a specific, verified starting point instead of an assumption.
4. Which pieces can stay in a degraded state while others recover
Not everything needs to be fully healthy before it's usable. A service that only reads reference data from the shared database can often be brought up in a degraded, read-only mode as soon as the database is verified, well before every other service has finished its own recovery. The service that depends on the external payment API is a good candidate to deliberately hold in a degraded state (accept and queue requests, decline new payment attempts) until its own dependency check passes, rather than blocking the rest of the sequence on an external system nobody on the recovery team controls.
What counts as an acceptable degraded state is a business call as much as a technical one: it depends on what each function actually needs to deliver during the outage, not on what's technically possible to bring online.
5. Handling a failed step
Define a bounded number of retries with a real timeout, not an indefinite retry loop; retrying against a dependency that isn't going to recover in that window just burns time. When the bound is hit, hold everything downstream of the failed step in its current state and escalate to a named decision point (the recovery coordinator, or the owner of the affected business function) who chooses between: wait longer, fail over to an alternate path if one exists, or proceed with dependents running in a degraded mode that doesn't require the failed step. Whatever is decided gets logged with the reasoning, same as the checkpoints, so a later review can tell what was planned and what was improvised under pressure.
Worked example
Map: shared-database (no dependencies); service-A and service-B (both depend only on shared-database, not on each other); service-C (depends on service-A); service-D (depends only on the external payment API).
Order: shared-database first. Then service-A and service-B in parallel, since neither depends on the other and both only need the database. service-C waits specifically for service-A's checkpoint to pass, and can start while service-B is still on its own track. service-D can be started any time after the database step but is deliberately held in a degraded (read-only or queued) mode until the external payment API's own health check clears, since that dependency is outside the team's control.
If service-A's checkpoint fails (health check passes but the row-count integrity check does not), service-C does not start. service-C's traffic is held, service-A is retried up to the defined bound, and if it doesn't clear, the recovery coordinator escalates: either fail service-A over to a secondary if one exists, or accept that service-C runs in a reduced-functionality mode that doesn't need service-A's data, whichever gets the organization to an acceptable state fastest.
Trade-offs & pitfalls. The most common wrong turn is treating "process is up" as "safe to build on top of": a liveness check is not a correctness check, and skipping the integrity checkpoint is how a downstream service inherits a silently corrupted starting point. Being too conservative (fully sequential, no parallel tracks) wastes recovery time when the graph has real independence in it; being too aggressive (parallelizing across an actual dependency edge) risks starting a service against data that isn't there yet. The graph should decide what's parallelizable, not intuition. And a dependency map that looks authoritative but isn't kept current is worse than no map at all, because it gives false confidence about what's actually safe to run in parallel.
You're partway through a sequenced recovery when a third-party dependency you were counting on stays down longer than expected. Which services do you bring online anyway, how do you handle the transactions that would normally rely on that dependency, and how do you communicate the degraded state to customers in the meantime?
Sample Answer
Direct answer. Bring online everything that doesn't strictly need the down dependency, and put a firm, visible hold on anything that does, rather than letting it silently degrade or silently retry. Decide what "handling a transaction" means without the dependency in a way that never risks a customer being charged, billed, or committed twice: hold, don't guess. And tell customers proactively, in plain language, before they have to ask.
1. Decide what comes online, using criticality, not convenience
This decision should trace back to the business impact analysis, the process that ranks business functions by how much an hour or a day of downtime actually costs, so it isn't made ad hoc mid-incident. Functions that don't touch the down dependency come online first. Functions that touch it only on a non-essential path (browsing, viewing account history) come online in a read-only or informational mode. Functions where the dependency is essential to the transaction itself (authorizing a new charge) stay explicitly gated, not silently attempted.
The authority to declare "we're operating in a degraded state" and to approve which functions run that way should be defined in the continuity plan ahead of time, not improvised in the room. That's usually someone senior enough to own the customer and regulatory risk of the decision, not just whoever is closest to the outage.
2. Handle the affected transactions without guessing
Accept the transaction, record the customer's intent, and hold it in a clearly marked pending state rather than attempting it against a dependency that isn't there, or silently retrying it in the background where the customer can't see what happened. Be honest about the state: "received and pending" is very different from letting a customer believe it went through, and that distinction is also what stops a customer from trying again themselves out of uncertainty.
When the dependency comes back, process the backlog in order and reconcile before declaring the incident closed: confirm, transaction by transaction, that what your system believes happened matches what the dependency's own records show, and follow up individually on anything that doesn't match rather than assuming the queue drained cleanly.
3. Communicate the degraded state to customers
Say something before they have to ask. A visible status message in the product itself, not only on a status page nobody checks mid-transaction, that names what's affected, sets expectations (their action is saved but pending, not lost), and gives a realistic timeframe, even a wide one, beats silence. Update it as the situation changes, and if any pending items need a customer to take action, reach out directly instead of leaving them to notice on their own. Keep the message consistent across channels so a customer who checks two of them doesn't get two different stories on top of the outage itself.
4. Protect service levels for what's still running
The functions still online need their own expectations reset for the duration. It's reasonable to temporarily relax targets for anything adjacent to the affected dependency (background jobs that would also normally touch it) so they don't cause a second incident by retrying aggressively against something that's down. That should be a deliberate, communicated decision, not something that just happens because nobody planned for it.
Worked example
A checkout flow depends on a third-party payment processor down for two hours, well past what anyone expected. Browsing, cart building, and order history don't touch the processor, so they stay fully online. Checkout is gated: a customer can complete every step through "place order," and the order is recorded as pending payment with the cart and the chosen payment method captured, but no charge is attempted. The product tells them directly: order saved, payment processing is delayed, confirmation will follow once it's back, expected within the hour. No background retry loop fires against the processor. When it recovers, pending orders are processed in the order placed, and each is reconciled against the processor's own transaction record before being marked complete, so a customer is never charged twice even if an earlier attempt is later discovered to have partially gone through.
Trade-offs & pitfalls. The tempting shortcut is to keep retrying the transaction in the background hoping the dependency returns soon; that's exactly how a customer ends up double-charged if the retry succeeds silently after they've already tried again themselves out of frustration. Hold and communicate, don't guess and retry. Silence is worse than an honest "we don't know exactly when," because customers who get no information assume the worst or attempt workarounds that make reconciliation harder afterward. And bringing too much online too fast, without a clear degraded-mode decision from someone with the authority to own that risk, is how a team ends up shipping a feature that looks live but isn't actually safe to depend on.
Walk me through how you'd actually conduct a Business Impact Analysis for a company's application portfolio. Who would you interview, what would you ask them to assess criticality, and what would the finished output look like?
Sample Answer
Direct answer
I'd run the business impact analysis (BIA), the exercise that quantifies what an outage costs each business function and sets its recovery targets, as a structured set of stakeholder interviews mapped against an inventory of the application portfolio, not a survey blast. For each application I'd talk to the business process owner (what breaks and how fast the harm grows), a downstream consumer of that application (what breaks for them), and a technical lead (what this application actually depends on, so the impact can be traced through the chain). The finished output is a BIA register: one row per application with its criticality tier, tolerance window, business-side recovery time objective (RTO, the target time to get a function working again) and recovery point objective (RPO, how much recent data the function can afford to lose, measured in time), and its upstream and downstream dependencies, signed off by the business owner.
Structured elaboration
1. Scope and inventory. Build the list of in-scope applications from the business side (which processes exist) rather than the IT asset list, so nothing gets excluded just because it lacks a system owner in the configuration management database (CMDB), the IT system of record for tracked assets.
2. Stakeholder interviews. Two distinct conversations per application, deliberately kept separate so the technical lead can't quietly set the business tolerance:
- With the business/process owner: "If this is unavailable for 1 hour, 4 hours, 24 hours, 3 days, what actually happens? Who notices, what's the financial or legal consequence, is there a seasonal or peak period where this gets worse, and is there a manual workaround?"
- With a downstream consuming team: "What do you depend on this for, and does your own tolerance window get eaten by the time this stays down?" This surfaces cascading impact that the primary owner often doesn't see.
- With the technical/ops lead: "What does this depend on upstream (data, other systems, third parties, specific teams) and who would need to be involved to bring it back?" This is a dependency map, not a design conversation; the goal is knowing what's connected, not deciding how to recover it.
3. Scoring and tiering. Combine a quantitative score (revenue or cost impact per elapsed-time band, contractual penalties) with a qualitative one (regulatory exposure, reputational harm, customer trust) using a shared rubric applied the same way across every interview, specifically to counter the fact that every owner tends to describe their own system as critical.
4. Output. A BIA register document: application name, business owner, criticality tier, maximum tolerable period of disruption (MTPD), target RTO/RPO, upstream and downstream dependencies, and the date of last review. This becomes the input to recovery sequencing and to the exercise program's scope.
Worked example
Three systems from a recent BIA at a mid-size company, showing why the interview-based approach catches things a pure system-criticality guess would miss:
| System | Owner's first guess | What the interviews revealed | Final tier |
|---|---|---|---|
| Customer-facing web application | "Obviously Tier 1" | Confirmed: drives signup and checkout revenue directly; no manual workaround exists | Tier 1, RTO 1 hour |
| Internal HR system | "Probably low priority" | Downstream interview with Payroll revealed payroll processing depends on data pulled from this system by day 3 of the month, and missing that window creates a wage-payment compliance issue in several jurisdictions | Tier 2, RTO 48 hours, with a hard deadline constraint tied to the pay cycle rather than a flat window |
| Analytics pipeline | "Low priority, it's just dashboards" | Confirmed genuinely low priority day-to-day, but the downstream interview with Finance revealed it feeds the monthly board reporting package, so its real constraint is a monthly deadline, not a continuous-availability one | Tier 3, RTO tied to reporting deadline (typically several days), not hours |
The HR and analytics examples are the reason the downstream-consumer interview is a separate step: the primary owner of both systems undersold their own criticality, and only the consuming team surfaced the real constraint.
Trade-offs and pitfalls
Interview-based BIAs are vulnerable to self-serving bias (every owner wants to be Tier 1 to guarantee resources), which is why a shared, applied-consistently rubric matters more than any individual interview. They're also vulnerable to missing cascading dependencies if you only interview the primary owner and skip the downstream-consumer conversation, as the HR and analytics examples show. A separate trap specific to a technical audience: it's tempting to let the technical-lead interview drift into "how would we actually achieve that RTO," which turns a business-impact exercise into a premature architecture discussion. Keep that conversation to what the system depends on, and leave the how-to-recover-it decision for the technical team afterward, informed by the tier the BIA assigned.
Before a disaster recovery event ever happens, what stakeholder roles and communication structure should already be defined? Walk through who owns which responsibility and how the update cadence should work.
Sample Answer
Direct answer
Before any event happens, the organization needs a named, rehearsed structure for who acts, who decides, and who talks to whom, because working this out for the first time mid-crisis is how communication chaos happens. At minimum: one clearly designated decision-maker per event, a communications owner who is a different person from whoever is doing the operational recovery work, and a pre-agreed update cadence for each audience.
Roles that need to exist before an event
- Continuity commander: single point of decision authority for the event's duration. Who this is, and how backup authority is documented if they're unreachable, is its own governance question, declaration authority (who has the standing to formally declare and activate the plan); what matters here is that the role exists and is named in advance.
- Communications lead: separate from the people fixing the problem, so the team doing operational recovery isn't also drafting customer-facing language under pressure.
- Function or service leads: one named owner per affected business function, reporting status and progress into the commander.
- Executive sponsor: kept informed, and consulted on higher-stakes trade-offs depending on severity (for example, accepting a longer outage versus authorizing emergency mitigation spend).
- Legal or compliance liaison: engaged whenever an event may carry regulatory notification obligations, contractual exposure, or litigation risk, so anything going external gets reviewed first.
- Scribe: keeps a timestamped decision log, separate from the commander's job, so no one is both making decisions and writing the record at the same time.
Communication structure by audience
| Audience | What they need | Typical cadence |
|---|---|---|
| Responders / internal technical teams | Granular, real-time status | Continuous, in the working channel |
| Executives | Business impact, ETA, decisions needed from them | Fixed interval (e.g., every 30-60 minutes) or on material change |
| Customers / external | Plain-language status, no internal detail | At defined checkpoints: acknowledgment, updates, resolution |
| Regulators / legal (when applicable) | Factual, reviewed language, notification-timeline aware | Per the applicable compliance obligation, not ad hoc |
Cadence gets committed to in advance, not decided live: pick a default interval per audience, and let the communications lead depart from the default when conditions genuinely warrant it, rather than negotiating frequency mid-event. What makes that cadence achievable under pressure is pre-built material: templated messages for each audience and status ("investigating," "mitigating," "resolved," "follow-up to come"), a pre-identified channel per audience, and a standing decision-log format the scribe isn't inventing on the fly.
Worked example
A mid-size retailer's checkout service fails during a promotional weekend. Because roles were pre-assigned, the on-call engineering lead isn't the one fielding calls from the VP of Sales; the communications lead handles that, posting a pre-written "investigating" update within minutes and following the pre-agreed cadence rather than waiting for engineering to have spare attention. The legal liaison is looped in immediately because payment processing is affected, even though no breach has occurred, purely because that trigger condition was defined in advance rather than judged in the moment.
Trade-offs and pitfalls
- If the same person both fixes the problem and writes the customer update, both jobs suffer. The split has to be a standing structure, not an improvisation invented under pressure.
- Over-specifying cadence for minor events creates alert fatigue and unnecessary overhead; scale the structure to the declared severity level instead of applying one rhythm to everything.
- A structure that only lives in a document nobody has read is no better than none. It has to be exercised so people know their role before they're asked to play it for real.
- Looping in legal or compliance "when needed," without a pre-defined trigger, tends to mean they're looped in too late. Define the trigger condition in advance, not in the moment.
Unlock Full Question Bank
Get access to all 20 Disaster Recovery and Business Continuity interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.