Disaster Recovery and Business Continuity Questions
Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.
Runbooks and continuity documentation go stale fast once systems, teams, and org structure keep changing. How do you keep them accurate over time? Cover ownership, versioning, and how you'd catch drift before it matters during a real event rather than after.
Sample Answer
Direct answer. Ownership has to be a named role tied to the business function or system the runbook covers, not "whoever wrote it last," with versioning that ties each revision to a specific trigger, not just a date. Catching drift before a real event means testing the runbook during a planned exercise instead of discovering it's wrong while executing it under pressure, paired with a lightweight review trigger whenever something the runbook depends on changes.
1. Ownership
Assign a named owner to every runbook, ideally the person or role accountable for the business function or system it covers, not a shared team inbox and not necessarily the original author, who may move on. Ownership means being responsible for the runbook's accuracy, not personally executing every step during a real event. Make it visible on the document itself (name, role, last-reviewed date) so anyone opening it during an incident can see who to ask if something looks wrong, and so a stale "owner" who's left the organization is an obvious red flag rather than a hidden one. When ownership changes hands, require an explicit handoff with a fresh review as part of the transition, not a silent reassignment.
2. Versioning
Treat runbooks as controlled documents: every change gets a version number, a date, who made it, and why, kept in the document's own change history rather than relying on file metadata or institutional memory. Tie versions to triggers, not just a calendar. A scheduled review is a reasonable baseline, but the more reliable trigger is linking review to the events that actually make a runbook wrong: the system it depends on changes, the org structure around it changes, or a real exercise or incident surfaces a gap. A runbook untouched for eighteen months right after the system it describes was rebuilt is far more suspect than one reviewed on schedule six months ago. Keep prior versions accessible, not just the latest, so a post-event review can tell whether a gap that showed up was already known and fixed but hadn't propagated to whoever was executing, or was genuinely new.
3. Catching drift before it matters
The strongest mechanism is using the runbook for real, on a schedule, inside an exercise (a tabletop discussion or a functional exercise that actually executes some steps), before a live event forces the discovery. A runbook nobody has walked through since it was written is a hypothesis, not a tested procedure. Build a lightweight checklist step into whatever process generates the kind of change that breaks runbooks: when a system change, a supplier change, or an org change is being planned, that process flags which runbooks might be affected, so the review happens close to the change instead of being discovered much later. Track a simple staleness signal across the whole runbook set (time since last review, time since the system or team it covers last changed) and report on it the way you'd report any other risk metric, so an accumulating backlog of unreviewed runbooks is visible to whoever owns the continuity program overall, not discovered runbook by runbook.
Worked example
A payments-processing runbook is owned by the payments platform lead, currently at version 4: v2 added a fallback provider, v3 corrected an escalation contact list after a reorg, v4 followed a functional exercise that found the documented rollback step no longer matched how the system actually rolled back. Six months after v4, the team migrates the payments platform to new infrastructure as part of an unrelated project. Because "infrastructure migration for a system with an owned runbook" sits on that project's change-review checklist, the payments platform lead is flagged automatically and confirms the runbook needs an update before the migration ships, catching the drift as part of the planned change rather than during the next incident. At the next scheduled functional exercise three months later, the team walks through the updated runbook end to end and finds one further step, a monitoring dashboard link, still pointing at the decommissioned infrastructure. That becomes a logged finding and v5.
Trade-offs & pitfalls. A purely calendar-based review cadence catches drift too slowly for a fast-changing system and too often for a stable one; tying review to real change triggers is more work to set up but matches effort to actual risk. Ownership without accountability for staleness (a name on a document nobody ever checks) is barely better than no ownership; the staleness signal has to surface to someone, not just exist. And testing a runbook only during a real event is the worst possible time to discover it's wrong: the exercise programme is what lets a team find the gap on an ordinary afternoon instead of during an actual outage.
A regulatory examiner is about to review your business continuity program. What would you want to have ready to show them, and how do you make sure your documentation actually reflects a tested, current program rather than a plan that's been sitting untouched since it was written?
Sample Answer
Direct answer. I'd want a small, coherent evidence package proving three things: the program identifies what actually matters to the business (a current business impact analysis and criticality ranking), the plan has been tested against reality on a regular cadence with findings tracked to closure, and it's actively maintained rather than signed once and shelved. What makes documentation credible to an examiner is version history, dated exercise records with named participants and findings, and open corrective actions with owners and target dates, not how polished the plan document itself looks.
1. Program foundation documents
A current business impact analysis, or BIA (the process that identifies which business functions are critical and quantifies the cost of their downtime over time), is what the rest of the program's priorities trace back to. "Current" means dated and tied to a defined refresh cadence, not a one-time artifact from years ago. Alongside it: the continuity plan itself at the business-function level, describing what each critical function needs to keep operating or recover and who's responsible; and a declaration and activation framework naming who has the authority to formally invoke the plan and the escalation path if that person is unreachable.
2. Evidence the plan has been tested, not just written
Records from an exercise programme across different depths: a tabletop exercise (a facilitated discussion where the team talks through a scenario without actually executing anything), a functional exercise (a partial, real execution of specific procedures), and periodically a full-scale exercise, each with a date, scope, participants, and a documented outcome. A hotwash or after-action report for each one (the structured debrief run immediately afterward, capturing what worked, what didn't, and why), not just a pass or fail note. A corrective action log tying each finding back to a specific exercise, with an owner and a target date, and a record of which items actually closed versus which are still open. An examiner is generally less interested in whether every exercise went flawlessly and more interested in whether findings get followed through on.
3. Evidence the documentation reflects the business today, not the business as it was
Version control on the plan: a change log showing when it was last updated, by whom, and why, ideally tied to specific triggers (an org restructure, a new critical system, a new key supplier) rather than a periodic rubber stamp. A management review record showing leadership periodically reviewing the program's health, including whether resourcing and risk-acceptance decisions were actually made. And training and awareness records for the people named in the plan, since a plan naming a role that no longer exists, or a person who's left, is itself evidence the program isn't current.
4. Regulatory and compliance framing
These are examples of different regulatory regimes, not a checklist to memorize in full: which one applies to you, if any, depends on your sector and jurisdiction. Different regulators expect different levels of prescriptiveness. In the US, financial-sector examiners work from the Federal Financial Institutions Examination Council (FFIEC)'s IT Examination Handbook, which describes practices examiners use to evaluate business continuity management (BCM) governance, resilience strategy, training, exercises, and maintenance, scaled to an institution's size and complexity, rather than imposing a fixed checklist of requirements. The EU's Digital Operational Resilience Act (DORA) is more prescriptive for financial entities: it requires maintaining and periodically testing information and communication technology (ICT) continuity plans, with standard testing performed at least annually and more intensive testing on a longer cycle for larger institutions, and requires records of that testing to be kept. ISO 22301, the international management-system standard for business continuity, lays out the shape a mature program has structurally regardless of sector: a policy, a scope, a BIA and risk assessment, documented strategies and plans, an exercising programme, and the ordinary management-system disciplines of internal audit, management review, and corrective action closed out rather than just logged. Building the evidence package to that shape tends to answer the question before an examiner from any of these regimes asks it.
Worked example
A mid-size organization preparing for an examiner visit assembles: a BIA refreshed nine months earlier and due for its annual refresh in three months, listing the top twelve critical functions and their maximum tolerable downtime; the continuity plan for each function, with a change log showing two revisions in the past year, one triggered by a platform migration and one by a reorganization of the team owning a critical function; records of two tabletop exercises and one functional exercise in the past twelve months, each with a hotwash summary; a corrective action log showing eleven findings raised across those exercises, nine closed with evidence and two still open with named owners and dates inside the next quarter; and quarterly risk-committee minutes showing the program's status was reviewed and the two open items were acknowledged with a committed resourcing decision. That combination, current inputs, a tested cadence, and traceable follow-through, is what makes a program read as alive rather than as a binder written once and never reopened.
Trade-offs & pitfalls. A beautifully written plan with no exercise history behind it is a bigger red flag to an experienced examiner than a rougher plan with a real testing record, since untested plans reliably fail in ways nobody predicted. Closing corrective actions by simply marking them closed without evidence is worse than leaving them visibly open, because an examiner who spot-checks one and finds it wasn't actually fixed calls the whole log's credibility into question. And treating the BIA as a one-time project rather than a maintained artifact is the most common gap of all: businesses change what's critical to them faster than a plan written once accounts for, and an examiner asking how you know this is still accurate is really asking whether a refresh cadence exists at all.
What does it take to make sure a business continuity program actually satisfies the regulatory obligations that apply to your industry, things like financial-services BCP mandates, healthcare contingency-planning rules, or SOC 2 continuity controls? Explain how those obligations shape what the program has to cover and document.
Sample Answer
Direct answer
Regulatory and compliance regimes almost never dictate a specific recovery architecture; they dictate what your continuity program has to contain, document, and be able to prove on demand. In practice that means: a written plan with defined scope, a documented business impact and criticality analysis, defined and justified recovery objectives, evidence that the plan has actually been tested (not just written), and, for most regimes that matter to a mid-size company, explicit coverage of the third parties your critical functions depend on. Different regimes add different specifics on top of that shared skeleton, and the program has to satisfy all of the ones that apply at once, not run a separate plan per regulation.
Structured elaboration
What auditors and examiners actually check. Across the regimes I've worked with, the common ask is: a documented plan naming its scope, a documented impact/criticality analysis, recovery objectives with a rationale for how they were set, dated records of exercises and their findings, evidence that findings got remediated, and a defined review and maintenance cadence. The plan is the artifact; a well-run internal system with no paper trail behind it often fails an audit anyway, because there's nothing to sample.
How named regimes each shape the program:
| Regime | What it specifically requires | What it does not require |
|---|---|---|
| HIPAA Security Rule, 45 CFR 164.308(a)(7) | A written contingency plan with a data backup plan, a disaster recovery plan, and an emergency-mode operations plan (all required); testing/revision procedures and an applications-and-data criticality analysis (both "addressable", meaning implement them or document a reasonable equivalent and the justification for not doing it as written) | A specific backup technology, replication topology, or failover pattern |
| ISO 22301 | A documented business continuity management system including an exercising and testing programme whose results feed back into plan revisions; internationally recognized structure many programs adopt to satisfy a customer's or regulator's "recognized standard" requirement | Certification is voluntary; it's a management-system standard, not a law |
| NIST SP 800-34 | A US federal contingency-planning framework that ties exercise depth to system impact level (a tabletop for lower-impact systems, scaling to functional and full-scale exercises for higher-impact ones); widely referenced as good practice even by non-federal programs | Not itself a legal mandate outside federal systems, but often cited as the bar an auditor expects to see approximated |
| DORA (Regulation (EU) 2022/2554, EU financial entities; entered into force 16 January 2023, applicable and enforceable from 17 January 2025 after a two-year transition) | An information and communication technology (ICT) risk management framework, incident reporting, resilience testing, and explicit oversight of critical ICT third-party providers including contractual audit and exit rights | Does not replace ISO 22301 or NIST-style testing frameworks; it layers third-party oversight and reporting obligations on top |
| The Federal Financial Institutions Examination Council (FFIEC) business continuity management (BCM) guidance (US banking examiners) | Examiner expectations that a bank's BCM program include its critical third-party service providers in the enterprise-wide exercise and testing program, scaled to how critical that provider is | Does not itself set a specific recovery time objective (RTO, the target time to get a function working again) or recovery point objective (RPO, how much recent data the function can afford to lose, measured in time); those targets are left to the institution's own business impact analysis (BIA), the exercise that quantifies downtime cost and assigns each function its own targets |
| SOC 2 (Availability criterion) | Not a regulation but a widely required customer attestation: documented continuity controls plus evidence an independent auditor can sample | Does not prescribe HIPAA- or DORA-style specifics; it checks that controls exist and operate as described |
What actually changes in the program because of these obligations. Scope has to be explicit (what's covered, what's deliberately out, and why), evidence has to be retained (dated exercise records, findings, remediation sign-off), third parties have to be inventoried and included proportional to criticality, and there has to be a named owner and review cadence for the plan itself, not just for the systems it covers.
Worked example
A mid-size regional hospital network under HIPAA needs a written contingency plan covering every system that touches electronic protected health information. Concretely: a data backup plan (required), a disaster recovery plan to restore lost data (required), and an emergency-mode operations plan describing how critical clinical functions, like verifying a patient's active medication list before a new prescription is administered, keep running safely if the electronic health record system is down (required). Testing and a criticality analysis are "addressable": the network can satisfy this by running an annual tabletop exercise and documenting the results, or by documenting a specific reason that approach isn't reasonable for a given system and describing the equivalent measure it uses instead. If the Office for Civil Rights investigates after an incident, what gets requested first is the plan itself, the criticality analysis, and the exercise records, not a description of the network's storage architecture.
A second, shorter example: a mid-size payments company subject to bank-examiner-style BCM oversight maintains a register of its critical ICT third-party providers (its card processor, its cloud host) that states each provider's own tested recovery capability, the contractual audit and notification rights the company holds, and how that provider is included in the company's own exercise program. That register, not a technical diagram, is what an examiner reviews.
Trade-offs and pitfalls
The most common and costly failure is treating compliance as a paperwork exercise: a plan that exists, was reviewed once at signing, and has never been tested satisfies no one, regulator or business, when a real disruption hits. A second failure is misreading HIPAA's "addressable" as "optional": it still requires either the specified control or a documented, reasoned equivalent, and "we didn't get to it" is not that. A third is scoping the plan too narrowly to reduce testing burden, which creates a gap exactly where the regulator (and the business) cares most. Finally, when an organization sits under more than one regime at once (a healthcare company that also processes card payments, for example), the efficient answer is one program whose documentation maps to each regime's specific asks, not parallel plans that drift out of sync with each other.
Design a communications plan for major incidents that affect availability. Cover who needs updates (executives, customers, engineering, legal), how often, what a status-page update should say, and when you escalate from an engineering update to an executive one. Sketch out what the first sixty minutes of communication would look like for a critical outage.
Sample Answer
Direct answer
I'd build the plan around three things: a stakeholder map that separates internal audiences (engineering, executives, legal) from external ones (customers, partners, press) since they need different content and tone, a RACI matrix (who is Responsible, Accountable, Consulted and Informed) so it's unambiguous who is allowed to say what to whom, and a defined escalation threshold for when a routine engineering update becomes an executive one. The first sixty minutes follow a fixed cadence: detect, declare, and get an initial acknowledgment out fast, even before the cause is known, then update on a regular clock rather than only when there's new information.
Structured elaboration
Stakeholder map: internal vs. external.
| Audience | Content | Tone |
|---|---|---|
| Engineering (internal) | Technical symptoms, current hypothesis, what's being tried, what's needed from them | Direct, technical, fast-moving |
| Executives (internal) | Business impact, decision points that need their authority, reputational or legal exposure | Concise, framed around impact and decisions, not implementation detail |
| Legal and compliance (internal) | Anything with regulatory or contractual exposure, review of external language before it ships | Precise, focused on obligations and risk |
| Customers and partners (external) | What's affected, current status, workaround if any, next update time | Plain language, no internal jargon, no speculation about root cause until confirmed |
| Press (external, if triggered) | A single approved statement, routed only through the designated spokesperson | Careful, legally reviewed, minimal |
RACI for the response.
| Role | RACI |
|---|---|
| Incident Commander | Responsible for coordinating the technical response and deciding when thresholds are crossed |
| Communications Lead | Responsible for drafting and publishing every external update; the single pen holding the message |
| Executive sponsor | Accountable for approving customer-facing language once the escalation threshold is crossed |
| Legal | Consulted before any public statement that touches liability, data exposure, or contractual terms |
| Customer support | Informed, and executes pre-approved response templates for inbound customer questions |
Escalation thresholds, engineering to executive. A routine update stays at the engineering level until it crosses a defined line, for example greater than 25% of customers affected, any customer-visible data integrity risk, a duration past 30 minutes with no root cause identified, or any event that spans multiple regions or products at once. Crossing that line triggers the executive sponsor and legal into the loop immediately, not at the next scheduled update.
Cross-timezone coordination for multi-region events. When an incident spans regions, assign a single Communications Lead of record per shift with a documented, timestamped handoff (what's been said publicly, what's pending, what's the next scheduled update time) so a shift change never produces two conflicting messages. All approved language lives in one shared document that both shifts write to, not in separate regional threads.
Status page content. Header with incident ID and affected services, a plain-language impact summary, current status (investigating, identified, mitigating, resolved), a short list of actions taken, and a concrete next-update time rather than a vague "soon."
Escalation flow.
flowchart TD
A[Incident detected] --> B{Meets major-incident threshold?}
B -- No --> C[Engineering tracks it, no exec notify]
B -- Yes --> D[Incident Commander declares major incident]
D --> E[Internal update: scope, IC, ETA]
D --> F[Status page: Investigating]
E --> G{Unresolved at 30 min or revenue-affecting?}
G -- No --> H[Engineering-led updates every 30 min]
G -- Yes --> I[Escalate to executives and legal]
I --> J[Exec liaison approves customer-facing language]
J --> K[Status page and customer comms updated]
H --> L[Resolution and postmortem scheduled]
K --> L
Worked example
First sixty minutes of a critical, multi-region outage at a mid-size payments company:
- Minute 0: Incident Commander declares the major incident based on the on-call alert.
- Minute 5: Communications Lead publishes an initial status page entry: "We're investigating reports of failed transactions. Next update in 15 minutes." No cause is stated yet.
- Minute 15: Internal update to the executive sponsor and legal: current hypothesis, estimated affected percentage, and whether the threshold has been crossed. In this case it has, since more than 25% of customers in two regions are affected.
- Minute 20: Status page updated with the confirmed impact scope, still no cause speculation externally.
- Minute 40: A mitigation is applied; the Communications Lead updates internally and externally with early signs of recovery, careful to say "error rates are declining" rather than "resolved," since a partial rollback can regress.
- Minute 60: If recovery is confirmed, a resolution update is published with a commitment to a postmortem within 72 hours; if not, the next update time is restated and the executive sponsor is briefed on the extended timeline.
Trade-offs and pitfalls
Too frequent updates without new information create alarm fatigue and can look like the team doesn't have control; too infrequent updates erode customer trust faster than the outage itself does. Publishing a root cause before it's confirmed is a common and expensive mistake, since retracting it externally is more damaging than staying vague a little longer. A single point of failure in the plan itself is common: if the Communications Lead is unreachable, there must be a named backup, otherwise the whole cadence stalls. Finally, this is a crisis-communications discipline, not an incident-tooling configuration exercise: the plan should specify who says what to whom and when, not which paging tool routes the alert.
Walk me through how you'd actually conduct a Business Impact Analysis for a company's application portfolio. Who would you interview, what would you ask them to assess criticality, and what would the finished output look like?
Sample Answer
Direct answer
I'd run the business impact analysis (BIA), the exercise that quantifies what an outage costs each business function and sets its recovery targets, as a structured set of stakeholder interviews mapped against an inventory of the application portfolio, not a survey blast. For each application I'd talk to the business process owner (what breaks and how fast the harm grows), a downstream consumer of that application (what breaks for them), and a technical lead (what this application actually depends on, so the impact can be traced through the chain). The finished output is a BIA register: one row per application with its criticality tier, tolerance window, business-side recovery time objective (RTO, the target time to get a function working again) and recovery point objective (RPO, how much recent data the function can afford to lose, measured in time), and its upstream and downstream dependencies, signed off by the business owner.
Structured elaboration
1. Scope and inventory. Build the list of in-scope applications from the business side (which processes exist) rather than the IT asset list, so nothing gets excluded just because it lacks a system owner in the configuration management database (CMDB), the IT system of record for tracked assets.
2. Stakeholder interviews. Two distinct conversations per application, deliberately kept separate so the technical lead can't quietly set the business tolerance:
- With the business/process owner: "If this is unavailable for 1 hour, 4 hours, 24 hours, 3 days, what actually happens? Who notices, what's the financial or legal consequence, is there a seasonal or peak period where this gets worse, and is there a manual workaround?"
- With a downstream consuming team: "What do you depend on this for, and does your own tolerance window get eaten by the time this stays down?" This surfaces cascading impact that the primary owner often doesn't see.
- With the technical/ops lead: "What does this depend on upstream (data, other systems, third parties, specific teams) and who would need to be involved to bring it back?" This is a dependency map, not a design conversation; the goal is knowing what's connected, not deciding how to recover it.
3. Scoring and tiering. Combine a quantitative score (revenue or cost impact per elapsed-time band, contractual penalties) with a qualitative one (regulatory exposure, reputational harm, customer trust) using a shared rubric applied the same way across every interview, specifically to counter the fact that every owner tends to describe their own system as critical.
4. Output. A BIA register document: application name, business owner, criticality tier, maximum tolerable period of disruption (MTPD), target RTO/RPO, upstream and downstream dependencies, and the date of last review. This becomes the input to recovery sequencing and to the exercise program's scope.
Worked example
Three systems from a recent BIA at a mid-size company, showing why the interview-based approach catches things a pure system-criticality guess would miss:
| System | Owner's first guess | What the interviews revealed | Final tier |
|---|---|---|---|
| Customer-facing web application | "Obviously Tier 1" | Confirmed: drives signup and checkout revenue directly; no manual workaround exists | Tier 1, RTO 1 hour |
| Internal HR system | "Probably low priority" | Downstream interview with Payroll revealed payroll processing depends on data pulled from this system by day 3 of the month, and missing that window creates a wage-payment compliance issue in several jurisdictions | Tier 2, RTO 48 hours, with a hard deadline constraint tied to the pay cycle rather than a flat window |
| Analytics pipeline | "Low priority, it's just dashboards" | Confirmed genuinely low priority day-to-day, but the downstream interview with Finance revealed it feeds the monthly board reporting package, so its real constraint is a monthly deadline, not a continuous-availability one | Tier 3, RTO tied to reporting deadline (typically several days), not hours |
The HR and analytics examples are the reason the downstream-consumer interview is a separate step: the primary owner of both systems undersold their own criticality, and only the consuming team surfaced the real constraint.
Trade-offs and pitfalls
Interview-based BIAs are vulnerable to self-serving bias (every owner wants to be Tier 1 to guarantee resources), which is why a shared, applied-consistently rubric matters more than any individual interview. They're also vulnerable to missing cascading dependencies if you only interview the primary owner and skip the downstream-consumer conversation, as the HR and analytics examples show. A separate trap specific to a technical audience: it's tempting to let the technical-lead interview drift into "how would we actually achieve that RTO," which turns a business-impact exercise into a premature architecture discussion. Keep that conversation to what the system depends on, and leave the how-to-recover-it decision for the technical team afterward, informed by the tier the BIA assigned.
Unlock Full Question Bank
Get access to all 20 Disaster Recovery and Business Continuity interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.