Disaster Recovery and Business Continuity Questions
Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.
A regulatory examiner is about to review your business continuity program. What would you want to have ready to show them, and how do you make sure your documentation actually reflects a tested, current program rather than a plan that's been sitting untouched since it was written?
Sample Answer
Direct answer. I'd want a small, coherent evidence package proving three things: the program identifies what actually matters to the business (a current business impact analysis and criticality ranking), the plan has been tested against reality on a regular cadence with findings tracked to closure, and it's actively maintained rather than signed once and shelved. What makes documentation credible to an examiner is version history, dated exercise records with named participants and findings, and open corrective actions with owners and target dates, not how polished the plan document itself looks.
1. Program foundation documents
A current business impact analysis, or BIA (the process that identifies which business functions are critical and quantifies the cost of their downtime over time), is what the rest of the program's priorities trace back to. "Current" means dated and tied to a defined refresh cadence, not a one-time artifact from years ago. Alongside it: the continuity plan itself at the business-function level, describing what each critical function needs to keep operating or recover and who's responsible; and a declaration and activation framework naming who has the authority to formally invoke the plan and the escalation path if that person is unreachable.
2. Evidence the plan has been tested, not just written
Records from an exercise programme across different depths: a tabletop exercise (a facilitated discussion where the team talks through a scenario without actually executing anything), a functional exercise (a partial, real execution of specific procedures), and periodically a full-scale exercise, each with a date, scope, participants, and a documented outcome. A hotwash or after-action report for each one (the structured debrief run immediately afterward, capturing what worked, what didn't, and why), not just a pass or fail note. A corrective action log tying each finding back to a specific exercise, with an owner and a target date, and a record of which items actually closed versus which are still open. An examiner is generally less interested in whether every exercise went flawlessly and more interested in whether findings get followed through on.
3. Evidence the documentation reflects the business today, not the business as it was
Version control on the plan: a change log showing when it was last updated, by whom, and why, ideally tied to specific triggers (an org restructure, a new critical system, a new key supplier) rather than a periodic rubber stamp. A management review record showing leadership periodically reviewing the program's health, including whether resourcing and risk-acceptance decisions were actually made. And training and awareness records for the people named in the plan, since a plan naming a role that no longer exists, or a person who's left, is itself evidence the program isn't current.
4. Regulatory and compliance framing
These are examples of different regulatory regimes, not a checklist to memorize in full: which one applies to you, if any, depends on your sector and jurisdiction. Different regulators expect different levels of prescriptiveness. In the US, financial-sector examiners work from the Federal Financial Institutions Examination Council (FFIEC)'s IT Examination Handbook, which describes practices examiners use to evaluate business continuity management (BCM) governance, resilience strategy, training, exercises, and maintenance, scaled to an institution's size and complexity, rather than imposing a fixed checklist of requirements. The EU's Digital Operational Resilience Act (DORA) is more prescriptive for financial entities: it requires maintaining and periodically testing information and communication technology (ICT) continuity plans, with standard testing performed at least annually and more intensive testing on a longer cycle for larger institutions, and requires records of that testing to be kept. ISO 22301, the international management-system standard for business continuity, lays out the shape a mature program has structurally regardless of sector: a policy, a scope, a BIA and risk assessment, documented strategies and plans, an exercising programme, and the ordinary management-system disciplines of internal audit, management review, and corrective action closed out rather than just logged. Building the evidence package to that shape tends to answer the question before an examiner from any of these regimes asks it.
Worked example
A mid-size organization preparing for an examiner visit assembles: a BIA refreshed nine months earlier and due for its annual refresh in three months, listing the top twelve critical functions and their maximum tolerable downtime; the continuity plan for each function, with a change log showing two revisions in the past year, one triggered by a platform migration and one by a reorganization of the team owning a critical function; records of two tabletop exercises and one functional exercise in the past twelve months, each with a hotwash summary; a corrective action log showing eleven findings raised across those exercises, nine closed with evidence and two still open with named owners and dates inside the next quarter; and quarterly risk-committee minutes showing the program's status was reviewed and the two open items were acknowledged with a committed resourcing decision. That combination, current inputs, a tested cadence, and traceable follow-through, is what makes a program read as alive rather than as a binder written once and never reopened.
Trade-offs & pitfalls. A beautifully written plan with no exercise history behind it is a bigger red flag to an experienced examiner than a rougher plan with a real testing record, since untested plans reliably fail in ways nobody predicted. Closing corrective actions by simply marking them closed without evidence is worse than leaving them visibly open, because an examiner who spot-checks one and finds it wasn't actually fixed calls the whole log's credibility into question. And treating the BIA as a one-time project rather than a maintained artifact is the most common gap of all: businesses change what's critical to them faster than a plan written once accounts for, and an examiner asking how you know this is still accurate is really asking whether a refresh cadence exists at all.
Runbooks and continuity documentation go stale fast once systems, teams, and org structure keep changing. How do you keep them accurate over time? Cover ownership, versioning, and how you'd catch drift before it matters during a real event rather than after.
Sample Answer
Direct answer. Ownership has to be a named role tied to the business function or system the runbook covers, not "whoever wrote it last," with versioning that ties each revision to a specific trigger, not just a date. Catching drift before a real event means testing the runbook during a planned exercise instead of discovering it's wrong while executing it under pressure, paired with a lightweight review trigger whenever something the runbook depends on changes.
1. Ownership
Assign a named owner to every runbook, ideally the person or role accountable for the business function or system it covers, not a shared team inbox and not necessarily the original author, who may move on. Ownership means being responsible for the runbook's accuracy, not personally executing every step during a real event. Make it visible on the document itself (name, role, last-reviewed date) so anyone opening it during an incident can see who to ask if something looks wrong, and so a stale "owner" who's left the organization is an obvious red flag rather than a hidden one. When ownership changes hands, require an explicit handoff with a fresh review as part of the transition, not a silent reassignment.
2. Versioning
Treat runbooks as controlled documents: every change gets a version number, a date, who made it, and why, kept in the document's own change history rather than relying on file metadata or institutional memory. Tie versions to triggers, not just a calendar. A scheduled review is a reasonable baseline, but the more reliable trigger is linking review to the events that actually make a runbook wrong: the system it depends on changes, the org structure around it changes, or a real exercise or incident surfaces a gap. A runbook untouched for eighteen months right after the system it describes was rebuilt is far more suspect than one reviewed on schedule six months ago. Keep prior versions accessible, not just the latest, so a post-event review can tell whether a gap that showed up was already known and fixed but hadn't propagated to whoever was executing, or was genuinely new.
3. Catching drift before it matters
The strongest mechanism is using the runbook for real, on a schedule, inside an exercise (a tabletop discussion or a functional exercise that actually executes some steps), before a live event forces the discovery. A runbook nobody has walked through since it was written is a hypothesis, not a tested procedure. Build a lightweight checklist step into whatever process generates the kind of change that breaks runbooks: when a system change, a supplier change, or an org change is being planned, that process flags which runbooks might be affected, so the review happens close to the change instead of being discovered much later. Track a simple staleness signal across the whole runbook set (time since last review, time since the system or team it covers last changed) and report on it the way you'd report any other risk metric, so an accumulating backlog of unreviewed runbooks is visible to whoever owns the continuity program overall, not discovered runbook by runbook.
Worked example
A payments-processing runbook is owned by the payments platform lead, currently at version 4: v2 added a fallback provider, v3 corrected an escalation contact list after a reorg, v4 followed a functional exercise that found the documented rollback step no longer matched how the system actually rolled back. Six months after v4, the team migrates the payments platform to new infrastructure as part of an unrelated project. Because "infrastructure migration for a system with an owned runbook" sits on that project's change-review checklist, the payments platform lead is flagged automatically and confirms the runbook needs an update before the migration ships, catching the drift as part of the planned change rather than during the next incident. At the next scheduled functional exercise three months later, the team walks through the updated runbook end to end and finds one further step, a monitoring dashboard link, still pointing at the decommissioned infrastructure. That becomes a logged finding and v5.
Trade-offs & pitfalls. A purely calendar-based review cadence catches drift too slowly for a fast-changing system and too often for a stable one; tying review to real change triggers is more work to set up but matches effort to actual risk. Ownership without accountability for staleness (a name on a document nobody ever checks) is barely better than no ownership; the staleness signal has to surface to someone, not just exist. And testing a runbook only during a real event is the worst possible time to discover it's wrong: the exercise programme is what lets a team find the gap on an ordinary afternoon instead of during an actual outage.
Design an exercise that tests whether your fallback routes for a critical third-party dependency, like a payment processor or an SMS provider, actually work when that dependency goes down during a simulated failover. What would success look like, and what's your rollback plan if the fallback itself misbehaves?
Sample Answer
Direct answer. Design it as a functional exercise, a partial, real execution of the actual failover rather than just a tabletop discussion, run against a simulated outage of the dependency in a scoped, non-production environment. Success means the fallback route delivered what the business actually needs (transactions completed or safely queued, no duplicates, customers not silently left in the dark) within criteria set before the exercise runs. The rollback plan is a predefined abort trigger that reverts to the primary path and treats anything the fallback processed as needing reconciliation, exactly like a real incident would.
1. Scope and objectives
Name the specific dependency and specific fallback route under test (a secondary SMS provider, or a manual and queued path for payment confirmation), and state the objective in business terms: can the organization keep this function operating, at what level of degradation, for how long, if this specific dependency is unavailable. Choose the exercise type deliberately: a tabletop validates whether people understand the plan; a functional exercise, actually triggering the fallback and running real transactions through it in a controlled setting, validates whether it technically and operationally works. For a fallback route you're specifically trying to prove works under failure, a functional exercise is the right depth; a tabletop alone would only tell you the plan sounds right.
2. Design the scenario
Simulate the outage realistically but under control: route the primary dependency's calls to a fault response (timeouts, errors, or an explicit "unavailable" condition) in a scoped test environment with synthetic transactions and test accounts, never live customer data or real charges. Run more than one failure shape if the plan is meant to handle more than one: a hard outage behaves differently from a slow degradation (rising latency, intermittent errors), and can expose different gaps, particularly around whether the system correctly detects "bad enough to fail over" instead of tolerating it too long. Include volume, not a single transaction: a batch of concurrent test transactions through the fallback route surfaces queuing, ordering, and duplicate-handling behavior that one transaction at a time won't.
3. Define success before running it, not after
Business-facing criteria: every transaction attempted during the exercise either completes successfully via the fallback or ends up in a clearly tracked pending state, with zero duplicates and zero silently lost transactions, confirmed by reconciling the exercise's own transaction log against what the fallback provider and the primary system each show afterward. Operational criteria: the team followed the documented runbook to trigger the fallback rather than an ad hoc workaround, within whatever time target the plan specifies for detecting and switching over. Communication criteria, where the plan calls for customer-facing messaging during a real event of this kind: confirm the message that would have gone out is accurate to the actual state of the exercise's transactions, not a generic template nobody checked against what happened.
4. Rollback plan if the fallback itself misbehaves
Define the abort trigger ahead of time, not improvised mid-exercise: a concrete condition (error rate above a set threshold, a data-integrity check failing, an evidently wrong amount) that stops the exercise and reverts test traffic to the primary path, rather than letting a misbehaving fallback keep running and generating more to clean up. Treat anything the fallback processed before the abort exactly like a real incident: reconcile it, don't discard it, since finding out whether reconciliation actually works is as much the point of the exercise as testing the happy path. Have a named exercise controller with the authority to call the abort, separate from whoever is executing the steps, so the decision to stop isn't made by someone invested in proving the fallback works.
5. Close the loop
Run a hotwash immediately after (a short structured debrief on what worked, what didn't, and why) and log findings the way a real incident's post-event review would, with owners and target dates, so a gap the exercise found gets fixed and is retested at the next exercise rather than only known about.
Worked example
An organization tests its fallback SMS provider for one-time passcode delivery. Objective: if the primary provider is unavailable, can the fallback deliver codes within the plan's 30-second target with no duplicate sends. Design: a functional exercise, with primary-provider calls routed to a fault-injection endpoint returning timeouts for a fixed 20-minute window, run against 50 synthetic test accounts requesting codes continuously through that window, in a test environment with a provider sandbox that never sends real messages to real numbers. Success criteria set beforehand: at least 49 of the 50 attempts (allowing one edge case at the exact failover boundary) each receive exactly one code via the fallback within 30 seconds, and zero receive two.
Result: 50 of 50 delivered exactly once, averaging 22 seconds. But the exercise log shows the system took 40 seconds to detect the simulated outage and switch to the fallback in the first place, which isn't a delivery failure but eats into the 30-second budget for any request arriving in that detection window, and is logged as a finding. Nothing crossed the predefined abort trigger (an error rate above 5%), so the exercise ran to completion without needing rollback; the detection-time finding is written up with an owner and a target date for the next exercise to confirm it's fixed.
Trade-offs & pitfalls. Testing only the happy path, one transaction, no volume, is the most common shortcut, and it's exactly what misses the ordering and duplicate-handling problems that show up under concurrent load. Setting the success bar as "the fallback works at all" instead of "the fallback meets the actual business requirement" lets a technically-functioning fallback pass an exercise while still being unacceptable in a real event. And skipping the predefined abort trigger on the assumption an exercise "shouldn't need it" is how a test turns into a real incident; the rollback plan is itself part of what's being tested, not a formality around it.
You've just opened a war room for a high-severity disaster recovery event involving network, database, application, and customer-support teams. How do you structure it? Cover the physical or virtual setup, communication channels, what you'd want visible on shared dashboards, and how decisions get documented in real time.
Sample Answer
Direct answer
A continuity command center needs a small decision-making core insulated from a larger information-sharing periphery: one room, physical or virtual, for the people actually making calls; a mirrored read-only stream for everyone who needs situational awareness without adding noise; and a single written record that becomes the account of what happened and why. Getting this wrong, everyone crammed into one unstructured call, turns a recoverable event into a communication failure layered on top of the original disruption.
Physical or virtual setup
- Primary decision room: in person or a video bridge, limited to the commander, function leads, and the scribe.
- Broadcast bridge or channel: a larger listen-only stream for stakeholders who need visibility (support, other business units) without talking over the working group.
- Private channel: reserved for anything sensitive, legal exposure or workforce safety, that shouldn't go into the general information stream.
Communication channels, matched to purpose
| Channel | Purpose |
|---|---|
| Working channel | The decision group's live coordination line |
| Broadcast channel | One-way status pushed to the wider organization at the agreed cadence |
| Escalation channel | A dedicated path to reach executives or backup decision-makers |
Who's in the room
Beyond the technical responders, for anything above a minor severity the room includes a legal representative (contractual and regulatory exposure), a finance representative (cost-of-mitigation trade-offs, insurance notification triggers), and a PR/communications representative (external messaging, media exposure), present from early activation rather than read in after a decision has already been made.
A formal escalation-level structure decides who's paged and required present, rather than leaving it to a judgment call live. For example:
| Level | Meaning | Who's required present |
|---|---|---|
| 1 | Contained, business-as-usual response | Function lead only |
| 2 | Material customer or functional impact | Commander, function leads, comms lead |
| 3 | Enterprise-wide impact | Above, plus legal, finance, PR, and executive sponsor |
What belongs on shared dashboards
The room needs what the business uses to decide, not raw technical telemetry: which business services or functions are impacted and to what degree, estimated time to the next update, current escalation level, open decisions and their owners, and a running list of customer or stakeholder commitments made so far, so nobody promises something twice or contradicts an earlier statement. Deep technical monitoring stays with the technical responders in their own tooling; the war room's shared view is the business-impact layer above it.
Real-time decision documentation
A single scribe, not the commander, not a function lead, maintains one shared, timestamped log for the whole event: each decision, who made it, the reasoning, and what was communicated as a result. This becomes the record for the post-event review and, when legal or regulatory exposure exists, the evidentiary record of due diligence.
Worked example
A regional retailer's order-management platform fails during a holiday sale. The event is declared Level 2, which automatically pages the legal liaison (payment processing is affected) and the finance lead (evaluating emergency vendor spend). The shared dashboard reads "Order Management: severely degraded, roughly 40% of order volume affected, next update in 20 minutes, escalation Level 2," not a stream of raw error-rate graphs, which stay with engineering's own monitoring. The scribe logs the decision to pause promotional email sends at 14:05 with the reasoning ("avoid driving more traffic into a degraded checkout"), so that decision doesn't have to be re-explained in the post-event review.
Trade-offs and pitfalls
- Putting everyone who might need information into the same working channel as the decision-makers drowns the room in cross-talk; the broadcast stream exists specifically to prevent this.
- Bringing legal, finance, or PR in only after a decision is made, rather than having them present from early activation, is the most common failure mode this structure guards against; by the time they're read in, commitments may already be hard to walk back.
- A dashboard full of technical detail feels informative but slows business decision-making by forcing non-technical stakeholders to interpret raw signals. Keep the shared view at the business-impact layer.
- A decision log reconstructed afterward from memory, instead of maintained live, loses the accuracy that makes it useful for both the postmortem and any external accountability.
Once you have BIA findings for a set of business services, how do you translate that into recovery-priority tiers, and what actually determines whether a service lands in the top tier versus the bottom one? Walk through how those tiers then drive budget and staffing decisions.
Sample Answer
Direct answer
Tiering translates BIA (business impact analysis) findings into a small number of recovery-priority bands, using the business's tolerance for downtime and data loss as the primary axis, not technical difficulty. A service lands in the top tier because sustained disruption threatens revenue, safety, legal or regulatory standing, or a large share of customers within a very short window, not because it happens to be technically easiest to protect. The tiers then become the lever finance and staffing use to decide where headroom, retainer contracts, and dedicated on-call coverage get funded.
From BIA findings to tier criteria
- MTPD / MAO (maximum tolerable period of disruption, also called maximum acceptable outage): the longest a service can be down before consequences become unacceptable to the business. This is the anchor input from the BIA.
- Combine MTPD with: financial loss per hour of disruption, regulatory or contractual exposure, safety impact, breadth of customers or users affected, and dependency fan-out (how many other services or business functions break if this one is down).
- Score, don't guess: weight the BIA findings for each service across these dimensions. Example weighting for illustration only (the actual weights are a business decision the BIA sponsor and finance sign off on, not something one team sets unilaterally): financial impact 35%, regulatory/legal exposure 25%, customer scope 25%, dependency fan-out 15%.
Translating scores into tiers
| Tier | Recovery expectation (MTPD-driven) | What lands here | Staffing / budget posture |
|---|---|---|---|
| 1 | Minutes to a few hours | Revenue-critical or safety/regulatory-critical, broad customer exposure | Dedicated on-call rotation, pre-funded standby resources, exercised quarterly |
| 2 | Several hours to one business day | Degraded operation tolerable briefly, moderate exposure | Named on-call owner, exercised semi-annually |
| 3 | One to several business days | Internal or narrow-scope, workaround exists | Best-effort recovery, reviewed annually |
| 4 | No fixed near-term deadline | Low-impact, batch, or back-office | Recovery scheduled opportunistically |
What actually pushes a service to the top isn't "sounds important," it's the combination of speed of consequence onset (does damage start accruing in minutes) and irreversibility (can the loss be made up later, like a delayed nightly batch, or is it gone for good, like a breached SLA credit or a safety incident). A service with high total impact but slow-accruing consequence, a monthly reporting job that only starts costing money after a week of delay, is often Tier 2, not Tier 1, even if its total dollar impact looks large on paper.
How tiers drive budget and staffing
- Rotas: Tier 1 justifies funded, cross-trained on-call rotations with named backups; Tier 3/4 rely on the general on-call pool at best effort.
- Standby spend: Tier 1 gets pre-approved budget for whatever standby resources the business decided it needs; this is where the tiering output hands off into the technical recovery-strategy conversation, but the tier itself is a business-impact call, not an architecture decision.
- Exercise cadence: higher tiers are drilled more often, because the cost of an untested plan scales with the cost of the outage it exists to prevent.
- Supplier spend: Tier 1 is where organizations pay for expedited third-party support contracts and priority escalation paths.
- Governance: tiers get reviewed at a fixed cadence (at least annually, or whenever the BIA is refreshed) and signed off by the business owner, because a service's tier is a statement about acceptable business risk, not just an IT classification. This is the same discipline external frameworks like the ISO 22301 business continuity standard expect: BIA and risk assessment as a documented, periodically reviewed input to recovery prioritization, without the standard itself prescribing specific weights or tier counts; those stay organization-specific.
Worked example
A mid-size B2B SaaS company scores three services 0-10 on each dimension using the weighting above (financial 0.35, regulatory 0.25, customer scope 0.25, dependency fan-out 0.15), with band cutoffs of Tier 1 at 7.5+, Tier 2 at 5.0-7.49, Tier 3 at 2.5-4.99, Tier 4 below 2.5:
- Payment processing: financial 9, regulatory 9, customer scope 8, dependency 7.
9(0.35)+9(0.25)+8(0.25)+7(0.15)=3.15+2.25+2.00+1.05=8.45
Composite 8.45 lands in Tier 1. - Customer support ticketing: financial 5, regulatory 4, customer scope 6, dependency 4.
5(0.35)+4(0.25)+6(0.25)+4(0.15)=1.75+1.00+1.50+0.60=4.85
Composite 4.85 lands in Tier 3. - Internal expense-reporting tool: financial 2, regulatory 1, customer scope 1, dependency 2.
2(0.35)+1(0.25)+1(0.25)+2(0.15)=0.70+0.25+0.25+0.30=1.50
Composite 1.50 lands in Tier 4.
Consequence for staffing and budget: payment processing (Tier 1) gets a funded dedicated on-call rotation and a pre-approved contract with a backup payment partner; the expense tool (Tier 4) is recovered on a best-effort basis by the general helpdesk queue with no dedicated budget line.
Trade-offs and pitfalls
- Letting engineering effort quietly redefine tiers ("it's already highly available so it must be Tier 1") inverts the logic. Tiering has to stay anchored to business impact, not to what's already been built.
- Too many tiers dilutes the prioritization signal; two to four is the usual practical range. Beyond that, tiering stops driving clear resourcing decisions.
- Scoring once and never revisiting is a common failure: business models shift (an internal tool becomes customer-facing) and tiering needs the same refresh cadence as the BIA itself, not a one-time exercise.
- Weighting choices are political. Finance, legal, and the business owner need to agree the weights before scores are trusted, or every re-score becomes a renegotiation.
- A score sitting near a boundary, like the ticketing example above, deserves judgment, not automatic sorting; treat cutoffs as a starting point, not a verdict.
Unlock Full Question Bank
Get access to all 20 Disaster Recovery and Business Continuity interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.