Disaster Recovery and Business Continuity Questions
Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.
Who should actually have the authority to declare a disaster and invoke the continuity plan? Design the governance around that decision: how backup declaration authority is documented if the primary decision-maker is unavailable, what gets recorded when the call is made, and how you'd build in an emergency-override path without losing accountability.
Sample Answer
Direct answer
Declaration authority should sit with a small, named list of people, in an explicit order of succession, not with whoever happens to be senior and reachable when something breaks. The plan documents a primary decision-maker plus one or two backups by role, spells out the specific conditions under which any of them can declare, and requires every declaration to be logged with who made the call, when, and why, so authority stays traceable even when the primary is unreachable.
Why a single named owner isn't enough
Disasters don't wait for the primary decision-maker's calendar. A continuity plan needs an order of succession: a ranked list of who can declare if the primary is unreachable, with each person's authority spelled out in advance rather than assumed. Order of succession is one specific component of a broader continuity of operations (COOP) program, not a synonym for it: a COOP program is the organization's overall plan for keeping its essential functions running, and getting them fully restored, through and after a disruption. FEMA's continuity guidance lists order of succession alongside essential functions (the small set of functions that cannot lapse without unacceptable harm) and delegations of authority (who can act on the primary decision-maker's behalf, and for what) as three of several required components; the rest cover facilities, communications, records, staffing, and exercising, and matter for a full COOP program but aren't the focus here. Two or three levels deep is usually enough; more layers slows the decision the structure exists to speed up.
What "declare" actually authorizes
Declaration isn't a status update. It's the trigger that activates the continuity plan itself: it mobilizes the response team, authorizes emergency spend, and starts the plan's communication and recovery sequencing. Because it triggers real commitments, the authority to make that call has to be explicit, not inferred from a job title.
Documenting the succession
- Name roles, not just people, since people change jobs and the document shouldn't need rewriting every reorg.
- Define the specific trigger conditions that justify declaration, tied back to the criticality tiering and BIA findings (for example: "any Tier 1 service confirmed down beyond the tier's trigger duration with no clear near-term fix").
- Define how the next person in line is notified and confirms they're stepping in if the primary can't be reached within an agreed window.
What gets recorded when the call is made
Timestamp, who declared, their role and authority basis, the specific trigger condition that was met, and the initial scope of the declaration. This record does two jobs: it lets the response team act without re-litigating whether they're authorized, and it becomes the accountability record if the declaration is later questioned as premature or too slow.
Emergency override without losing accountability
The hard part is letting someone act fast in a genuine emergency without creating a loophole where authority quietly evaporates. The standard resolution has two parts:
- Anyone in the pre-defined succession list can declare unilaterally under defined trigger conditions, no committee vote required, because waiting for consensus is itself a cost.
- Every emergency declaration requires mandatory retroactive review by the executive sponsor or a governance body within a fixed window (commonly 24-72 hours) to confirm it was within policy; anything outside the pre-defined triggers gets flagged specifically for that review.
The override buys speed. The retroactive review is what keeps it from becoming an accountability gap. This is also the kind of structure external frameworks like the ISO 22301 business continuity standard expect organizations to have documented, though the exact succession depth and trigger conditions are always organization-specific, not prescribed by the standard itself.
Worked example
As the on-call incident commander for a regional outage at 2 a.m., I'm second in the documented order of succession behind the VP of Engineering, who isn't reachable within the 15-minute window the plan specifies. The signals I weigh before declaring: is the affected service confirmed Tier 1 per the BIA-derived tiering, has the outage crossed the pre-defined trigger duration, and is there a credible near-term fix that would make declaration premature. No single signal is enough on its own. Once it's Tier 1 and the trigger duration has passed with no credible near-term fix, I declare. I don't wait for the VP's sign-off first, because the plan already gave me that authority under exactly this condition. I log the declaration with the timestamp, the trigger met, and my reasoning, and the VP reviews it retroactively at the next check-in. Balancing speed against risk here means trusting the pre-agreed trigger conditions rather than re-deriving the judgment call from scratch under pressure; the trigger conditions are where the careful thinking already happened, in a calmer moment, before the outage.
Trade-offs and pitfalls
- A succession list naming specific individuals rather than roles goes stale the moment someone changes jobs. Tie authority to role, and keep the roster of who currently holds each role actually maintained.
- Vague trigger conditions ("declare when it seems bad enough") push the real decision back onto individual judgment under pressure, exactly what the structure is meant to remove.
- Making the retroactive review punitive discourages people from declaring when in doubt, which is worse than an occasional over-declaration. The review should confirm policy adherence, not assign blame.
- Too shallow a succession list, only one backup, recreates the single-point-of-failure problem the structure exists to solve; too deep a list slows the decision with unnecessary hierarchy.
You've just opened a war room for a high-severity disaster recovery event involving network, database, application, and customer-support teams. How do you structure it? Cover the physical or virtual setup, communication channels, what you'd want visible on shared dashboards, and how decisions get documented in real time.
Sample Answer
Direct answer
A continuity command center needs a small decision-making core insulated from a larger information-sharing periphery: one room, physical or virtual, for the people actually making calls; a mirrored read-only stream for everyone who needs situational awareness without adding noise; and a single written record that becomes the account of what happened and why. Getting this wrong, everyone crammed into one unstructured call, turns a recoverable event into a communication failure layered on top of the original disruption.
Physical or virtual setup
- Primary decision room: in person or a video bridge, limited to the commander, function leads, and the scribe.
- Broadcast bridge or channel: a larger listen-only stream for stakeholders who need visibility (support, other business units) without talking over the working group.
- Private channel: reserved for anything sensitive, legal exposure or workforce safety, that shouldn't go into the general information stream.
Communication channels, matched to purpose
| Channel | Purpose |
|---|---|
| Working channel | The decision group's live coordination line |
| Broadcast channel | One-way status pushed to the wider organization at the agreed cadence |
| Escalation channel | A dedicated path to reach executives or backup decision-makers |
Who's in the room
Beyond the technical responders, for anything above a minor severity the room includes a legal representative (contractual and regulatory exposure), a finance representative (cost-of-mitigation trade-offs, insurance notification triggers), and a PR/communications representative (external messaging, media exposure), present from early activation rather than read in after a decision has already been made.
A formal escalation-level structure decides who's paged and required present, rather than leaving it to a judgment call live. For example:
| Level | Meaning | Who's required present |
|---|---|---|
| 1 | Contained, business-as-usual response | Function lead only |
| 2 | Material customer or functional impact | Commander, function leads, comms lead |
| 3 | Enterprise-wide impact | Above, plus legal, finance, PR, and executive sponsor |
What belongs on shared dashboards
The room needs what the business uses to decide, not raw technical telemetry: which business services or functions are impacted and to what degree, estimated time to the next update, current escalation level, open decisions and their owners, and a running list of customer or stakeholder commitments made so far, so nobody promises something twice or contradicts an earlier statement. Deep technical monitoring stays with the technical responders in their own tooling; the war room's shared view is the business-impact layer above it.
Real-time decision documentation
A single scribe, not the commander, not a function lead, maintains one shared, timestamped log for the whole event: each decision, who made it, the reasoning, and what was communicated as a result. This becomes the record for the post-event review and, when legal or regulatory exposure exists, the evidentiary record of due diligence.
Worked example
A regional retailer's order-management platform fails during a holiday sale. The event is declared Level 2, which automatically pages the legal liaison (payment processing is affected) and the finance lead (evaluating emergency vendor spend). The shared dashboard reads "Order Management: severely degraded, roughly 40% of order volume affected, next update in 20 minutes, escalation Level 2," not a stream of raw error-rate graphs, which stay with engineering's own monitoring. The scribe logs the decision to pause promotional email sends at 14:05 with the reasoning ("avoid driving more traffic into a degraded checkout"), so that decision doesn't have to be re-explained in the post-event review.
Trade-offs and pitfalls
- Putting everyone who might need information into the same working channel as the decision-makers drowns the room in cross-talk; the broadcast stream exists specifically to prevent this.
- Bringing legal, finance, or PR in only after a decision is made, rather than having them present from early activation, is the most common failure mode this structure guards against; by the time they're read in, commitments may already be hard to walk back.
- A dashboard full of technical detail feels informative but slows business decision-making by forcing non-technical stakeholders to interpret raw signals. Keep the shared view at the business-impact layer.
- A decision log reconstructed afterward from memory, instead of maintained live, loses the accuracy that makes it useful for both the postmortem and any external accountability.
How do you factor a critical vendor's own continuity posture into your DR and business continuity planning? Cover how you'd assess whether they can actually meet the recovery commitments you're relying on, and what you do when a single vendor's downtime would take down something you promised to keep running.
Sample Answer
Direct answer
I treat a critical vendor as an extension of the organization's own criticality tiering, not as someone else's problem once a contract is signed. That means verifying the vendor's own continuity posture with evidence rather than a sales claim, classifying how much concentration risk it represents (how many of our functions have no alternative path if it fails), negotiating contractual continuity obligations up front, and deciding, before anything breaks, what the business needs to happen when the vendor is down. That last decision belongs to the business-process owner who depends on the vendor, not to whichever engineer is on call when the outage starts.
Structured elaboration
Verifying the vendor's own continuity posture. Ask for evidence, not assurance: the date and outcome of their most recent continuity or DR test, a relevant attestation (SOC 2 Type II, ISO 22301 or 27001 certification), their own sub-vendor dependencies (a vendor who is itself entirely dependent on one cloud region is a risk you've now inherited), and their actual incident history rather than their marketing page's uptime claim.
Classifying concentration risk, "single point of failure" in the supplier sense. This means a vendor whose outage stops multiple business functions or customers because there is no alternative path, the same idea as a technical single point of failure, just applied to a supplier relationship rather than a system component. Tier it the way a business impact analysis (BIA), the exercise that quantifies what an outage costs each business function and sets its recovery targets, tiers internal functions:
| Tier | Definition | Example posture |
|---|---|---|
| Tier 1 | No alternative path exists; outage stops a revenue-critical or safety-critical function immediately | Requires the deepest due diligence, a tested manual workaround, and often a qualified secondary vendor |
| Tier 2 | Alternative exists but is degraded or manual | Documented workaround required, secondary vendor optional depending on cost |
| Tier 3 | Redundant paths already exist, or the function can tolerate extended downtime | Standard contract terms, lighter-touch monitoring |
Contractual levers. Recovery commitments (the vendor's own stated recovery time objective (RTO, the target time to get a function working again) and recovery point objective (RPO, how much recent data the function can afford to lose, measured in time) as contract terms), a notification SLA for when the vendor itself has an incident, audit or right-to-test clauses, and exit or data-portability terms so switching away from the vendor is actually possible rather than theoretical. SLA credits are worth including but shouldn't be mistaken for protection: a financial credit compensates the contract, it doesn't recover a lost transaction or a missed regulatory deadline.
The business-level fallback decision. For each Tier 1 or Tier 2 vendor, the business-process owner decides, in advance, one of: accept the risk (documented, with a name and a date attached to that decision so it's revisited, not forgotten), qualify and periodically re-test a secondary vendor, stand up a documented manual or degraded-mode workaround with a stated capacity limit, or insource the function. What the technology needs to deliver follows from that decision; the decision itself is a business call informed by, not delegated to, engineering.
Worked example
A mid-size payments company depends on a single external identity-verification vendor to open new customer accounts. The concentration risk is total: if new-account signups drive roughly $80,000/day in first-year revenue and identity verification is the only path to opening an account, then every hour of vendor downtime puts a proportional share of that revenue at risk:
$80,000/24≈$3,333 per hour at riskGiven that, the business-process owner (not engineering) sets the fallback: a manual review process run by the compliance team, using existing documentation to verify identity for a capped number of applications per hour, explicitly rated to absorb up to 4 hours of vendor downtime before the queue backs up faster than compliance can clear it. That 4-hour figure is the business's own RTO for this dependency, arrived at the same way a BIA sets any other recovery target, and it's what tells engineering what the fallback integration actually needs to support, not the other way around.
Trade-offs and pitfalls
SLA credits create a false sense of protection: a contractual credit worth a few hundred dollars does nothing to offset a day with $3,333/hour of exposure sitting behind it. Auditing every vendor to the same depth is unrealistic, which is exactly why the concentration-risk tiering has to gate how much due-diligence effort a given vendor gets. A "qualified backup vendor" that was tested once at onboarding and never revisited is a paper mitigation, not a real one, since vendor capabilities and your own dependency on them both drift over time. The trap specific to a technical audience: letting engineering quietly build a technical workaround to a second vendor without the contractual, compliance, and business-tolerance work behind it creates a shadow dependency that nobody has actually assessed or approved.
Before a disaster recovery event ever happens, what stakeholder roles and communication structure should already be defined? Walk through who owns which responsibility and how the update cadence should work.
Sample Answer
Direct answer
Before any event happens, the organization needs a named, rehearsed structure for who acts, who decides, and who talks to whom, because working this out for the first time mid-crisis is how communication chaos happens. At minimum: one clearly designated decision-maker per event, a communications owner who is a different person from whoever is doing the operational recovery work, and a pre-agreed update cadence for each audience.
Roles that need to exist before an event
- Continuity commander: single point of decision authority for the event's duration. Who this is, and how backup authority is documented if they're unreachable, is its own governance question, declaration authority (who has the standing to formally declare and activate the plan); what matters here is that the role exists and is named in advance.
- Communications lead: separate from the people fixing the problem, so the team doing operational recovery isn't also drafting customer-facing language under pressure.
- Function or service leads: one named owner per affected business function, reporting status and progress into the commander.
- Executive sponsor: kept informed, and consulted on higher-stakes trade-offs depending on severity (for example, accepting a longer outage versus authorizing emergency mitigation spend).
- Legal or compliance liaison: engaged whenever an event may carry regulatory notification obligations, contractual exposure, or litigation risk, so anything going external gets reviewed first.
- Scribe: keeps a timestamped decision log, separate from the commander's job, so no one is both making decisions and writing the record at the same time.
Communication structure by audience
| Audience | What they need | Typical cadence |
|---|---|---|
| Responders / internal technical teams | Granular, real-time status | Continuous, in the working channel |
| Executives | Business impact, ETA, decisions needed from them | Fixed interval (e.g., every 30-60 minutes) or on material change |
| Customers / external | Plain-language status, no internal detail | At defined checkpoints: acknowledgment, updates, resolution |
| Regulators / legal (when applicable) | Factual, reviewed language, notification-timeline aware | Per the applicable compliance obligation, not ad hoc |
Cadence gets committed to in advance, not decided live: pick a default interval per audience, and let the communications lead depart from the default when conditions genuinely warrant it, rather than negotiating frequency mid-event. What makes that cadence achievable under pressure is pre-built material: templated messages for each audience and status ("investigating," "mitigating," "resolved," "follow-up to come"), a pre-identified channel per audience, and a standing decision-log format the scribe isn't inventing on the fly.
Worked example
A mid-size retailer's checkout service fails during a promotional weekend. Because roles were pre-assigned, the on-call engineering lead isn't the one fielding calls from the VP of Sales; the communications lead handles that, posting a pre-written "investigating" update within minutes and following the pre-agreed cadence rather than waiting for engineering to have spare attention. The legal liaison is looped in immediately because payment processing is affected, even though no breach has occurred, purely because that trigger condition was defined in advance rather than judged in the moment.
Trade-offs and pitfalls
- If the same person both fixes the problem and writes the customer update, both jobs suffer. The split has to be a standing structure, not an improvisation invented under pressure.
- Over-specifying cadence for minor events creates alert fatigue and unnecessary overhead; scale the structure to the declared severity level instead of applying one rhythm to everything.
- A structure that only lives in a document nobody has read is no better than none. It has to be exercised so people know their role before they're asked to play it for real.
- Looping in legal or compliance "when needed," without a pre-defined trigger, tends to mean they're looped in too late. Define the trigger condition in advance, not in the moment.
Design an exercise that tests whether your fallback routes for a critical third-party dependency, like a payment processor or an SMS provider, actually work when that dependency goes down during a simulated failover. What would success look like, and what's your rollback plan if the fallback itself misbehaves?
Sample Answer
Direct answer. Design it as a functional exercise, a partial, real execution of the actual failover rather than just a tabletop discussion, run against a simulated outage of the dependency in a scoped, non-production environment. Success means the fallback route delivered what the business actually needs (transactions completed or safely queued, no duplicates, customers not silently left in the dark) within criteria set before the exercise runs. The rollback plan is a predefined abort trigger that reverts to the primary path and treats anything the fallback processed as needing reconciliation, exactly like a real incident would.
1. Scope and objectives
Name the specific dependency and specific fallback route under test (a secondary SMS provider, or a manual and queued path for payment confirmation), and state the objective in business terms: can the organization keep this function operating, at what level of degradation, for how long, if this specific dependency is unavailable. Choose the exercise type deliberately: a tabletop validates whether people understand the plan; a functional exercise, actually triggering the fallback and running real transactions through it in a controlled setting, validates whether it technically and operationally works. For a fallback route you're specifically trying to prove works under failure, a functional exercise is the right depth; a tabletop alone would only tell you the plan sounds right.
2. Design the scenario
Simulate the outage realistically but under control: route the primary dependency's calls to a fault response (timeouts, errors, or an explicit "unavailable" condition) in a scoped test environment with synthetic transactions and test accounts, never live customer data or real charges. Run more than one failure shape if the plan is meant to handle more than one: a hard outage behaves differently from a slow degradation (rising latency, intermittent errors), and can expose different gaps, particularly around whether the system correctly detects "bad enough to fail over" instead of tolerating it too long. Include volume, not a single transaction: a batch of concurrent test transactions through the fallback route surfaces queuing, ordering, and duplicate-handling behavior that one transaction at a time won't.
3. Define success before running it, not after
Business-facing criteria: every transaction attempted during the exercise either completes successfully via the fallback or ends up in a clearly tracked pending state, with zero duplicates and zero silently lost transactions, confirmed by reconciling the exercise's own transaction log against what the fallback provider and the primary system each show afterward. Operational criteria: the team followed the documented runbook to trigger the fallback rather than an ad hoc workaround, within whatever time target the plan specifies for detecting and switching over. Communication criteria, where the plan calls for customer-facing messaging during a real event of this kind: confirm the message that would have gone out is accurate to the actual state of the exercise's transactions, not a generic template nobody checked against what happened.
4. Rollback plan if the fallback itself misbehaves
Define the abort trigger ahead of time, not improvised mid-exercise: a concrete condition (error rate above a set threshold, a data-integrity check failing, an evidently wrong amount) that stops the exercise and reverts test traffic to the primary path, rather than letting a misbehaving fallback keep running and generating more to clean up. Treat anything the fallback processed before the abort exactly like a real incident: reconcile it, don't discard it, since finding out whether reconciliation actually works is as much the point of the exercise as testing the happy path. Have a named exercise controller with the authority to call the abort, separate from whoever is executing the steps, so the decision to stop isn't made by someone invested in proving the fallback works.
5. Close the loop
Run a hotwash immediately after (a short structured debrief on what worked, what didn't, and why) and log findings the way a real incident's post-event review would, with owners and target dates, so a gap the exercise found gets fixed and is retested at the next exercise rather than only known about.
Worked example
An organization tests its fallback SMS provider for one-time passcode delivery. Objective: if the primary provider is unavailable, can the fallback deliver codes within the plan's 30-second target with no duplicate sends. Design: a functional exercise, with primary-provider calls routed to a fault-injection endpoint returning timeouts for a fixed 20-minute window, run against 50 synthetic test accounts requesting codes continuously through that window, in a test environment with a provider sandbox that never sends real messages to real numbers. Success criteria set beforehand: at least 49 of the 50 attempts (allowing one edge case at the exact failover boundary) each receive exactly one code via the fallback within 30 seconds, and zero receive two.
Result: 50 of 50 delivered exactly once, averaging 22 seconds. But the exercise log shows the system took 40 seconds to detect the simulated outage and switch to the fallback in the first place, which isn't a delivery failure but eats into the 30-second budget for any request arriving in that detection window, and is logged as a finding. Nothing crossed the predefined abort trigger (an error rate above 5%), so the exercise ran to completion without needing rollback; the detection-time finding is written up with an owner and a target date for the next exercise to confirm it's fixed.
Trade-offs & pitfalls. Testing only the happy path, one transaction, no volume, is the most common shortcut, and it's exactly what misses the ordering and duplicate-handling problems that show up under concurrent load. Setting the success bar as "the fallback works at all" instead of "the fallback meets the actual business requirement" lets a technically-functioning fallback pass an exercise while still being unacceptable in a real event. And skipping the predefined abort trigger on the assumption an exercise "shouldn't need it" is how a test turns into a real incident; the rollback plan is itself part of what's being tested, not a formality around it.
Unlock Full Question Bank
Get access to all 12 Disaster Recovery and Business Continuity interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.