Backup and Disaster Recovery Questions
Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.
Define Recovery Time Objective (RTO) and Recovery Point Objective (RPO). Then give two practical examples mapping specific RTO and RPO values to backup types and frequencies: one for a small e-commerce transactional database, and one for a log or analytics pipeline where some data loss is acceptable.
Sample Answer
Direct answer
Recovery Time Objective (RTO) is the maximum acceptable length of time between a failure or data-loss event and the moment the affected system is back up and functioning normally: it answers "how long can we tolerate being down." Recovery Point Objective (RPO) is the maximum acceptable amount of data loss, measured as a span of time between the last valid recoverable state and the moment of failure: it answers "how much recent data can we afford to lose." Both are business decisions first (driven by what downtime or data loss actually costs the business) that then drive the technical backup design, not numbers an engineering team should invent unilaterally. Below are two worked examples that map specific RTO/RPO values to specific backup types and frequencies, one for a small, low-volume e-commerce transactional database, and one for a log or analytics pipeline where some data loss is tolerable.
Structured elaboration
Recovery Time Objective (RTO). Measured on the downtime axis. If a database goes down at 2:00pm and the RTO is 1 hour, the business has committed to being back online by 3:00pm at the latest. RTO is achieved through how fast you can restore or fail over, not through how often you back up.
Recovery Point Objective (RPO). Measured on the data-loss axis. If the last successful backup or log-shipment happened at 1:55pm and the database fails at 2:00pm, you've lost 5 minutes of writes: an RPO of 5 minutes means that gap is the maximum you've committed to tolerating. RPO is achieved through how frequently you capture recoverable state (backup cadence or continuous log shipping), not through restore speed.
Example 1: small e-commerce transactional database. Context: an online store taking live customer orders and payments continuously, at a modest volume, say roughly 200 orders/day on average (this is stated as a daily average for illustration; actual peak-hour order rate would be higher than this average implies, so treat it as a lower bound, not a worst case).
- RPO target: 5 minutes. Basis: 5 minutes is the maximum gap allowed between the last captured recoverable state and a failure. At the stated average order rate, a 5-minute gap risks under one order's worth of lost data on average (200 orders/day works out to about one order every 7 minutes on average, so a 5-minute window covers roughly 0.7 of that spacing), though an actual burst of orders arriving close together could lose more than the average implies, which the business judged acceptable to re-key or reconcile manually if it happens. Achieved via continuous transaction-log shipping (point-in-time recovery, PITR) capturing changes roughly every 5 minutes, layered on top of a daily full backup as the base to restore from.
- RTO target: 1 hour. Basis: 1 hour is the time from "database unreachable" to "store accepting orders again," achieved via an automated failover runbook: promote a standby replica in a separate availability zone, or restore the latest full backup and replay the shipped logs into a fresh instance, with connection-string or DNS cutover automated rather than done by hand under pressure.
Example 2: log or analytics pipeline (some data loss acceptable). Context: a pipeline ingesting application or clickstream logs into a warehouse used for internal dashboards and offline analysis, not customer-facing in real time.
- RPO target: 24 hours. Basis: 24 hours is the gap between the last completed nightly backup checkpoint and a failure. This is acceptable specifically because the upstream source (e.g. an event queue or a landing-zone object store) separately retains 7 days of raw events, so a lost day of processed warehouse data can be reprocessed from the untouched raw source rather than being permanently gone; the real cost of missing this RPO is a delay in reprocessing, not irreversible loss. Achieved via a nightly full or incremental backup of the warehouse tables; continuous log shipping isn't warranted here because the 24-hour RPO doesn't require it.
- RTO target: 8 business hours. Basis: this is measured in business hours within a single working day, not 8 calendar hours overnight, because dashboards are consulted during business hours and a pipeline down overnight causes no visible impact. Achieved via restoring the prior night's backup and re-running that day's batch ETL job once the source system is reachable again.
Why these numbers, and not others. Every value above is stated on a single, named basis: RPO is a time gap between capture events, RTO is a wall-clock duration to restored service, and the order-rate and retention figures used to justify them are explicitly labeled as averages or fixed-window retentions so they aren't mistaken for guarantees. The e-commerce example gets the tighter numbers on both axes because live revenue and payment records are on the line; the analytics example gets looser numbers on both because the underlying raw data is independently recoverable and the consumers are internal and delay-tolerant, not customer-facing.
Design a pattern for backing up and restoring petabytes of object storage data across regions, minimizing transfer costs while still meeting a 24-hour restore SLA for your most critical objects. How would you make a selective restore of just the objects you need fast, without restoring everything?
Sample Answer
Direct answer
The pattern that satisfies both constraints, transfer cost and a 24-hour restore SLA for the most critical objects, is tiering by criticality rather than replicating everything uniformly, backed by an independent, always-available metadata catalog that lets you resolve exactly which objects an incident requires without scanning the whole store. Cost scales with what you actually replicate and at what storage tier, not with the size of the total archive, which is the lever that makes petabyte-scale DR affordable at all.
Structured elaboration
Minimizing transfer cost. Most object stores at this scale have a long tail: a small fraction of objects account for most access, and a large share of the rest is genuinely cold. Replicate incrementally at the object level, only new or changed objects (detectable via version ID, ETag, a hash-based fingerprint of the object's contents that changes whenever the object does, or last-modified timestamp, avoiding re-transferring unchanged data), and place the replicated copies of cold objects in a cheaper archive storage class in the DR region rather than a uniformly warm one. That lowers ongoing storage cost even though it raises the retrieval latency for those specific objects, which is acceptable because the 24-hour SLA is explicitly scoped to the most critical objects, not the whole archive.
Worked cost example (basis: the size of the replicated critical subset, not the total store). Suppose the critical subset is 10 TB out of a 5 PB total store. Replicating just that 10 TB continuously costs egress on 10 TB, not on 5,000 TB: at an illustrative $0.02/GB egress rate, that's 10,000 GB x $0.02 = about $200 one-time (plus small ongoing costs for deltas), versus replicating the full 5 PB at the same rate, 5,000,000 GB x $0.02 = about $100,000, roughly a 500x difference. This is exactly why uniform replication doesn't scale at petabyte size and tiering by criticality does.
Meeting the 24-hour restore SLA for critical objects specifically. Keep the critical subset in a warm storage class in the DR region, immediately readable with no thaw or retrieval delay, so its restore time is bounded by network transfer and API call time, not by an archive-retrieval wait. This matters because some archive storage classes have retrieval delays ranging from minutes to tens of hours depending on the tier chosen; placing critical objects in a deep-archive tier by mistake could consume the entire 24-hour SLA before any data even starts moving, independent of network speed. The storage class chosen for the critical subset is itself an SLA-determining decision, separate from and at least as important as network throughput.
Selective fast restore without restoring everything. This requires a metadata catalog, an index of object key, version, checksum, storage tier and location, backup timestamp, and criticality classification, maintained independently as a small, always-hot database, not embedded inside the bulk archive itself (the same principle as keeping a backup catalog out of the data center it describes: if the catalog only lives inside the store that just had an incident, you can't find anything to restore). The catalog resolves "which objects does this incident actually require" to a precise list quickly, and that list drives targeted, parallelized restore requests for exactly those objects, since object stores support high per-object parallelism, rather than the much slower default of restoring a whole time-ordered prefix and scanning through it to find what's needed.
Putting it together. The pattern is: replicate incrementally and selectively rather than uniformly; place the critical subset in a warm, immediately-readable tier and everything else in cheaper archive tiers; maintain an independent, fast metadata catalog that turns "restore what we need" into a precise, parallel operation instead of a full-archive scan. Validate the design by periodically restoring a sample from the critical tier and timing it against the 24-hour SLA, and separately confirming the catalog itself stays available and accurate even during a regional incident.
Design a disaster recovery plan for a stateful Postgres deployment running in the cloud (RDS or self-managed on EBS). Include target RPO and RTO, backup cadence and retention, cross-region replication options, failover procedures, validation and automated DR testing, and how you'd restore production traffic in a controlled way.
Sample Answer
Direct answer
The design splits into two layers that solve different problems: continuous replication (streaming or WAL shipping, WAL being the write-ahead log, an append-only record of every change Postgres makes before applying it) gets you a warm standby you can promote quickly, meeting RTO; and independent base backups plus archived WAL get you point-in-time recovery (PITR) and protection against the kind of failure replication doesn't cover, like a bad DROP TABLE that a synchronous or asynchronous standby would faithfully replicate right along with the primary. For illustration I'll target RPO (Recovery Point Objective, the maximum acceptable data loss measured in time) of roughly 30 seconds and RTO (Recovery Time Objective, the maximum acceptable downtime) of roughly 15 minutes, both stated as example targets since the question leaves them to the candidate, not as universal numbers.
Structured elaboration
RDS vs self-managed on EBS. On RDS (a managed Postgres service), automated backups continuously stream WAL to object storage, giving PITR out of the box with retention up to 35 days, and a cross-region read replica (asynchronous by default) can be promoted for regional DR. On self-managed Postgres on EBS (Elastic Block Store, cloud block storage volumes), the same capability has to be built: pg_basebackup for periodic base backups, continuous WAL archiving to object storage (commonly via wal-g or pgBackRest) for PITR, and physical streaming replication to a standby in another region. EBS snapshots can also back the base-backup layer, but need care: a snapshot taken mid-write without coordinating with Postgres (e.g. via pg_start_backup) can capture an inconsistent volume state, so either use a backup tool that handles this coordination or take the base backup through pg_basebackup instead of a raw disk snapshot.
Cross-region replication: sync vs async. Synchronous replication (the primary waits for the standby to acknowledge before committing) gives an RPO close to zero but adds the cross-region round-trip time to every write's latency, commonly tens to well over a hundred milliseconds depending on the region pair, and risks stalling the primary if the standby becomes unreachable. Asynchronous replication accepts a small replication lag (commonly seconds under normal load) in exchange for no write-latency tax; given a 30-second RPO target, async has ample headroom and is the more cost-effective default. Reserve synchronous replication for a stricter RPO tier if one is ever needed.
Failover procedure. Promote the standby (pg_ctl promote, or the managed failover API on RDS), update the client-facing endpoint (DNS or a connection proxy) to point at the new primary, and fence the old primary, revoking its ability to accept writes, so it can't diverge from the newly promoted primary if it later comes back online (a split-brain scenario, where two nodes both believe they're primary and accept conflicting writes).
Worked RTO breakdown (basis: wall-clock minutes from declared incident to first successful write, summed and checked against the 15-minute target, not each phase in isolation). Detection and decision to fail over: about 5 minutes (health-check failure threshold plus confirmation). Promotion of the standby: about 2 to 5 minutes. Endpoint cutover and client reconnect: about 5 minutes (DNS propagation and application retry/reconnect behavior). Smoke validation before declaring the incident resolved: about 3 to 5 minutes. Summed, that's roughly 15 to 20 minutes; if the total lands over the 15-minute target, the endpoint-cutover step (usually the least automatable) is the first place to look for savings, e.g. via a connection proxy that can be repointed faster than DNS TTL propagation allows.
Validation and automated DR testing. Run scheduled drills that promote the DR replica into an isolated network (not the live traffic path), run automated read/write smoke tests and a row-count or checksum comparison against the primary before the drill promotes, and time every phase against the RTO/RPO targets. Separately, test the PITR path itself (restore a base backup plus archived WAL to a specific past timestamp) on its own schedule, since it's a different recovery mechanism from replica promotion and can silently break (a missing WAL segment, an expired retention window) without the replication drills ever exercising it.
Controlled restore of production traffic. After promoting the DR replica and validating it, don't cut 100% of traffic over immediately: shift a small percentage first via weighted routing, watch error rates and replication lag on the newly promoted primary, then ramp to full traffic. Keep the old primary around, fenced but not destroyed, as a forensic copy in case anything needs to be recovered from it, and rebuild a fresh standby in a healthy region afterward so the system isn't left without DR coverage once the immediate incident is over.
Design an architecture for continuous log-based recovery for a high-throughput transactional database producing millions of writes per minute and requiring near-zero data loss (RPO on the order of seconds). Describe components for log capture (CDC/WAL), durable transport, storage, indexing for quick restore, and techniques to minimize impact on primary performance.
Sample Answer
Direct answer
A near-zero recovery point objective (RPO measured in seconds) for a database taking millions of writes per minute requires capturing the database's own transaction log as it's produced rather than querying it, shipping that log continuously to a separately durable store, indexing it for fast targeted replay, and doing all of this in a way that adds essentially no extra load to the primary, since a design that protects data at the cost of primary write performance has traded one production risk for another.
Structured elaboration
Log capture (CDC/WAL). Capture changes by tapping the database's native transaction log (the write-ahead log, WAL, for many relational systems, or a binary log for others) or by using change data capture (CDC), a mechanism that reads the log stream and emits a structured feed of row-level changes, rather than by periodically querying tables for what changed. This matters because the database writes its log as an integral part of committing a transaction anyway, so reading that log stream is asynchronous to the transaction path and doesn't add read load to the primary the way repeated table scans or polling queries would.
Durable transport. Continuously ship captured log records off the primary, as they're generated, to a separate, durable transport layer (a distributed log or queue system built for exactly this kind of ordered, durable streaming), so "captured" and "durably stored somewhere else" happen close together in time. This decoupling is what protects against a primary crash immediately after a commit: if the log record already made it into the durable transport, it survives the primary's loss even though the primary's own local copy is gone. The transport layer needs to preserve commit order within each shard or partition and support backpressure, so that if a downstream consumer slows down, the transport absorbs the backlog rather than that slowdown propagating back and blocking the primary's own commit path.
Storage. The shipped log stream lands in an append-only, durable store, organized by time or by log sequence number so replay can start from any point in the retained history. To avoid ever needing to replay from the very beginning of recorded history, the system periodically materializes full checkpoints, complete snapshots of state at a given point, so a restore only needs to replay from the nearest checkpoint forward to the target point, not from day one.
Indexing for quick restore. Index checkpoints and log segments by timestamp and log sequence number, so a request to "recover to time T" can jump directly to the right checkpoint and the right range of log segments instead of scanning through everything retained. This indexing is what keeps restore time bounded and predictable even as the total volume of retained log history grows over months, rather than restore time slowly degrading as history accumulates.
Minimizing impact on primary performance. Read the log through the mechanism the database already provides for this purpose (its native streaming replication or log-shipping protocol) rather than a heavier approach like frequent full-table dumps or repeated large queries. Where the database supports it, capture from a replica's log stream rather than directly from the primary, keeping essentially all of the extra read load off the primary entirely. And apply backpressure and rate-limiting on the capture path itself, so that if the downstream transport or storage layer slows down, that slowdown is absorbed by buffering in the capture and transport layers rather than ever blocking or throttling the primary's own commit path; a near-zero RPO commitment must never come at the cost of primary write availability, since that would trade a data-loss risk for an equally serious availability risk.
Worked example
A payments database processes several million writes per minute and commits an RPO target on the order of seconds. Log capture reads directly from the database's native WAL streaming protocol, connected to a replica rather than the primary, keeping all of the capture load off the primary entirely. Captured log records stream continuously into a durable, ordered transport layer with steady-state lag typically sitting in the low hundreds of milliseconds, well inside the seconds-scale RPO target, leaving real headroom for brief spikes rather than running right at the edge of the commitment. A full checkpoint materializes every 15 minutes; log segments and checkpoints are indexed by log sequence number and wall-clock time. When a targeted restore is later requested for a specific timestamp, the index resolves directly to the nearest preceding checkpoint and the small range of log segments between that checkpoint and the target, avoiding a replay of the full multi-month retained history, and completes in a small, bounded amount of time regardless of how much total log history the system has accumulated.
Trade-offs and pitfalls
- Capturing changes by polling tables instead of reading the native transaction log adds real, avoidable read load directly to the primary and typically can't achieve seconds-scale freshness anyway, since polling has to run frequently enough to approximate continuous capture, working against the very goal it's meant to serve.
- Setting the replication or shipping lag target exactly equal to the RPO commitment, rather than well below it, leaves no margin for a transient spike, which is exactly when a real failure is also more likely to be happening.
- Skipping periodic checkpoints to save storage means every restore, even a recent one, has to replay from further back in history, directly working against the goal of keeping restore time bounded and fast as retained log volume grows.
Design a retention lifecycle for healthcare records that must meet HIPAA compliance with auditable controls and 7-year retention. Include archival tiering, deletion policies, legal hold handling, integrity checks, and cost-control measures while ensuring records can be retrieved when needed.
Sample Answer
Direct answer
Design the lifecycle around one governing rule and one override: records age from hot to warm to cold storage automatically on a schedule tied to the 7-year retention clock, and a legal hold can freeze that clock for an individual record at any point, overriding the schedule entirely, with every transition, access, and deletion logged so the whole lifecycle is independently auditable. It is worth noting as a framing point, not a correction to the question, that the specific "7 years" figure is typically driven by state medical-records law or an organization's own retention policy rather than by HIPAA itself, which principally sets minimum standards for security, privacy, and audit controls rather than a fixed record-retention duration; the design below treats 7 years as the given business requirement regardless of which regulation ultimately drives that number.
Archival tiering
- Hot (roughly years 0-2): actively-referenced records, ongoing patient care, fast retrieval expected, standard low-latency storage.
- Warm (roughly years 2-4): infrequent access, still fast enough for a routine records request, moved to a cheaper storage class via an automated lifecycle policy keyed off record creation or last-access date.
- Cold/archive (roughly years 5-7): rarely accessed, retained principally for compliance and legal-discovery readiness, cheapest storage tier, retrieval latency of hours is an acceptable and explicit trade-off, not an accident.
Encryption is maintained across every tier, not just the hot tier, since a "cheaper, rarely accessed" tier is not a "less protected" tier.
Deletion policies
Each record gets a computed expiry date at the 7-year mark from its creation or applicable retention trigger. Deletion is a two-phase process: mark for deletion, hold for a short verification window, then irreversible purge, with an audit log entry proving what was deleted and when. This matters in both directions: an organization must be able to prove it retained records for the required period, and equally must be able to prove it actually deleted records once the retention period expired, since over-retention (keeping protected health information longer than justified) is its own compliance and breach-exposure risk, not a conservatively safe default.
Legal hold handling
A legal_hold flag, independent of the computed expiry date, overrides normal expiration entirely: a record under hold cannot be deleted regardless of its age, applied when litigation, an audit, or a compliance investigation requires preserving specific records. Removing a hold is itself a controlled, logged, authorized action, ideally requiring approval from someone other than whoever applied it, so a hold cannot be silently lifted by the same party that has an interest in a record disappearing.
Integrity checks
Periodic checksum verification of archived records, since bit rot and media degradation are real risks over a 7-year dormancy window in cold storage, not a theoretical concern. HIPAA's Security Rule also expects audit trails of who accessed what, so access logging runs continuously across all tiers, not only the hot tier where access is more common. On the same cadence as a general restore-drill program, periodically sample-restore records from cold storage to confirm they are actually retrievable, not merely present and un-corrupted on paper.
Cost-control measures
Tiering itself is the primary lever: moving 5-year-old, rarely-accessed records to the cheapest appropriate tier rather than leaving everything in hot storage indefinitely out of caution. Compress archived records where the storage format allows it. Deduplicate where legally and technically permissible. Most importantly, retain to the true legal minimum rather than defaulting to "keep everything forever" as a safety instinct, since over-retention is not free: it expands the blast radius of any future breach to include data the organization had no ongoing obligation to hold.
Worked example
Patient record R-48213 is created on 2016-01-10, giving it a computed expiry of 2023-01-10. It moves through the normal tier schedule on that basis: hot through early 2018, warm through roughly 2020, cold from around 2021 onward, all driven purely by its age, independent of anything that happens later. In 2022-08-01, an unrelated compliance audit requires preserving the record, and a legal_hold flag is set on it. The record's computed expiry of 2023-01-10 arrives while the hold is still active: the scheduled deletion sweep checks the flag, finds it set, skips deletion, and logs the skip rather than purging on schedule. The hold stays in place for over a year past that expiry date. On 2024-11-15, the audit concludes and someone other than the person who originally set the hold releases it, logging the release as its own authorized action. With no hold protecting it and its expiry already well in the past, the next scheduled sweep marks R-48213 for deletion, holds it for the short verification window, then purges it, writing an audit log entry that records the original creation date, the computed expiry, the hold window, and the actual deletion date, so an auditor can see exactly why a record with a 2023 expiry was not actually removed until 2024.
Ensuring records remain retrievable when needed
Even cold-tier records need a bounded, documented retrieval SLA agreed with legal and compliance stakeholders (for example, any record retrievable within 24 to 48 hours), since audit responses and legal discovery requests carry their own deadlines that "eventually retrievable" does not satisfy. "Cold" should mean slower and cheaper, never "practically unreachable."
Auditable controls, tying it together
Every mechanism above (tier transition, access, hold applied or removed, deletion) needs to leave an audit trail satisfying the general audit-controls expectation under the HIPAA Security Rule (the exact regulatory citation is not asserted here with full confidence and should be confirmed against current legal guidance before this design is treated as a compliance sign-off), so the full lifecycle is demonstrably governed, not just technically implemented.
Unlock Full Question Bank
Get access to all Backup and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.