Backup and Disaster Recovery Questions
Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.
For an active-active distributed database (e.g., CockroachDB or Galera), explain strategies to capture consistent backups. Cover snapshot coordination, logical backups, point-in-time recovery options, and how to handle concurrent writes across nodes to ensure a usable restore.
Sample Answer
Direct answer
For an active-active cluster, "a consistent backup" cannot mean independently snapshotting each node on its own schedule, since writes land concurrently across nodes and uncoordinated per-node snapshots would capture different, mutually inconsistent moments, the same cross-node consistency problem as backing up unrelated volumes independently, just at database-node scale. The fix is to pick one single, cluster-wide logical point, an MVCC (multi-version concurrency control) timestamp for CockroachDB, or a specific binary-log position on a temporarily-desynced node for Galera, and reconcile every node's data to that one point, rather than trying to freeze the whole cluster's writes to get a clean moment.
Snapshot coordination
Use the cluster's own distributed consistency mechanism rather than coordinating snapshots externally. CockroachDB's native BACKUP command leverages the cluster's MVCC timestamp ordering (built on a hybrid logical clock, which combines wall-clock time with a counter so timestamps stay close to real time but are still guaranteed to always increase, even across nodes with imperfect clock sync) to take a backup as of a single, cluster-wide-consistent timestamp across every range and node, without needing to pause writes anywhere. Galera (synchronous multi-master replication for MySQL/MariaDB) has no equivalent built-in distributed-snapshot primitive as clean as CockroachDB's; the common approach is to use Galera's "desync" mode to temporarily pause one specific node from applying new replicated writes without blocking the rest of the cluster, take a physical or logical backup of that now-quiescent node, then let it resync and catch back up to the cluster.
Logical backups
A logical backup (schema plus row-level data) is portable across versions and storage engines and supports selective restore (a single table, for example), at the cost of being generally slower to produce and restore than a physical backup. For a large dataset, a logical dump only represents one true consistent instant if the tool captures it from a single, shared consistent read view; a naive per-table dump without that guarantee does not, and would reintroduce the same cross-table inconsistency problem the whole exercise is trying to avoid. Both CockroachDB's backup mechanism and a properly-configured logical dump against one fixed snapshot or transaction achieve this correctly.
Point-in-time recovery (PITR) options
CockroachDB supports incremental backups layered on a full backup, and because it retains MVCC history for a configurable garbage-collection window, it can restore to an arbitrary timestamp within that window, similar in spirit to WAL-based PITR. Galera, built on MySQL/MariaDB, can layer standard binary-log-based PITR on top of a physical backup taken from a desynced node: enable and retain binlog for at least the desired recovery window, then replay binlog events forward from the backup's known position to the target time, mechanically similar to WAL replay in PostgreSQL.
Handling concurrent writes across nodes for a usable restore
The property both approaches depend on is anchoring to a single logical point (an MVCC/HLC timestamp, or a specific binlog position) that every node's data can be reconciled against, rather than pausing the entire cluster to manufacture a clean moment, which would undermine the whole reason for running an active-active topology. After restore, explicitly verify cross-node and cross-shard consistency rather than assuming the backup tool handled it perfectly: a relationship spanning two ranges or shards could, in principle, be captured at slightly different micro-instants if the chosen consistency point was not truly atomic across every node, so this needs a targeted check (referential integrity across the specific tables or ranges most likely to span nodes), not a blanket assumption.
Worked example
Using illustrative timestamps, not measured values: node A commits a write to range R1 at HLC timestamp 100.002 (wall-clock component 100, logical counter 2); at nearly the same wall-clock instant, node B commits an unrelated write to range R2 at HLC timestamp 100.001. Neither node knows about the other's commit, but both timestamps are drawn from the same cluster-wide ordering, so they can still be compared. An operator runs BACKUP ... AS OF SYSTEM TIME targeting HLC timestamp 100.005: because that is later than both writes' timestamps, the backup includes the R1 write at 100.002 and the R2 write at 100.001, and excludes a later write node A had not yet committed at 100.010. Every range on every node resolves against that same 100.005 mark, so the backup reflects one logical instant across the whole cluster rather than each node's own most-recent-committed state. For Galera, the equivalent trace is simpler: node C is desynced at binlog position 4812003; a physical backup taken from node C in that state is tagged with that exact position, and restoring it then replaying binlog events forward from 4812003 reaches the same target point the rest of the cluster can be reconciled to.
Design a Kubernetes restore process that includes restoring etcd, persistent volumes for statefulsets, and re-creating cluster resources so an application can come back online in a known-good state. Include handling of secrets (KMS/rotations), storageclass differences across providers, and steps to bootstrap cluster services after data restore.
Sample Answer
Direct answer
Restoring a Kubernetes cluster is really two coordinated restores that have to be sequenced correctly: the cluster's control-plane state (etcd) has to come back first since nothing else can be reconciled without it, while the actual application data on persistent volumes is restored at the storage layer and then re-linked to the cluster objects that expect it, and secrets, storage classes, and bootstrap order each carry their own specific failure modes that a generic "restore everything" plan misses.
Structured elaboration
Restoring etcd. etcd is the key-value store holding the entire desired-state of the cluster (every Deployment, ConfigMap, Secret, and other object definition); the standard recovery path is restoring from an etcd snapshot onto fresh data directories for every etcd member, which always forms a new etcd cluster rather than rejoining the old one. All members must be restored from the exact same snapshot; mixing snapshots taken at different points across members produces an internally inconsistent cluster state. Once etcd is restored, point the API server(s) at it and confirm the API server comes up healthy before proceeding.
Persistent volumes for statefulsets. Restoring the volumes backing StatefulSet pods happens at the storage layer, independently of etcd, since it's usually a completely different system (a cloud block-storage service, an in-cluster storage system) with its own backup and restore mechanism. This independence is exactly what makes ordering tricky: if etcd is restored to a point that doesn't match what actually exists at the storage layer after volumes are restored, especially if restored volumes come back under new identifiers, Kubernetes' record of which PersistentVolume (PV, the actual storage resource) backs which PersistentVolumeClaim (PVC, a pod's request for storage that gets bound to a PV) can point at a volume ID that either doesn't exist anymore or, worse, now refers to different data than the cluster's restored state expects. The restore plan needs an explicit remapping step: after restoring volumes at the storage layer, update or recreate the PersistentVolume objects so their volume handles correctly point at the restored volumes, rather than assuming the old references still resolve correctly.
Re-creating cluster resources. After etcd and the API server are healthy, worker node kubelets (the per-node agent that runs and manages pods on that node) reconnect and re-register, and the scheduler and controllers begin reconciling the restored desired state against actual running pods. Some things aren't fully captured by an etcd snapshot alone and may need separate recreation alongside it: infrastructure-managed resources like external DNS records or cloud load balancer bindings that a service or ingress object depends on, which are often owned by infrastructure-as-code outside the cluster and need their own restore or reconciliation step run in parallel.
Secrets: KMS and rotation. Kubernetes Secrets are commonly encrypted at rest in etcd using an encryption-at-rest configuration backed by a key management service (KMS). The restored etcd snapshot contains ciphertext, so the exact key (or key version) that was active at backup time must still exist and be accessible at restore time, or those secrets are permanently unreadable regardless of how successful the rest of the restore was. This is a real, easy-to-miss gap: key rotation policy has to explicitly retain old key versions for at least the full backup retention window, or re-encrypt all secrets under a new key before an old one is destroyed, and it's worth periodically testing a restore against an intentionally older backup specifically to catch a rotated-and-destroyed key before it's discovered during a real incident.
StorageClass differences across providers. If the restore target is a different cluster, or the same cluster on a different cloud provider or storage backend than where the backup was taken, StorageClass names and their underlying provisioners often don't match; a PersistentVolumeClaim bound via a specific provider's StorageClass on the source may reference a provisioner that doesn't exist, or exists with different parameters, on the target. Restore tooling needs to remap StorageClass names to their equivalents on the target and verify performance-tier parity explicitly (a "fast" tier on the source silently landing on a slower equivalent on the target is not something a "restore succeeded" status check will catch).
Bootstrap steps. A reasonable sequence: restore etcd across all members from the same snapshot, start the API server(s) against the restored etcd and confirm cluster health, restore or remap the persistent volumes at the storage layer and update PersistentVolume objects to match, verify KMS access for encrypted secrets before relying on anything that depends on them, allow kubelets to reconnect, bring up cluster-critical system components first (networking/CNI, meaning the Container Network Interface, the plugin that provides pod networking, plus DNS, ingress controller) since application pods depend on those being functional, then allow StatefulSet pods to bind to the restored, remapped volumes, and finally run an application-level smoke test before declaring the cluster recovered rather than trusting Kubernetes-level "pod is Running" status alone.
Worked example
A team restores a cluster after a regional outage destroyed the original. They restore all 5 etcd members from the same hourly snapshot and bring the API server up first, confirming kubectl get nodes and core objects look correct before touching anything else. Storage-layer volume restore runs in parallel: block volumes come back under new provider-assigned IDs, so the team runs a remapping script that updates each PersistentVolume object's volume handle to the new ID before letting any StatefulSet pod schedule. KMS access is verified next: the backup was taken before a key rotation two weeks ago, and because the team's rotation policy explicitly retains prior key versions for 90 days, the old key version is still available and Secrets decrypt correctly, avoiding what would otherwise have been a silent, hard-to-diagnose failure only surfacing when an application tried to read a secret. CNI and CoreDNS come up first once kubelets reconnect; only after those report healthy does the team allow StatefulSet pods to schedule against the remapped volumes, followed by an application-level health check before declaring the restore complete.
Trade-offs and pitfalls
- Restoring etcd from a snapshot taken at a slightly different time than the storage-layer volume backups is a common, subtle failure: the cluster's record of what should exist and the actual data on restored volumes can silently disagree, which is why volume remapping needs to be an explicit, verified step, not an assumption.
- A restore drill that always uses the most recently rotated encryption key will never catch a retention gap in older key versions; testing restores against a deliberately older backup point is the only way to catch that specific failure mode before it happens for real.
- Bringing StatefulSet pods online before cluster-critical networking and DNS components are healthy can cause those pods to crash-loop or come up in a degraded state that looks like a data problem but is actually a sequencing problem, wasting time on the wrong root cause during an already-stressful recovery.
Design a backup and retention plan for a three-tier web application: stateless frontend web servers, application servers, a PostgreSQL primary with replicas, and S3-like object storage for user uploads. Business requirements: web/app RTO 2 hours, DB RPO 15 minutes, retention: 30 days for user uploads hot tier and 1 year for transactions. Describe backup types, cadence, storage tiers, and recovery order.
Sample Answer
Direct answer
Treat the two RTO/RPO numbers as two different engineering problems, not one: the 2-hour RTO for the stateless web/app tier is an infrastructure-provisioning problem solved with immutable images and IaC (no backup of the tier itself is needed, only of its definition), while the 15-minute RPO for PostgreSQL is a data-continuity problem that a nightly or even hourly backup cannot satisfy on its own and requires continuous WAL (write-ahead log) archiving or streaming replication. Recovery order follows the dependency graph: infrastructure and the database restore in parallel where possible, but the app tier cannot serve correct traffic until the database is verified, so the database restore is the critical path inside the 2-hour budget, not an independent 2-hour budget of its own.
Backup types by component
- Frontend and app servers (stateless): no data backup at all. "Backup" here means an immutable machine image or container image plus the Infrastructure-as-Code (IaC) that defines how many instances, what config, and what network wiring, stored in version control. Recovery is redeployment, not restore.
- PostgreSQL primary with replicas: a weekly or nightly full base backup, plus continuous WAL archiving (shipping WAL segments to durable storage as they are generated, not on a fixed interval) to support point-in-time recovery (PITR). Separately, keep at least one warm standby replica for fast failover, since PITR from a cold base backup is a different, slower recovery path than promoting an already-current replica.
- S3-like object storage for user uploads: object storage is typically already durable (multi-AZ replicated by the provider), so "backup" here means versioning plus cross-region replication to protect against accidental deletion, overwrite, or a regional failure, not protection against media failure.
Same-basis check on the RPO number
15 minutes RPO means: after any failure, at most 15 minutes of committed transactions may be unrecoverable. A nightly full backup alone gives an RPO of up to 24 hours, which does not meet this requirement by two orders of magnitude, so cadence has to be re-derived from the RPO target, not assumed. Continuous WAL archiving with a shipping interval well under 15 minutes (WAL segments are typically archived every few seconds to a couple of minutes under normal load) comfortably meets the target with margin; if only a fixed-interval WAL push is available rather than continuous streaming, set that interval to something like 5 minutes, not 15, because the RPO has to account for the interval plus shipping latency plus detection time, and setting the interval exactly at the SLA boundary leaves no margin for any of those.
Storage tiers
- User uploads, 30-day hot tier: keep the most recent 30 days in a standard, low-latency object storage class since that is the window most likely to be accessed or need a fast individual-object restore (accidental deletion, corruption). Interpreting the requirement as "uploads are retained in the hot tier for 30 days" (rather than "uploads are deleted after 30 days"), lifecycle older objects to a cheaper infrequent-access or archive tier rather than deleting them, unless the business has separately confirmed 30 days is the full retention period; this is stated as an assumption because the question does not specify what happens to uploads after day 30.
- Transactions, 1-year retention: recent PostgreSQL backups (say the last 1-2 weeks of base backups plus WAL) stay in fast storage for quick PITR; older backups within the 1-year window move to a cheaper archive tier where restore latency of hours is acceptable, because a 300-day-old backup is being kept for audit or compliance reasons, not for a 2-hour RTO scenario.
Recovery order and a worked RTO check
- Provision network and load balancer baseline via IaC. This has no data dependency and can start immediately, in parallel with step 2.
- Restore the PostgreSQL primary (the critical path): promote an existing warm standby if one survived the failure (fast, typically single-digit minutes), or restore the latest base backup and replay WAL to the most recent consistent point if the standby is also lost (slower, and must be timed).
- Verify database integrity (checksums, row counts, application-level sanity queries) before pointing anything at it.
- Deploy app servers from the pinned image/IaC pointed at the verified database; this is typically fast (minutes) since it is just infrastructure provisioning with no data to move.
- Deploy frontend web servers, same mechanism.
- Warm any caches, then cut traffic over.
Worked check on whether cold restore fits the 2-hour budget: suppose the database is 500 GB and cold restore-plus-replay throughput from the backup store is, as an illustrative ESTIMATE, 500 MB per minute. 500000/500=1000 minutes, about 16.7 hours, which massively breaches the 2-hour RTO even before the stateless tiers are provisioned. This is the concrete reason a cold base-backup restore should not be the primary recovery path for the app tier's 2-hour RTO: keep a warm standby that can be promoted in minutes as the primary recovery mechanism, and treat cold PITR restore as the fallback for a failure mode a standby cannot cover, such as logical corruption replicated to the standby, where the RTO target may need to be renegotiated with the business since a large cold restore genuinely cannot fit inside 2 hours at typical restore throughput.
Where the numbers do not directly compare
The RTO (2 hours, for web/app) and RPO (15 minutes, for the database) are not the same measurement and should not be added or averaged: RTO is "how long until service is back," RPO is "how much data could be lost." A design can meet both independently (standby promotion gives a low RTO; continuous WAL archiving gives a low RPO) but meeting one does not imply anything about the other, and a plan that only discusses one while citing both numbers has not actually answered the question.
Design a retention lifecycle for healthcare records that must meet HIPAA compliance with auditable controls and 7-year retention. Include archival tiering, deletion policies, legal hold handling, integrity checks, and cost-control measures while ensuring records can be retrieved when needed.
Sample Answer
Direct answer
Design the lifecycle around one governing rule and one override: records age from hot to warm to cold storage automatically on a schedule tied to the 7-year retention clock, and a legal hold can freeze that clock for an individual record at any point, overriding the schedule entirely, with every transition, access, and deletion logged so the whole lifecycle is independently auditable. It is worth noting as a framing point, not a correction to the question, that the specific "7 years" figure is typically driven by state medical-records law or an organization's own retention policy rather than by HIPAA itself, which principally sets minimum standards for security, privacy, and audit controls rather than a fixed record-retention duration; the design below treats 7 years as the given business requirement regardless of which regulation ultimately drives that number.
Archival tiering
- Hot (roughly years 0-2): actively-referenced records, ongoing patient care, fast retrieval expected, standard low-latency storage.
- Warm (roughly years 2-4): infrequent access, still fast enough for a routine records request, moved to a cheaper storage class via an automated lifecycle policy keyed off record creation or last-access date.
- Cold/archive (roughly years 5-7): rarely accessed, retained principally for compliance and legal-discovery readiness, cheapest storage tier, retrieval latency of hours is an acceptable and explicit trade-off, not an accident.
Encryption is maintained across every tier, not just the hot tier, since a "cheaper, rarely accessed" tier is not a "less protected" tier.
Deletion policies
Each record gets a computed expiry date at the 7-year mark from its creation or applicable retention trigger. Deletion is a two-phase process: mark for deletion, hold for a short verification window, then irreversible purge, with an audit log entry proving what was deleted and when. This matters in both directions: an organization must be able to prove it retained records for the required period, and equally must be able to prove it actually deleted records once the retention period expired, since over-retention (keeping protected health information longer than justified) is its own compliance and breach-exposure risk, not a conservatively safe default.
Legal hold handling
A legal_hold flag, independent of the computed expiry date, overrides normal expiration entirely: a record under hold cannot be deleted regardless of its age, applied when litigation, an audit, or a compliance investigation requires preserving specific records. Removing a hold is itself a controlled, logged, authorized action, ideally requiring approval from someone other than whoever applied it, so a hold cannot be silently lifted by the same party that has an interest in a record disappearing.
Integrity checks
Periodic checksum verification of archived records, since bit rot and media degradation are real risks over a 7-year dormancy window in cold storage, not a theoretical concern. HIPAA's Security Rule also expects audit trails of who accessed what, so access logging runs continuously across all tiers, not only the hot tier where access is more common. On the same cadence as a general restore-drill program, periodically sample-restore records from cold storage to confirm they are actually retrievable, not merely present and un-corrupted on paper.
Cost-control measures
Tiering itself is the primary lever: moving 5-year-old, rarely-accessed records to the cheapest appropriate tier rather than leaving everything in hot storage indefinitely out of caution. Compress archived records where the storage format allows it. Deduplicate where legally and technically permissible. Most importantly, retain to the true legal minimum rather than defaulting to "keep everything forever" as a safety instinct, since over-retention is not free: it expands the blast radius of any future breach to include data the organization had no ongoing obligation to hold.
Worked example
Patient record R-48213 is created on 2016-01-10, giving it a computed expiry of 2023-01-10. It moves through the normal tier schedule on that basis: hot through early 2018, warm through roughly 2020, cold from around 2021 onward, all driven purely by its age, independent of anything that happens later. In 2022-08-01, an unrelated compliance audit requires preserving the record, and a legal_hold flag is set on it. The record's computed expiry of 2023-01-10 arrives while the hold is still active: the scheduled deletion sweep checks the flag, finds it set, skips deletion, and logs the skip rather than purging on schedule. The hold stays in place for over a year past that expiry date. On 2024-11-15, the audit concludes and someone other than the person who originally set the hold releases it, logging the release as its own authorized action. With no hold protecting it and its expiry already well in the past, the next scheduled sweep marks R-48213 for deletion, holds it for the short verification window, then purges it, writing an audit log entry that records the original creation date, the computed expiry, the hold window, and the actual deletion date, so an auditor can see exactly why a record with a 2023 expiry was not actually removed until 2024.
Ensuring records remain retrievable when needed
Even cold-tier records need a bounded, documented retrieval SLA agreed with legal and compliance stakeholders (for example, any record retrievable within 24 to 48 hours), since audit responses and legal discovery requests carry their own deadlines that "eventually retrievable" does not satisfy. "Cold" should mean slower and cheaper, never "practically unreachable."
Auditable controls, tying it together
Every mechanism above (tier transition, access, hold applied or removed, deletion) needs to leave an audit trail satisfying the general audit-controls expectation under the HIPAA Security Rule (the exact regulatory citation is not asserted here with full confidence and should be confirmed against current legal guidance before this design is treated as a compliance sign-off), so the full lifecycle is demonstrably governed, not just technically implemented.
Design a quarterly restore-drill program for a global company. Specify drill frequency, scope (sample workloads versus full restore), success criteria, participants, metrics to collect, and which parts of the drill can or should be automated without risking production.
Sample Answer
Direct answer
A quarterly restore-drill program should have two speeds: a lightweight, largely automated check that runs far more often than quarterly (so backup corruption is caught in days, not months), and the human-facing quarterly drill itself, which validates the process, the runbook, and the people, not just the bytes. Conflating "quarterly" with "the only time we check anything" is the mistake to avoid; quarterly is the cadence for the expensive, cross-functional exercise, not for basic restorability.
Drill frequency
- Continuous/automated restorability checks: daily or weekly automated restore-and-verify of a rotating sample (see Automation below), catching silent backup corruption fast.
- Quarterly drill (as asked): a scheduled, semi-manual exercise that exercises the actual runbook and people, not just the restore mechanics.
- Annual full-scale exercise: once a year, a genuine end-to-end DR exercise including network/DNS cutover for the critical path, not just isolated volume restores, since a quarterly drill's smaller scope will not by itself validate the full cutover choreography.
Scope: sample versus full
Rotate coverage rather than trying to fully test everything every quarter, which is not affordable at any real scale. A reasonable split: Tier-0/critical systems get drilled every quarter (small enough set that full-restore is affordable); the broader estate is covered via a statistically meaningful random sample each quarter (for example 10-20% of remaining workloads), such that over a rolling year most systems have been exercised at least once, with Tier-1 systems targeted for coverage roughly every two quarters and Tier-2 systems at least annually.
Success criteria
- Restored within the system's stated RTO, measured end to end from declared drill start to verified functional, not to "data copy complete" (a restore that copies bytes quickly but takes another hour to pass validation has not actually met RTO, and measuring only the copy phase overstates performance).
- Data loss within the system's stated RPO, measured as the age of the restored data relative to the drill's simulated failure point.
- Zero unresolved data-integrity failures (checksum, referential integrity, or application-level acceptance tests, depending on system tier).
- The runbook as written was sufficient. Any manual workaround the operator had to invent on the spot is a runbook gap, and gets logged as a defect even if the drill otherwise "passed."
Participants
Quarterly drills for most of the estate can be lean: the backup/infrastructure engineers doing the technical restore, plus the owning application team validating data and functionality. The annual full-scale exercise should be cross-functional: on-call/incident-commander practicing the coordination role under drill conditions, security validating no reintroduced compromise, and business stakeholders observing so the exercise is credible evidence to leadership and auditors, not just an internal engineering checkbox.
Metrics to collect
- Actual restore time versus RTO target, per system, tracked over time (a restore time that is quietly creeping upward quarter over quarter is an early warning before it actually breaches the SLA).
- Actual data loss versus RPO target.
- Pass/fail rate of validation tests, trended over time, not just per-drill.
- Number of manual interventions required (each one is a candidate for either automation or a runbook fix).
- Mean time from "drill declared" to "system verified functional," which is the metric stakeholders actually care about.
What can and should be automated without risking production
Safe to automate: provisioning an isolated recovery environment via Infrastructure as Code (IaC), executing the restore itself, running checksum and smoke-test validation, and tearing the environment down afterward. Keep human-in-the-loop: the decision to actually cut real production traffic over, even in the annual exercise (simulate the cutover with a controlled, small blast radius rather than flipping real DNS for real customers without a rollback plan ready), and the final review of drill results, since a near-miss on timing or a borderline pass is exactly the kind of judgment call that should not be auto-approved.
Unlock Full Question Bank
Get access to all Backup and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.