Backup and Disaster Recovery Questions

Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.

MediumTechnical
65 practiced

For an active-active distributed database (e.g., CockroachDB or Galera), explain strategies to capture consistent backups. Cover snapshot coordination, logical backups, point-in-time recovery options, and how to handle concurrent writes across nodes to ensure a usable restore.

HardTechnical
59 practiced

Design a Kubernetes restore process that includes restoring etcd, persistent volumes for statefulsets, and re-creating cluster resources so an application can come back online in a known-good state. Include handling of secrets (KMS/rotations), storageclass differences across providers, and steps to bootstrap cluster services after data restore.

MediumSystem Design
78 practiced

Design a backup and retention plan for a three-tier web application: stateless frontend web servers, application servers, a PostgreSQL primary with replicas, and S3-like object storage for user uploads. Business requirements: web/app RTO 2 hours, DB RPO 15 minutes, retention: 30 days for user uploads hot tier and 1 year for transactions. Describe backup types, cadence, storage tiers, and recovery order.

HardTechnical
74 practiced

Design a retention lifecycle for healthcare records that must meet HIPAA compliance with auditable controls and 7-year retention. Include archival tiering, deletion policies, legal hold handling, integrity checks, and cost-control measures while ensuring records can be retrieved when needed.

MediumTechnical
76 practiced

Design a quarterly restore-drill program for a global company. Specify drill frequency, scope (sample workloads versus full restore), success criteria, participants, metrics to collect, and which parts of the drill can or should be automated without risking production.

Unlock Full Question Bank

Get access to all Backup and Disaster Recovery interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.