Backup and Disaster Recovery Questions
Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.
Design a monitoring and alerting plan for backup systems: what would tell you a backup silently failed, and what would you put in front of an on-call engineer versus what a manager needs to see on a dashboard?
Sample Answer
Direct answer
A silent backup failure is one where the job scheduler reports success, or reports nothing at all,
while the actual outcome is wrong, so detecting it needs signals that are independent of the job's
own self-reported status. What goes in front of an on-call engineer and what goes on a manager's
dashboard should differ, because the two audiences need to do different things with what they see:
one needs to act right now, the other needs to judge overall posture. The right level of tooling to
build any of this with scales with team size; the underlying questions do not.
Structured elaboration
Signals that reveal a silent failure, from basic to more advanced.
A small team with limited tooling can start with a short, cheap checklist and it is a legitimate,
cost-appropriate answer, not a lesser version of a larger build:
- Missing-job detection. Alert if a job that is supposed to run on a schedule simply did not
run, or reported nothing at all, by its expected completion time, sometimes called a dead man's
switch. This catches failures upstream of the job even starting, a down scheduler, a
disconnected agent, which a "job failed" alert alone would never catch, since the job never got
the chance to fail. - Size or duration anomaly. Alert if a backup's output size or duration is drastically
different from its recent normal range. A backup that usually writes 50 GB and one night writes
500 MB most likely did not actually capture the data it thinks it did, even though the job
reports success. - Basic capacity threshold. Alert on backup storage utilization crossing a defined threshold,
the same check a small team might otherwise only do manually on a weekly review, turned into an
automated alert.
A team with more monitoring maturity adds depth on top of the same three ideas:
- Checksum and catalog completeness results feeding an alert, rather than only the job
scheduler's own status. - Automated restore or synthetic verification results treated as a first-class monitored
signal, with its own alert if a check fails or has not run recently. - Trend-based alerting, not just threshold-based, catching statistically meaningful drift, for
example backup duration creeping upward over weeks, which often precedes an outright failure,
rather than waiting for a hard threshold to be crossed.
What goes in front of an on-call engineer. Actionable, specific, real-time information: which
job or system failed or looks anomalous, what kind of failure it is (missed schedule, suspicious
size, verification failure), a direct link or runbook to the remediation steps, and enough context
to triage severity, a single low-priority file share versus a production database, without having
to dig for it. Filter aggressively: alert fatigue is a real risk, so on-call should see things that
need action now, not every warning-level signal.
What a manager needs to see on a dashboard. Aggregate, trend-level information, not per-job
noise: overall backup success rate over a time window, stated clearly as a percentage of scheduled
jobs, not of all systems, since those are different denominators and easy to conflate; restore-test
pass rate, arguably more important than raw job success rate, since a high job-success rate alone
does not prove data is actually recoverable; open or unresolved incidents and their age; and
capacity or cost trending, whether a budget increase is coming. The manager view answers "should I
be worried about our overall recoverability posture." The on-call view answers "what do I need to
fix right now."
How this scales with team size
For a small shop, the entire monitoring system might legitimately be a short checklist, someone
glancing at the basic-tier signals above on a regular cadence, plus a simple weekly rollup for
whoever is the de facto backup owner, rather than a purpose-built platform, and that is a
reasonable choice at that scale. For a larger team with dedicated tooling, the same underlying
questions, detecting silent failure, giving on-call and management different views, get answered
with dedicated monitoring, automated verification pipelines, and separate dashboards. The
principles do not change between the two; only the tooling maturity applied to them does.
Explain the differences between full, incremental, and differential backups. For each type, describe how the restore operation actually works, typical storage and I/O characteristics, how complex the recovery chain gets, and a realistic scenario where you'd prefer that type over the others.
Sample Answer
Direct answer
The single decision that determines everything else about a backup type is what it actually captures: a full backup captures every byte, an incremental backup captures only what changed since the last backup of any kind, and a differential backup captures everything that changed since the last full backup (ignoring any differentials in between). That one difference cascades into how restore works, how much storage and I/O each consumes, how long and fragile the recovery chain gets, and which scenario each fits. Point-in-time recovery via transaction-log shipping is a related but separate mechanism and isn't one of the three types asked about here.
Structured elaboration
Full backups.
- What it captures: a complete copy of all data at the moment the backup runs.
- Restore mechanics: apply a single file. No assembly, no ordering to get right.
- Storage and I/O: highest of the three, every run duplicates the entire dataset; backup-time I/O is also the highest since the whole dataset is read and written every time.
- Recovery chain complexity: none, it's a chain of length one.
- Realistic scenario: as the periodic anchor point underneath a mixed strategy (a weekly full backing incremental or differential backups on the days in between), or as a one-off just before a risky operation like a schema migration, where you want the simplest possible restore path available if it goes wrong, not a clever one.
Incremental backups.
- What it captures: only the data changed since the immediately preceding backup, whatever type that was (the last full, or the last incremental).
- Restore mechanics: restore the last full, then apply every incremental since it, strictly in order. If backups were taken Mon (full), Tue, Wed, Thu, Fri, Sat, Sun (incrementals), restoring to Sunday means applying all six incrementals on top of the full in sequence.
- Storage and I/O: lowest per individual backup run, since only changed data is captured each time, so this is the cheapest option in steady-state storage and the fastest to actually take.
- Recovery chain complexity: the highest of the three, and it grows every day since the last full. A single corrupted or missing incremental in the middle of that chain breaks the restore for every day after it, not just that one day.
- Realistic scenario: very large datasets with a tight backup window and a real storage-cost driver, e.g. a multi-terabyte warehouse where a nightly full backup wouldn't finish before the maintenance window closes and daily full copies would be prohibitively expensive to store, and where actual restores are rare enough that a longer, multi-file restore assembly is an acceptable trade for the ongoing savings.
Differential backups.
- What it captures: everything changed since the last full backup (not since the last differential), so each day's differential is a superset of the previous day's.
- Restore mechanics: restore the last full, then apply only the single most recent differential. Restoring to Sunday in the same Mon-Sun example above means applying the full plus only Sunday's differential, the Tue through Sat differentals are never needed.
- Storage and I/O: between full and incremental, and it grows every day since the last full as more cumulative change gets folded into each new differential; by the day before the next full, the differential can approach the size of a full backup of just the changed rows.
- Recovery chain complexity: fixed at two, full plus latest differential, regardless of how many days have passed since the last full. Losing an older differential doesn't matter because only the newest one is ever used.
- Realistic scenario: systems where restore speed and restore reliability matter more than minimizing backup-time storage, e.g. a transactional system with a demanding recovery-time target where you don't want six independent incremental files all needing to be intact to get back online; you're trading a bit more storage for a restore path that can't be broken by losing an old file.
Worked example, basis of every number stated. Assume a 500 GB database with a constant 10 GB/day of changed data (this is the daily delta measured in changed bytes, treated as constant for illustration; a real system's WAL (write-ahead log, the database's own transaction log) volume can exceed the net-changed-bytes figure because of engine-level effects like full-page writes (the database re-logging an entire disk page rather than just the changed bytes, the first time that page changes after a checkpoint), so treat 10 GB/day as the backup-relevant change volume, not raw log volume). A full backup runs Monday: 500 GB. Over the following six days:
- Incremental: each day backs up only that day's 10 GB, so Tue through Sun each write a 10 GB file, roughly 60 GB total across the week. Restoring to Sunday means applying the 500 GB full plus six 10 GB incrementals in order: seven files, seven steps.
- Differential: each day backs up everything since Monday's full, so the differentials are 10 GB (Tue), 20 GB (Wed), 30 GB (Thu), 40 GB (Fri), 50 GB (Sat), and 60 GB (Sun), since each day rolls in one more day's worth of change on top of the same base. Restoring to Sunday means applying the 500 GB full plus only the 60 GB Sunday differential: two files, two steps, even though the Sunday differential file itself is as large as the entire incremental week combined.
That last line is the trade-off in one sentence: incremental keeps each individual backup small at the cost of a restore chain that gets longer and more fragile every day; differential keeps the restore chain fixed at two files at the cost of each individual backup getting larger every day since the last full.
Explain the difference between filesystem or volume snapshots and traditional backups. Discuss scenarios where snapshots are sufficient (e.g., rapid rollback) and cases where snapshots are not a replacement for backups (e.g., offsite, long-term retention, or provider failures).
Sample Answer
Direct answer
A snapshot is a point-in-time reference inside the same storage system it was taken from, typically
recording only what has changed since the snapshot point and depending on the original data still
being present as its base. A backup is an independent, complete copy stored in a separate location
or system, with no ongoing dependency on the original surviving. That single distinction, where the
data physically lives relative to the source, decides everything about what each one can and cannot
survive.
Structured elaboration
When snapshots are sufficient. Rapid rollback is the clearest case: undoing a bad deployment,
a failed patch, or an accidental change within minutes, where near-instant reversal matters more
than surviving a total loss of the storage system itself. Snapshots are typically much faster to
create and restore from than a full backup, since they do not move or copy the full dataset, only
the changes since the reference point, which is also what makes them cheap enough to take very
often, hourly or more, as a frequent safety net between less frequent full backups.
Offsite protection. A snapshot that lives on the same storage array or system as the original
data provides no protection if that array, its site, or its data center is lost, damaged, or
compromised, a fire, a flood, or ransomware that reaches the array's own management plane. A backup
stored somewhere genuinely independent, a different system, a different location, ideally a
different administrative boundary, protects against exactly this class of failure, which a snapshot
structurally cannot, because it never leaves the source system in the first place.
Long-term retention. Because most snapshot implementations rely on tracking changes relative to
the live volume, retaining very old snapshots for years is usually neither practical nor storage
efficient: the underlying change-tracking metadata and retained delta chains grow large and
unwieldy over long periods, and many snapshot systems have practical limits on how many, or how old,
snapshots they can efficiently keep. Backups, especially on a cheap archive tier, are the
appropriate mechanism for retention measured in months to years rather than days.
Provider or system failures. If the failure is the storage platform itself, a hardware fault, a
bug in the storage controller, or, in a cloud context, a regional outage or account-level issue
with the provider hosting the snapshots, then any snapshot stored on or by that same system is
unavailable right along with everything else. Only a genuinely independent backup copy, a different
system and ideally a different provider, account, or region boundary, survives a failure of the
platform that hosted both the original data and its snapshots.
Why real disaster-recovery strategy uses both
Snapshots handle fast, frequent, cheap local recovery, the "undo that" case. Backups handle genuine
disaster recovery, offsite protection, and long-term retention, the "the building is gone" or "we
need this from three years ago" case. Relying on snapshots alone is a common and genuinely
dangerous shortcut precisely because snapshots look like backups, both let you go back to an
earlier point, while having a structurally different failure-survival profile: a snapshot survives
everything except losing the system it lives on, and losing that system is exactly the scenario a
backup exists for.
A small company needs a backup retention policy: daily backups kept 30 days, weekly backups kept 6 months, yearly backups kept 7 years. Walk through how you'd actually structure and enforce a policy like this so it holds up over time, including what should happen when a legal request requires keeping something past its normal deletion date.
Sample Answer
Direct answer
A tiered retention policy like this (daily kept 30 days, weekly kept 6 months, yearly kept 7 years) only holds up over time if each tier is its own independently-scheduled backup job with its expiry computed and enforced automatically, deletion is checked against a legal-hold registry before it ever runs, and a hold, when placed, overrides the normal schedule entirely until it's explicitly released.
Structured elaboration
Structuring the tiers. Run daily, weekly, and yearly backups as three separate, independently-scheduled jobs, each tagging its own backups with its tier at creation time, rather than trying to derive "the weekly backup" or "the yearly backup" after the fact by picking out a specific daily run (for example, "whichever daily happens to land on Sunday"). Deriving tiers after the fact is fragile: if the daily job that was supposed to double as that week's designated weekly backup happens to fail on exactly that day, you silently lose that tier's retention point even though other daily backups ran fine that week. Each tier's backup carries its own expiry date, computed at creation from the policy (30 days from creation for a daily, 6 months for a weekly, 7 years for a yearly).
Enforcing it. A scheduled retention job runs regularly, finds backups whose computed expiry has passed, and deletes them, but only after checking each one against a legal-hold registry first; a held backup is skipped regardless of how far past its normal expiry it is. Where the storage platform supports it, backing this with an immutability mechanism (a retention lock the backup system itself can't casually override) makes the policy something the system enforces structurally, not just something a script is trusted to get right every time.
What should happen for a legal hold. When a legal request requires preserving something past its normal deletion date, that request creates an entry in a hold registry, keyed to the specific backups or to a broader query (a date range, a specific system) that the request covers. This hold is logged: who placed it, what matter it relates to, and when it's expected to be reviewed. From that point, the retention job's deletion step treats a hold as a hard veto that overrides the normal expiry clock entirely, a held backup is not deleted at 30 days, 6 months, or 7 years even though its normal schedule would otherwise expire it. When the hold is explicitly released by whoever has authority to release it, the backup re-enters the normal expiry queue; if its normal expiry date already passed while it was held, the sensible default is to delete it promptly after release rather than let it linger indefinitely just because it happened to be held for a while.
What happens over time. As backups age through their tiers, the total retained footprint stays bounded rather than growing without limit, because older daily and weekly backups expire and roll off while the yearly tier captures a much sparser set of long-term points; the daily tier is never trying to hold seven years of daily granularity, only the yearly tier holds that much time depth, at a much coarser interval. A periodic audit report, showing exactly what was retained, what was deleted and when, and what's currently under hold, gives the organization something concrete to show if a regulator or a court later asks whether the policy was actually followed, rather than just describing what the policy says on paper.
Worked example
A company generates roughly 500 GB of backup data per full backup point (using this as a simplifying assumption for the calculation below; real deduplication would lower this in practice, and that's a real limitation of the estimate worth stating rather than treating it as exact). Using GB throughout for a consistent basis: the daily tier at 30 days of retention holds 30 x 500 = 15,000 GB. The weekly tier at roughly 6 months (about 26 weekly backups) holds 26 x 500 = 13,000 GB. The yearly tier at 7 years holds 7 x 500 = 3,500 GB. The steady-state total across all three tiers is 15,000 + 13,000 + 3,500 = 31,500 GB, about 31.5 TB. Compare that to what naively keeping every single daily backup for the full 7 years would require: 7 years x 365 days x 500 GB = 1,277,500 GB, about 1,277.5 TB, roughly 1.28 PB. The tiered structure holds onto only about 31,500 / 1,277,500, roughly 2.5%, of that naive figure, a reduction of about 97.5%, which is the concrete payoff of rolling old daily granularity off in favor of sparser weekly and yearly points rather than keeping everything at full daily resolution forever. Midway through year 3, a litigation hold arrives covering a specific customer's data for a defined date range. The hold registry gets an entry for every backup (across all three tiers) whose contents include that customer's data in that range; the next scheduled retention run finds several of those backups already past their normal daily or weekly expiry and, seeing the active hold, skips deleting them, logging each skip against the hold's ID.
Trade-offs and pitfalls
- Deriving a tier's backup from an existing daily run after the fact (rather than scheduling it as its own independent job) is a common, subtle way to silently lose a retention point on exactly the day it mattered; each tier should be its own job.
- Holding backups indefinitely by default once a hold is placed, with no process to release them, is how legal-hold storage quietly grows without bound over years; holds need an owner and an expected review point, even if that review sometimes concludes the hold should continue.
- An enforcement job that isn't backed by an actual immutability guarantee is only as reliable as the script and the access controls around it; a compromised credential or an operator mistake could delete something the policy was supposed to protect, which is exactly the gap an immutability or object-lock mechanism closes.
Design a backup and disaster recovery system for 200 TB of production block storage spread across multiple data centers. Requirements: daily incremental backups, weekly full backups, RTO < 4 hours for critical datasets, RPO < 1 hour for highest-priority data, and retention/compliance policies. Detail architecture (snapshot vs block-level copy vs agent), cataloging, verification, restore runbooks, and how SREs should operate and test the system.
Sample Answer
Direct answer
At 200 TB, two things break the naive reading of the requirements and have to be designed around explicitly. First, "weekly full backup" cannot mean literally re-reading and re-copying all 200 TB every week: at a realistic sustained backup throughput of 500 MB/s, copying 200 TB takes roughly 200,000,000 MB / 500 MB/s ≈ 400,000 seconds ≈ 111 hours, well over four days, so a literal weekly full would never finish before the next one starts. The fix is snapshot-based, copy-on-write "synthetic full" backups: a full recovery point that is constructed from an incremental chain without re-reading the whole dataset. Second, RPO < 1 hour (Recovery Point Objective: the maximum acceptable data loss, measured as time) is stated only for the highest-priority subset of data, not the whole 200 TB; daily incrementals alone give an RPO of up to 24 hours, so that highest-priority subset needs a separate, more frequent capture mechanism layered on top of the daily/weekly baseline, not a tighter version of the same schedule applied everywhere.
Structured elaboration
Architecture: snapshot vs block-level copy vs agent, and why a hybrid.
- Storage-level snapshots (copy-on-write snapshots at the SAN, Storage Area Network, a dedicated network of shared block-storage devices, array, or cloud block-storage layer) are fast to create, capture only changed blocks, and are the right mechanism for the frequent, high-priority tier: they don't require reading the whole volume, so an hourly snapshot on a 200 TB estate is cheap regardless of total size. They are typically crash-consistent by default (consistent with an ungraceful power-off) and need an application-aware quiesce step (flushing writes, briefly freezing the filesystem) to be application-consistent for databases.
- Block-level backup/replication (deduplicated, changed-block backup software, or asynchronous block replication to the second data center) is what actually gets data across data centers for disaster recovery, since a local snapshot alone doesn't survive losing the data center it lives in. This is the mechanism that should carry the highest-priority subset's near-real-time copy to the DR site.
- Agent-based, file-level backup (an agent running inside each host or VM that understands the application, e.g. taking an application-consistent dump via a pre-freeze hook) is the most flexible but the slowest to scan at this scale, and shouldn't be the primary mechanism for the bulk of 200 TB. Reserve it for stateful applications (databases, message queues) that need an application-consistent capture the block layer can't provide on its own, layered on top of block-level backup for the rest.
- Recommended hybrid: block-level, copy-on-write snapshots as the default for all volumes (daily, feeding synthetic weekly fulls); asynchronous block-level replication to the second data center for the highest-priority subset, running frequently enough (e.g. every 15 to 30 minutes) to comfortably clear the 1-hour RPO with margin; agent-based, application-consistent backups only for databases and similar stateful services that need a coordinated quiesce, at the same cadence as their tier.
Worked restore-time example (basis: a single critical dataset's volume, not the full 200 TB). Suppose the highest-priority tier is a 10 TB subset. At an aggregate parallel restore throughput of 1 GB/s (achievable by restoring many volumes or shards concurrently rather than serially), restoring 10 TB takes 10,000 GB / 1 GB/s = 10,000 s ≈ 2.8 hours, comfortably inside the 4-hour RTO (Recovery Time Objective: the maximum acceptable downtime) for that dataset. This number is scoped to the 10 TB critical subset; restoring the full 200 TB estate at the same throughput would take roughly 200,000 GB / 1 GB/s ≈ 55.6 hours, which is why the RTO commitment applies per-dataset-tier, not as a promise to restore everything in 4 hours.
Cataloging. Maintain a backup catalog (an index of what was backed up, when, where, and with what checksum) as a small, independently replicated service, not embedded inside the 200 TB of bulk data itself. If the catalog only exists in the data center that just failed, nobody can find anything to restore even though the data may be intact elsewhere; replicate the catalog to every data center that could need to drive a restore.
Verification. Automated, scheduled restore tests: sample a set of volumes weekly, restore them into an isolated environment, and compare checksums against what the catalog recorded at backup time. Separately, run periodic bit-rot scans against the backup storage itself, since 200 TB sitting mostly cold for months can silently degrade without ever being touched by a restore. Neither of these is optional at this scale: an untested backup is a hypothesis, not a backup.
Restore runbooks. Document, per tier: who declares an incident and authorizes a restore, the priority order (highest-priority datasets first, since restore bandwidth is a shared, finite resource across a 200 TB estate and can't restore everything in parallel at full speed), the exact automated restore procedure (scripted, not ad hoc commands typed under pressure), and the rollback point if the restore itself needs to be aborted.
How SREs operate and test it. Track backup success rate, backup duration, and restore-verification pass rate as monitored SLOs (Service Level Objectives), with alerts on any missed window. Run quarterly full DR game days that actually fail a critical dataset over to the second data center and time every runbook phase against the 4-hour and 1-hour targets, rather than trusting the design math alone.
Unlock Full Question Bank
Get access to all Backup and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.