Backup and Disaster Recovery Questions
Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.
Write a POSIX-compatible Bash script that scans a directory of backup files, computes a SHA256 checksum for each new file, stores checksums in a manifest file in the form 'sha256 filename', and reports mismatches between the manifest and current files. Assume filenames include timestamps and the manifest lives beside the backups.
Sample Answer
Direct answer
The script below treats the manifest as an append-only ledger of trusted hashes: on each run it hashes every file currently in the directory, compares against any hash already recorded for that filename, appends a line for genuinely new files, reports (but does not silently "fix") any file whose current hash disagrees with its recorded hash, and separately reports any file the manifest still lists that is no longer on disk. It never overwrites a recorded hash once written, since silently updating a mismatched hash would defeat the entire point of tamper detection.
The script
#!/bin/sh
# backup-checksum-manifest.sh
# POSIX-compatible integrity manifest for a directory of backup files.
# Usage: ./backup-checksum-manifest.sh <backup_dir>
set -eu
BACKUP_DIR="${1:?usage: backup-checksum-manifest.sh <backup_dir>}"
MANIFEST_NAME="manifest.sha256"
MANIFEST="$BACKUP_DIR/$MANIFEST_NAME"
[ -d "$BACKUP_DIR" ] || { echo "error: $BACKUP_DIR is not a directory" >&2; exit 1; }
[ -f "$MANIFEST" ] || : > "$MANIFEST"
# Pick an available SHA256 tool (GNU coreutils, BSD/macOS, or openssl fallback).
sha256_of() {
if command -v sha256sum >/dev/null 2>&1; then
sha256sum "$1" | awk '{print $1}'
elif command -v shasum >/dev/null 2>&1; then
shasum -a 256 "$1" | awk '{print $1}'
else
openssl dgst -sha256 "$1" | awk '{print $NF}'
fi
}
FILE_LIST="$(mktemp "${TMPDIR:-/tmp}/filelist.XXXXXX")"
NEW_LINES="$(mktemp "${TMPDIR:-/tmp}/newlines.XXXXXX")"
trap 'rm -f "$FILE_LIST" "$NEW_LINES"' EXIT
find "$BACKUP_DIR" -maxdepth 1 -type f ! -name "$MANIFEST_NAME" | sort > "$FILE_LIST"
new_count=0
mismatch_count=0
ok_count=0
while IFS= read -r file; do
fname=$(basename "$file")
current_hash=$(sha256_of "$file")
recorded_hash=$(awk -v f="$fname" '$2 == f { print $1; exit }' "$MANIFEST")
if [ -z "$recorded_hash" ]; then
printf '%s %s\n' "$current_hash" "$fname" >> "$NEW_LINES"
new_count=$((new_count + 1))
printf 'NEW: %s -> recorded %s\n' "$fname" "$current_hash"
elif [ "$current_hash" = "$recorded_hash" ]; then
ok_count=$((ok_count + 1))
else
mismatch_count=$((mismatch_count + 1))
printf 'MISMATCH: %s manifest=%s actual=%s\n' "$fname" "$recorded_hash" "$current_hash"
fi
done < "$FILE_LIST"
# Files the manifest still lists that are no longer on disk.
missing_count=0
while IFS= read -r line; do
[ -z "$line" ] && continue
mfname=$(printf '%s\n' "$line" | awk '{print $2}')
[ -f "$BACKUP_DIR/$mfname" ] || {
missing_count=$((missing_count + 1))
printf 'MISSING: %s is in the manifest but not on disk\n' "$mfname"
}
done < "$MANIFEST"
# Append newly discovered files to the manifest.
if [ -s "$NEW_LINES" ]; then
cat "$NEW_LINES" >> "$MANIFEST"
fi
printf '\nSummary: %d new, %d ok, %d mismatched, %d missing\n' \
"$new_count" "$ok_count" "$mismatch_count" "$missing_count"
[ "$mismatch_count" -eq 0 ] && [ "$missing_count" -eq 0 ]
Worked example: run cold, in a fresh directory, real output
This exact script was extracted and run unmodified in an empty scratch directory on a real machine (macOS/BSD userland, no GNU coreutils sha256sum on PATH, so the script's shasum -a 256 fallback path is the one actually exercised). Two files were created first:
$ printf 'fake tarball contents v1\n' > backups/db-2026-08-10T02-00-00Z.tar.gz
$ printf 'fake tarball contents v2\n' > backups/db-2026-08-11T02-00-00Z.tar.gz
First run, no manifest exists yet, actual output:
$ sh ./backup-checksum-manifest.sh backups
NEW: db-2026-08-10T02-00-00Z.tar.gz -> recorded 62dec6cc1e4b0f9fd450c8050bfb18d2d0aff312327cab071e10297754ff5d5c
NEW: db-2026-08-11T02-00-00Z.tar.gz -> recorded 864d22f4d6c0310333126b63ba757eb1a29b1cc0387ef58d9718ca3e23f6b8be
Summary: 2 new, 0 ok, 0 mismatched, 0 missing
exit code: 0
The manifest file written to disk:
62dec6cc1e4b0f9fd450c8050bfb18d2d0aff312327cab071e10297754ff5d5c db-2026-08-10T02-00-00Z.tar.gz
864d22f4d6c0310333126b63ba757eb1a29b1cc0387ef58d9718ca3e23f6b8be db-2026-08-11T02-00-00Z.tar.gz
Second run, nothing changed, actual output confirms idempotency (no NEW/MISMATCH/MISSING lines, everything counted as ok):
$ sh ./backup-checksum-manifest.sh backups
Summary: 0 new, 2 ok, 0 mismatched, 0 missing
exit code: 0
Third run, after mutating one file's contents, deleting a second, and adding a third new file:
$ printf 'fake tarball contents v1 CHANGED\n' > backups/db-2026-08-10T02-00-00Z.tar.gz
$ rm backups/db-2026-08-11T02-00-00Z.tar.gz
$ printf 'fake tarball contents v3\n' > backups/db-2026-08-12T02-00-00Z.tar.gz
$ sh ./backup-checksum-manifest.sh backups
MISMATCH: db-2026-08-10T02-00-00Z.tar.gz manifest=62dec6cc1e4b0f9fd450c8050bfb18d2d0aff312327cab071e10297754ff5d5c actual=d1391900c0ad53a29bd0602e2daea24bc901a31e12110f3a4706d38bdafa17a6
NEW: db-2026-08-12T02-00-00Z.tar.gz -> recorded 3385e1df335ad631fb58fc8bb6220f9e07795bdb0181336dd7413cb8cebbc882
MISSING: db-2026-08-11T02-00-00Z.tar.gz is in the manifest but not on disk
Summary: 1 new, 0 ok, 1 mismatched, 1 missing
exit code: 1
This confirms all three detection paths (new, mismatched, missing) fire correctly on the same run, the manifest gained only the genuinely new file's line rather than overwriting the mismatched entry, and the exit code turns nonzero specifically when integrity problems exist, which is what lets this be wired into cron with a simple || alert rather than requiring a human to read the log every time.
Trade-offs and design notes
- The manifest deliberately never rewrites a mismatched or missing entry automatically. If it did, running the script again after tampering would silently "launder" the tampered file into looking trusted again, which defeats the purpose. Updating a recorded hash after investigating a mismatch is a decision left to a human (or a separate, explicitly-named remediation step), not something this script does implicitly.
- The exit code is meaningful.
0only when there are zero mismatches and zero missing files, so the script composes cleanly with cron and monitoring (backup-checksum-manifest.sh /backups || send_alert). - Filenames with embedded timestamps need no special handling here: the manifest matches purely by exact filename string, so
db-2026-08-10T02-00-00Z.tar.gzis just an opaque key, sorting and matching correctly regardless of the timestamp format chosen. - POSIX portability was the actual constraint that mattered: the naive version of this script piped
find | while readdirectly, which runs the loop in a subshell (a copy of the shell's environment spawned to run the piped command) under a strict POSIX shell and silently loses thenew_count/mismatch_counttotals accumulated inside it. The shipped version instead writes the file list to a temp file first and redirects it into thewhile readloop (done < "$FILE_LIST"), which does not fork a subshell, so the counters set inside the loop are still visible in the final summary line.
Explain immutable backups and how they help mitigate ransomware and accidental deletion. Compare technology implementations such as WORM tape, object-store immutability (e.g., S3 Object Lock governance vs compliance), and immutable storage-layer snapshots. Discuss operational limitations and management practices.
Sample Answer
Direct answer
An immutable backup is one that, once written, cannot be modified, overwritten, or deleted, even
by an administrator account, until a defined retention period expires. The protection is enforced
by the storage layer itself rather than by an access-control policy that a compromised credential
could simply bypass, which is exactly what makes it effective against ransomware and accidental
deletion: it removes deletion and overwrite as an available action for any credential during the
lock period, whether the attempt is malicious or a mistake.
Structured elaboration
Why it mitigates ransomware. Ransomware that has obtained administrator or backup-operator
credentials can typically delete or encrypt over any backup those credentials can reach, defeating
the plan to simply restore. Immutability closes that path: even a fully compromised administrator
account cannot destroy a copy under an active lock, turning "the attacker also destroyed the
backups" from a real threat into a non-issue for anything currently locked.
Why it mitigates accidental deletion. The same mechanism protects against a human running the
wrong delete command or a buggy automation script wiping more than intended. The lock does not
distinguish malicious intent from an honest mistake; it simply refuses the action either way.
Comparing three implementations.
- Write Once, Read Many (WORM) tape. Media that is mechanically and logically write-once, and
typically air-gapped, physically disconnected from the network when not actively being written
to or read from. That physical disconnection means it is immune to network-based ransomware
entirely, not just logically protected. The trade-off is speed and handling: sequential access,
slower restores that involve locating and mounting the correct tape, and a real operational
discipline required to rotate and store it correctly. - Object-store immutability, for example AWS S3's Object Lock feature, which offers a
governance mode and a stricter compliance mode. Governance mode locks objects against deletion
for ordinary users but leaves an emergency override available to specially privileged accounts,
useful for legitimate corrections, though that override is itself an attack surface if the
privileged credential holding it is compromised. Compliance mode locks the object against
everyone, including the account's own administrators, strictly enforcing the retention clock with
no override at all; the trade-off is that a mistakenly long retention setting cannot be undone
either, so the retention length has to be a deliberate, carefully considered decision made up
front. - Immutable storage-layer snapshots, for instance on a storage array or backup appliance. Fast,
since it is typically block-level and local rather than tape, and integrates naturally with an
existing snapshot-based backup workflow. The trade-off is that this storage usually remains
network-reachable, unlike air-gapped tape, so its protection depends entirely on the immutability
enforcement genuinely being unbypassable at the storage layer and on the management interface
itself being hardened against privileged misuse.
Operational limitations and management practices. Locked data continues to consume and bill
for storage for the entire lock duration; it cannot be cleaned up early even if you want to, so the
retention length has to be chosen carefully, too long wastes budget, too short can be waited out by
a patient attacker. Immutability protects against deletion and overwrite, but it does not protect
against writing bad or already-compromised data as a new immutable object in the first place: if
ransomware encrypts source data and a backup job faithfully backs up the already-encrypted files as
that cycle's backup, the resulting immutable copy is preserved exactly as intended, and exactly
useless. Combine immutability with real restore verification (checking which restore points
actually predate the compromise) rather than assuming any locked copy is automatically clean, and
choose a retention length that reaches back far enough to cover the likely period an attacker sat
undetected before triggering encryption, which security research generally describes as ranging
from days to weeks depending on the case, though the exact figure varies considerably and should
not be treated as a fixed constant.
Problem-solving: Given a constrained backup budget, describe how you would schedule backups and retention for different data classes (transaction logs, databases, large file shares, VM images) to balance cost and business risk. Provide concrete suggestions (frequency, storage tier, and retention) and explain trade-offs.
Sample Answer
Direct answer
Do not optimize each data class in isolation. Model total cost across all of them, then push the
biggest, lowest-risk class as aggressively as possible toward a cheap tier and short retention,
because that is where a constrained budget actually gets freed up, and spend the savings on
keeping the small but critical classes on a fast tier with a tight schedule.
Structured elaboration and concrete plan
- Transaction logs. Frequency: continuous or near-continuous shipping, or every 5 to 15
minutes for systems that cannot ship continuously. Storage tier: standard or hot, since logs
exist specifically to support fast, fine-grained point-in-time recovery alongside a full or
differential backup. Retention: short, roughly 7 to 14 days, since logs are only useful in
combination with the full and differential chain they support; keeping logs long after the
backups they pair with have expired protects nothing. - Databases (the full and differential backups themselves). Frequency: nightly full or a
weekly full with nightly differentials. Storage tier: standard while recent, moving older
restore points to a cheaper tier as they age. Retention: roughly 30 to 90 days for operational
recovery, longer only where compliance specifically requires it. - Large file shares. Frequency: daily incremental with a weekly full, or an incremental-forever
model given the typical size. Storage tier: cool or infrequent-access, since restores of file
shares are comparatively rare. Retention: 30 to 90 days of operational retention, with an
optional monthly or yearly snapshot moved to a cold archive tier for anything that needs to be
kept longer at a much lower ongoing cost. - VM images. Frequency: daily image-level backup. Storage tier: standard for only the most
recent restore point (fast recovery after a bad deploy or patch), cold or cool for older restore
points. Retention: a grandfather-father-son style scheme, for example daily for one to two weeks,
weekly for a couple of months, monthly for a year.
Worked example: where the budget actually goes
Take illustrative footprints (the resulting storage under the retention choices above, not a
separate derivation): databases 5,000 GB, transaction logs 200 GB, file shares 50,000 GB, VM images
20,000 GB, for 75,200 GB total. Using illustrative prices of $0.02 per GB-month on a standard tier
and $0.01 per GB-month on a cool tier (order-of-magnitude realistic ratios, not live vendor
pricing), keeping everything on the standard tier costs 75,200 GB times $0.02 per GB-month, or
$1,504 per month.
A tiered plan instead keeps databases and logs (5,200 GB) on standard, since they need fast
restores: 5,200 times $0.02 is $104 per month. File shares (50,000 GB) move entirely to the cool
tier: 50,000 times $0.01 is $500 per month. VM images split, with roughly the most recent
restore point (about 1,400 GB) staying on standard for fast recovery and the remaining 18,600 GB
on cool: 1,400 times $0.02 is $28, plus 18,600 times $0.01 is $186, for $214 per month.
Total under the tiered plan: 104 plus 500 plus 214, or $818 per month, versus $1,504 per month if
everything stayed on standard, a saving of $686 per month, roughly 46%. Of that $686 in savings,
$500, about 73%, comes purely from moving the file shares to the cool tier. That is the point: file
shares are the largest volume and the lowest per-byte business risk of the four classes, so tiering
them is the single highest-leverage lever, and it is what funds keeping the small, high-value
database and log tiers fast and frequent.
Trade-offs
Cheaper tiers add retrieval latency and sometimes retrieval fees. The transaction logs and the most
recent database and VM restore points must not be pushed to a cold tier even though it is cheaper,
because the entire point of frequent logging and recent-point backups is fast recovery; a
budget-constrained plan protects recovery time and recovery point for the highest-risk classes
first, then applies cost savings to the classes that can genuinely absorb slower restores.
Explain deduplication and compression in backup systems. Discuss trade-offs between source-side dedupe and target-side dedupe, CPU and memory requirements, impact on restore performance, and practical scenarios where you might disable or tune these features.
Sample Answer
Direct answer
Deduplication removes redundancy across chunks of data by storing each unique chunk once and
replacing later occurrences with a small reference. Compression removes redundancy within a single
chunk by encoding it more compactly. They solve different problems and are typically used together,
not as alternatives. Where they trade off is in where their CPU and memory cost lands (on the
production source or on the backup repository), and in a restore-performance cost that dedup
specifically introduces and that is easy to underestimate.
Structured elaboration
Source-side versus target-side dedup. Source-side dedup hashes and deduplicates data on the
client or source machine before transmitting it, so only unique, new chunks travel over the
network. That saves real bandwidth, which matters a lot for large datasets or slow, expensive
links, but it spends CPU and memory on the production source machine itself, competing directly
with the workload that machine is actually there to run. Target-side dedup sends the full,
undeduplicated data across the network and deduplicates it at the backup repository. That protects
production resources, but uses full bandwidth on every backup with no network savings, and shifts
the CPU and memory burden to the repository, which needs to be sized for it: dedup at scale,
especially global dedup across many clients, needs a large, fast-lookup index of chunk hashes, and
an undersized index causes dedup throughput to collapse.
CPU and memory requirements. Dedup fundamentally needs a hash index sized to the total number
of unique chunks in the repository; larger dedup pools need more memory, or a fast SSD-backed
index, to keep lookups fast, and hashing itself is CPU-bound. Compression is CPU-bound too:
stronger compression ratios generally cost more CPU time per byte processed, which is its own
basis-labeled trade-off, compression ratio (bytes saved as a percentage of the original) against
CPU-seconds spent per GB, and where the right point on that curve sits depends on whether your
actual bottleneck is storage cost or backup-window time.
Impact on restore performance. This is the most commonly underestimated cost. Deduplicated
data is, by construction, no longer stored as one contiguous file; it is scattered chunks that must
be reassembled by following references, often from many different physical locations in the
repository. That fragmentation can show up as noticeably slower sequential restore throughput
compared to reading an undeduplicated backup, especially for older data whose chunks have been
shared and moved around over a long time. Compression's impact on restore is smaller and more
predictable: decompression is generally fast and cheap relative to compression, so it rarely
dominates restore time the way dedup reassembly overhead can.
Practical scenarios to disable or tune. Reduce or disable dedup for data that is already
compressed or encrypted at the source, media files, pre-compressed archives, or encrypted database
backups, since such data has little to no duplicate-chunk redundancy left to find; dedup CPU and
memory cost is then spent for near-zero storage benefit, sometimes negative once the dedup engine's
own metadata overhead is counted. Prefer target-side dedup, or reduce source-side dedup intensity,
for CPU or memory-constrained production servers that cannot spare cycles for backup-time hashing
during business hours. Prefer source-side dedup where the bottleneck is a genuinely constrained or
expensive network link and the source machines have spare CPU and memory headroom. Tune the most
restore-time-critical systems toward less aggressive dedup, or a periodic consolidation pass that
reduces fragmentation, specifically because heavy fragmentation's restore-speed cost can violate a
recovery time commitment even while it saves storage; make that trade-off deliberately, tier by
tier, rather than applying a single global default everywhere.
Worked example
Take 100 client machines each backing up the same 2 GB OS image, for a naive undeduplicated total
of 100 times 2 GB, or 200 GB. If 95% of each image is identical across clients, the shared portion,
2 GB times 0.95, or 1.9 GB, is stored once, and each client's unique 5%, 2 GB times 0.05, or 0.1 GB,
is stored per client: 0.1 GB times 100 clients, or 10 GB. Total physical storage is 1.9 GB plus 10
GB, or 11.9 GB, versus the 200 GB naive total, a reduction ratio of roughly 200 divided by 11.9, or
about 16.8 times. That is why dedup matters enormously for a fleet of similar systems, homogeneous
VM images or OS installs in particular, while offering little benefit for a set of already unique,
already compressed files.
Design an automated restore verification harness that periodically restores backups into isolated test environments and runs application-level acceptance tests across multiple stacks (web app, database, message queue). Describe orchestration, environment provisioning (IaC), test selection, data obfuscation for PII, pass/fail criteria, and reporting/alerting integration.
Sample Answer
Direct answer
The harness is a pipeline, not a script: an orchestrator triggers on a schedule or after each backup completes, provisions a fully isolated environment via Infrastructure as Code (IaC), restores the target backup into it, masks any PII before anything or anyone else can touch the data, runs a tiered set of tests from cheap-and-fast to expensive-and-thorough, evaluates explicit pass/fail criteria, and reports results with alerting wired to page on failure, then tears the environment down regardless of outcome so cost stays bounded.
Orchestration
A scheduler or pipeline service drives the whole run as an explicit state machine (provision, restore, mask, test, report, teardown), triggered either on a cadence (nightly for critical systems, weekly for the rest) or immediately after each new backup job completes, so a bad backup is caught close to when it was created rather than discovered weeks later during an actual incident. Each stage has its own timeout and retry policy so a single stuck restore does not block the entire pipeline indefinitely.
Environment provisioning (IaC)
Every run gets a fully isolated environment: a dedicated network segment or namespace with no route to production, provisioned via Terraform, Ansible, or Kubernetes manifests, sized close enough to production for the test to be meaningful. Teardown is automatic and unconditional (pass or fail), and teardown itself should be verified with a periodic orphaned-resource sweep, since a harness that silently leaks infrastructure on failed runs quietly becomes an expensive habit nobody notices until the cloud bill does.
Test selection, tiered by cost
- Cheapest, every run: did the restore complete, did the system boot or start successfully.
- Fast, every run: data integrity checks (checksums, row counts, referential integrity).
- Moderate, every run: application-level smoke tests, can the app start, connect to its restored database, and serve a handful of key endpoints.
- Expensive, rotating subset: the same full acceptance test suite used in CI, run against the restored stack, on a schedule lighter than nightly (for example weekly) or specifically before a run is counted as an official quarterly drill, since running the full suite on every single nightly restore is usually not affordable at scale.
Data obfuscation for PII
A restored environment is, by construction, a full copy of production data landing in a less-trusted test context, so masking has to happen automatically, as a gate, before any test or human gets access to the restored data, not as an optional cleanup step afterward. Use deterministic tokenization, the simpler default choice, or format-preserving encryption, worth the extra complexity only when the masked value must still pass a format validator downstream, for fields that need to stay joinable across tables (a customer_id that multiple tables reference), and irreversible masking for free-text or clearly sensitive fields (SSNs, free-text notes). No query against the restored environment, automated or human, should be possible until masking has run and been verified complete.
Pass/fail criteria
Restore completed within the target time (measured end to end from trigger to verified functional, on the same basis as the system's RTO, not from "data copy complete," which understates the real recovery time by omitting validation). Zero integrity-check failures. All application smoke tests passing. On the rotating full-suite tier, an acceptance-test pass rate above a defined threshold. Any failure marks that specific backup generation as suspect and pages the on-call rather than silently retrying until it happens to pass, which would hide a real, recurring problem.
Reporting and alerting integration
Every run writes a structured result (pass/fail per stage, timing, logs) to a results store so trends are visible over time, for example restore time creeping upward month over month well before it breaches SLA. Failures alert immediately, paging for Tier-0 systems and filing a ticket for lower tiers, with enough context (which stage failed, relevant logs) attached that triage does not require re-running the whole pipeline first. A periodic rollup (monthly, say) reports restore-success rate to leadership as a real reliability metric, distinct from and more meaningful than "backups completed," which only proves bytes were written, not that they are usable.
Unlock Full Question Bank
Get access to all Backup and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.