Backup and Disaster Recovery Questions
Keeping data durable and recoverable when systems fail: backup design (full, incremental, differential, and snapshot strategies; point-in-time recovery), backup verification and restore testing, retention and archival policy (including compliance retention and legal holds), encryption and key management for backups, and disaster-recovery planning measured against recovery-time and recovery-point objectives (RTO/RPO). Tests whether a candidate can design a backup strategy that actually restores, choose the right retention tier for a business's downtime and data-loss tolerance, and operate backup systems safely under failure, compliance, and ransomware threats. Distinct from code-level fault-tolerance patterns (circuit breakers, retries, bulkheads) and multi-region failover architecture, which belong to high-availability-and-disaster-recovery.
Design a retention lifecycle for healthcare records that must meet HIPAA compliance with auditable controls and 7-year retention. Include archival tiering, deletion policies, legal hold handling, integrity checks, and cost-control measures while ensuring records can be retrieved when needed.
Sample Answer
Direct answer
Design the lifecycle around one governing rule and one override: records age from hot to warm to cold storage automatically on a schedule tied to the 7-year retention clock, and a legal hold can freeze that clock for an individual record at any point, overriding the schedule entirely, with every transition, access, and deletion logged so the whole lifecycle is independently auditable. It is worth noting as a framing point, not a correction to the question, that the specific "7 years" figure is typically driven by state medical-records law or an organization's own retention policy rather than by HIPAA itself, which principally sets minimum standards for security, privacy, and audit controls rather than a fixed record-retention duration; the design below treats 7 years as the given business requirement regardless of which regulation ultimately drives that number.
Archival tiering
- Hot (roughly years 0-2): actively-referenced records, ongoing patient care, fast retrieval expected, standard low-latency storage.
- Warm (roughly years 2-4): infrequent access, still fast enough for a routine records request, moved to a cheaper storage class via an automated lifecycle policy keyed off record creation or last-access date.
- Cold/archive (roughly years 5-7): rarely accessed, retained principally for compliance and legal-discovery readiness, cheapest storage tier, retrieval latency of hours is an acceptable and explicit trade-off, not an accident.
Encryption is maintained across every tier, not just the hot tier, since a "cheaper, rarely accessed" tier is not a "less protected" tier.
Deletion policies
Each record gets a computed expiry date at the 7-year mark from its creation or applicable retention trigger. Deletion is a two-phase process: mark for deletion, hold for a short verification window, then irreversible purge, with an audit log entry proving what was deleted and when. This matters in both directions: an organization must be able to prove it retained records for the required period, and equally must be able to prove it actually deleted records once the retention period expired, since over-retention (keeping protected health information longer than justified) is its own compliance and breach-exposure risk, not a conservatively safe default.
Legal hold handling
A legal_hold flag, independent of the computed expiry date, overrides normal expiration entirely: a record under hold cannot be deleted regardless of its age, applied when litigation, an audit, or a compliance investigation requires preserving specific records. Removing a hold is itself a controlled, logged, authorized action, ideally requiring approval from someone other than whoever applied it, so a hold cannot be silently lifted by the same party that has an interest in a record disappearing.
Integrity checks
Periodic checksum verification of archived records, since bit rot and media degradation are real risks over a 7-year dormancy window in cold storage, not a theoretical concern. HIPAA's Security Rule also expects audit trails of who accessed what, so access logging runs continuously across all tiers, not only the hot tier where access is more common. On the same cadence as a general restore-drill program, periodically sample-restore records from cold storage to confirm they are actually retrievable, not merely present and un-corrupted on paper.
Cost-control measures
Tiering itself is the primary lever: moving 5-year-old, rarely-accessed records to the cheapest appropriate tier rather than leaving everything in hot storage indefinitely out of caution. Compress archived records where the storage format allows it. Deduplicate where legally and technically permissible. Most importantly, retain to the true legal minimum rather than defaulting to "keep everything forever" as a safety instinct, since over-retention is not free: it expands the blast radius of any future breach to include data the organization had no ongoing obligation to hold.
Worked example
Patient record R-48213 is created on 2016-01-10, giving it a computed expiry of 2023-01-10. It moves through the normal tier schedule on that basis: hot through early 2018, warm through roughly 2020, cold from around 2021 onward, all driven purely by its age, independent of anything that happens later. In 2022-08-01, an unrelated compliance audit requires preserving the record, and a legal_hold flag is set on it. The record's computed expiry of 2023-01-10 arrives while the hold is still active: the scheduled deletion sweep checks the flag, finds it set, skips deletion, and logs the skip rather than purging on schedule. The hold stays in place for over a year past that expiry date. On 2024-11-15, the audit concludes and someone other than the person who originally set the hold releases it, logging the release as its own authorized action. With no hold protecting it and its expiry already well in the past, the next scheduled sweep marks R-48213 for deletion, holds it for the short verification window, then purges it, writing an audit log entry that records the original creation date, the computed expiry, the hold window, and the actual deletion date, so an auditor can see exactly why a record with a 2023 expiry was not actually removed until 2024.
Ensuring records remain retrievable when needed
Even cold-tier records need a bounded, documented retrieval SLA agreed with legal and compliance stakeholders (for example, any record retrievable within 24 to 48 hours), since audit responses and legal discovery requests carry their own deadlines that "eventually retrievable" does not satisfy. "Cold" should mean slower and cheaper, never "practically unreachable."
Auditable controls, tying it together
Every mechanism above (tier transition, access, hold applied or removed, deletion) needs to leave an audit trail satisfying the general audit-controls expectation under the HIPAA Security Rule (the exact regulatory citation is not asserted here with full confidence and should be confirmed against current legal guidance before this design is treated as a compliance sign-off), so the full lifecycle is demonstrably governed, not just technically implemented.
Design a bring-your-own-key (BYOK) backup encryption scheme where only the customer holds the decryption keys, stored in an enterprise KMS/HSM. Walk through the key lifecycle, and be specific: what happens to already-encrypted backups when you rotate a key, and what happens if the customer loses it?
Sample Answer
Direct answer
Bring-your-own-key (BYOK) means the customer generates and owns the root encryption key, typically
in their own key management service (KMS, a managed service for creating, storing, and rotating
cryptographic keys) or hardware security module (HSM, tamper-resistant hardware that performs
cryptographic operations without ever exposing the raw key material). The vendor's ability to
encrypt or decrypt the customer's backups depends entirely on being granted narrow, revocable
access to that key, not on holding a key of its own. That single fact determines everything about
rotation and loss below.
Key lifecycle
- Generation. The customer generates a root key in their own KMS or HSM.
- Envelope encryption. Encrypting large backup volumes directly with a tightly controlled root
key every time would be slow and hard to audit at scale. Instead, the backup software generates
a random, per-backup data encryption key (DEK), encrypts the actual backup data with that DEK,
and then encrypts, or wraps, the small DEK itself with the customer's root key. The root key
never directly touches the bulk data; it only ever wraps and unwraps small DEKs, which limits how
much data any single root-key operation exposes and keeps root-key usage cheap enough to call
frequently. - Authorization. The vendor is granted a narrow, auditable permission to call encrypt and
decrypt against the customer's key, not the raw key material itself, so every use is logged in
the customer's own KMS audit trail. - Escrow. Some organizations choose to escrow, securely deposit a backup copy of the root key
with a trusted third party or an internal break-glass vault, specifically as a hedge against the
loss scenario below. This is optional and a genuine trade-off: an escrow copy is another location
a determined attacker or insider could target, weighed against the risk of permanently
unrecoverable data with no escrow at all. - Rotation. The root key is periodically replaced with a new one, a standard security practice
and sometimes a compliance requirement. - Revocation. The customer can revoke the vendor's permission to use the key at any time, for
example when offboarding the vendor or responding to a suspected compromise. Revocation prevents
future encrypt and decrypt calls immediately; it does not retroactively undo anything already
decrypted before revocation.
What happens to already-encrypted backups when you rotate a key
Rotation does not re-encrypt the bulk backup data. Because of envelope encryption, the root key
only ever wrapped the small per-backup DEKs, never the data itself, so rotating it means new
backups going forward get DEKs wrapped with the new key version, while existing backups keep their
DEKs wrapped with the old version. The KMS must therefore keep honoring the old key version, not
truly delete it, for as long as any backup encrypted under it still needs to be restorable. In
practice, key rotation in a KMS usually means versioning, retaining every prior version available
for unwrapping, rather than true replacement. The consequence is a basis point worth stating
explicitly: your key-retention policy must be at least as long as your data-retention policy, on
the same time basis, or you will make an old but still-required backup permanently unreadable.
What happens if the customer loses the key
Because BYOK means the vendor never independently held usable key material, losing the root key,
with no escrow copy and no other recoverable path to the wrapped DEKs, makes the encrypted backup
data permanently unrecoverable. This is by design, not a bug: no one, including the vendor, holds
a master key capable of decrypting it. That is precisely the trade-off of true customer-held keys:
it is the strongest available guarantee against the vendor, or anyone who compromises the vendor,
reading your data, and it is equally strong against your own ability to recover if the key is
mismanaged. Practical mitigation is a mandatory escrow or break-glass procedure with strict,
multi-party access controls, or accepting the loss risk as a documented, deliberate trade-off in
exchange for the security guarantee.
Functional limitations worth weighing alongside the security benefit
BYOK has a real functional cost beyond the lifecycle mechanics above. Cross-customer or
cross-backup deduplication, finding identical chunks of data across different backups to store
them once, generally stops working across customers using different keys, since two customers'
identical files now encrypt to different ciphertext under different keys and look like unrelated
random data to a deduplication engine, which can materially increase storage cost in a
multi-tenant repository. Similarly, any server-side search or indexing feature that lets a customer
search inside backed-up content without a full restore generally cannot operate on data it cannot
decrypt, so BYOK customers often lose that convenience unless the vendor has specifically
architected search to run inside a boundary the customer still controls. Both are genuine,
sometimes underappreciated costs of BYOK that deserve to be weighed against its security benefit
rather than treated as free.
Explain immutable backups and how they help mitigate ransomware and accidental deletion. Compare technology implementations such as WORM tape, object-store immutability (e.g., S3 Object Lock governance vs compliance), and immutable storage-layer snapshots. Discuss operational limitations and management practices.
Sample Answer
Direct answer
An immutable backup is one that, once written, cannot be modified, overwritten, or deleted, even
by an administrator account, until a defined retention period expires. The protection is enforced
by the storage layer itself rather than by an access-control policy that a compromised credential
could simply bypass, which is exactly what makes it effective against ransomware and accidental
deletion: it removes deletion and overwrite as an available action for any credential during the
lock period, whether the attempt is malicious or a mistake.
Structured elaboration
Why it mitigates ransomware. Ransomware that has obtained administrator or backup-operator
credentials can typically delete or encrypt over any backup those credentials can reach, defeating
the plan to simply restore. Immutability closes that path: even a fully compromised administrator
account cannot destroy a copy under an active lock, turning "the attacker also destroyed the
backups" from a real threat into a non-issue for anything currently locked.
Why it mitigates accidental deletion. The same mechanism protects against a human running the
wrong delete command or a buggy automation script wiping more than intended. The lock does not
distinguish malicious intent from an honest mistake; it simply refuses the action either way.
Comparing three implementations.
- Write Once, Read Many (WORM) tape. Media that is mechanically and logically write-once, and
typically air-gapped, physically disconnected from the network when not actively being written
to or read from. That physical disconnection means it is immune to network-based ransomware
entirely, not just logically protected. The trade-off is speed and handling: sequential access,
slower restores that involve locating and mounting the correct tape, and a real operational
discipline required to rotate and store it correctly. - Object-store immutability, for example AWS S3's Object Lock feature, which offers a
governance mode and a stricter compliance mode. Governance mode locks objects against deletion
for ordinary users but leaves an emergency override available to specially privileged accounts,
useful for legitimate corrections, though that override is itself an attack surface if the
privileged credential holding it is compromised. Compliance mode locks the object against
everyone, including the account's own administrators, strictly enforcing the retention clock with
no override at all; the trade-off is that a mistakenly long retention setting cannot be undone
either, so the retention length has to be a deliberate, carefully considered decision made up
front. - Immutable storage-layer snapshots, for instance on a storage array or backup appliance. Fast,
since it is typically block-level and local rather than tape, and integrates naturally with an
existing snapshot-based backup workflow. The trade-off is that this storage usually remains
network-reachable, unlike air-gapped tape, so its protection depends entirely on the immutability
enforcement genuinely being unbypassable at the storage layer and on the management interface
itself being hardened against privileged misuse.
Operational limitations and management practices. Locked data continues to consume and bill
for storage for the entire lock duration; it cannot be cleaned up early even if you want to, so the
retention length has to be chosen carefully, too long wastes budget, too short can be waited out by
a patient attacker. Immutability protects against deletion and overwrite, but it does not protect
against writing bad or already-compromised data as a new immutable object in the first place: if
ransomware encrypts source data and a backup job faithfully backs up the already-encrypted files as
that cycle's backup, the resulting immutable copy is preserved exactly as intended, and exactly
useless. Combine immutability with real restore verification (checking which restore points
actually predate the compromise) rather than assuming any locked copy is automatically clean, and
choose a retention length that reaches back far enough to cover the likely period an attacker sat
undetected before triggering encryption, which security research generally describes as ranging
from days to weeks depending on the case, though the exact figure varies considerably and should
not be treated as a fixed constant.
Explain what an air-gapped backup is, why organizations use air gaps as part of a defense-in-depth strategy, and describe two practical ways to implement air-gapped backups in either a cloud or hybrid environment while keeping them usable for restores.
Sample Answer
Direct answer
An air-gapped backup is a copy of data kept physically or logically disconnected from the production network and from anything a production compromise could reach, so an attacker (or a runaway automated process) who has taken over production still cannot delete, encrypt, or tamper with that copy, because there is no live path to it at the time of the attack. Organizations use air gaps as the outermost layer of defense-in-depth because modern ransomware routinely targets backup infrastructure directly, and an air gap is the one layer whose protection does not depend on any other layer (perimeter, endpoint, identity) having held.
Why air gaps matter for defense-in-depth
The classic guidance here is the 3-2-1-1 rule: 3 copies of data, on 2 different media types, with 1 copy offsite, and 1 copy air-gapped or otherwise immutable. Every layer before the air-gapped copy can fail: perimeter defenses can be bypassed, endpoint protection can miss a novel payload, and identity can be compromised via a phished or stolen privileged credential. An attacker who reaches domain-admin-equivalent access in a normally-connected environment can typically also reach and destroy normally-connected backups, which is exactly why real-world ransomware incidents increasingly include the backup repository as a target, not just production. The air gap's value is that it does not rely on any of those defenses continuing to hold; it removes the network path itself.
Two practical implementations
- Logical air gap via immutable, access-isolated object storage. Use object storage with a write-once retention lock (for example, an object-lock feature in compliance mode) so that once written, an object cannot be deleted or modified until its retention period expires, not even by an account with otherwise-broad permissions. Combine this with a credential and account boundary that production's normal operating identities never hold: a separate account or tenant, deny-by-default cross-account access, and a strictly one-way replication path (production can push new backups in, nothing, including a fully compromised production credential, can delete or modify what is already there). This is "logical" rather than physical: there is no literal cable being unplugged, but access is architecturally and cryptographically isolated such that compromising production does not grant a path to alter the vaulted copy.
- Physical or scheduled-connection air gap. Traditional tape or removable-media rotation, still common in hybrid environments, where media is physically disconnected from any network once a write completes. A cloud-native equivalent is a vaulting pattern where a secondary environment is connected only during a scheduled replication window: a nightly job establishes a one-way connection, pushes data across, then the connection and its credentials are torn down or rotated, so for the overwhelming majority of the day there is no live network path at all, minimizing the window an attacker could exploit even if they somehow reached the vault's edge.
Keeping air-gapped backups usable for restores
The practical trap with a true air gap is that "hard to reach for an attacker" often also means "hard to reach quickly during a legitimate emergency." Mitigate this with a documented, drilled break-glass retrieval process (not invented for the first time during an actual incident), retrieval credentials staged in a separate, monitored, multi-factor-gated vault distinct from production's normal credential store, and periodic test restores from the air-gapped copy on the same cadence as regular restore drills, so the air-gapped copy does not quietly become the one backup class nobody has ever actually verified is restorable.
A lawsuit requires placing a legal hold on 20 TB of customer data that must be preserved for 10 years. Design a cost-effective, searchable retention architecture that guarantees integrity and auditability and supports e-discovery retrieval within 48 hours. Explain indexing, storage tiering, immutability, and validation strategies.
Sample Answer
Direct answer
Separate the content from the index. Put the bulk 20 TB of held data on the cheapest tier that
still meets the 48-hour recovery target with real margin, and keep a much smaller, always-available
search index of what is inside it on fast, queryable storage. The 48-hour window only has to cover
retrieving the specific subset a request actually asks for, not scanning all 20 TB unfiltered, and
that distinction is what makes both cost-effectiveness and the SLA achievable at once.
Structured elaboration
Indexing. At the moment the legal hold is placed and the 20 TB dataset is captured, extract and
index metadata (custodian, meaning the person or system that held the data, plus dates, file types, source system) and, where feasible, full text into a
search index. This index is a much smaller structure than the source data, illustratively on the
order of a low single-digit percentage of the source volume for combined metadata and text
indexing, an order-of-magnitude example, not a measured constant. Counsel then searches this index
to identify the specific subset of the 20 TB actually responsive to a request, rather than ever
needing to search all 20 TB directly. That is what makes the 48-hour SLA realistic: the SLA needs
to cover retrieving a few hundred gigabytes of responsive content, not scanning the full 20 TB from
a cold tier plus doing legal review, all inside the same window.
Storage tiering. Since this is a legal hold, must-preserve, rarely accessed, and held for a
full decade, put the bulk content on a cold or archive tier. Ten years of standard-tier pricing on
data that is essentially never read would be needlessly expensive, and the economics here follow
the same pattern as any rarely accessed, long-retention dataset: infrequent access strongly favors
cold tiering, with a retrieval-fee trade-off, over continuous hot storage. Keep the search index
itself on a persistently available, queryable warm tier, since it must always be searchable on
demand; because it is a small fraction of the total volume, its cost is not the budget's dominant
line item.
Immutability. Apply a compliance-style immutable lock to both the bulk content and, just as
importantly, the index and metadata records describing it, for the full hold duration. A legal
hold's entire purpose is guaranteeing the data, and the record of what it originally was, cannot be
altered, deleted, or subject to spoliation (the legal term for destroying or altering evidence),
whether by an attacker, an automated retention-cleanup job that is not hold-aware, or an insider
trying to make something disappear. Auditability requires every access and every administrative
action on the held data to be logged in a tamper-evident way, who searched, who retrieved, and
when, so a clean chain of custody can be demonstrated if challenged in court.
Validation strategies. On a scheduled interval, for example annually or on any storage
migration, verify checksums of the archived content against the values recorded at ingestion, since
proving nothing has silently corrupted or been altered matters over a span as long as 10 years.
Separately, periodically run an actual end-to-end retrieval drill: pick a sample custodian or date
range, search the index, retrieve from cold storage, and confirm the retrieved content matches the
original and completes inside the 48-hour target, rather than assuming the pipeline still works
after years without a real e-discovery request (e-discovery is the legal process of finding and producing electronically stored data in litigation). This validates the pipeline against the specific
48-hour commitment, not just whether the bytes eventually come back.
Worked example: why cold tiering makes this cost-effective
Using illustrative prices, 20,000 GB on an archive tier at $0.002 per GB-month costs 20,000 times
0.002, or $40 per month, $480 per year, and $4,800 across 10 years. The same 20,000 GB on a
standard tier at $0.02 per GB-month costs $400 per month, $4,800 per year, and $48,000 across 10
years, roughly 10 times more. The search index, at an illustrative 2% of source volume, 400 GB, on
the warm tier at $0.02 per GB-month adds $8 per month, $96 per year, and $960 across 10 years, a
small addition next to the archive tier's $4,800 that does not change the dominant cost driver:
where the 20 TB of bulk content sits.
Unlock Full Question Bank
Get access to all 8 Backup and Disaster Recovery interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.