Storage Systems and Infrastructure Questions
The physical storage substrate that sits beneath databases and other data-intensive systems: disk and volume management, RAID levels and disk-redundancy trade-offs (capacity vs fault tolerance vs rebuild time), and cloud block and instance-store performance characteristics (IOPS, throughput, latency), including how ephemeral instance-store volumes differ from persistent attached ones. Also covers storage-service tiering: hot, warm, cold, and archival lifecycle policies, the retrieval-cost-versus-latency trade-off, and designing the automation that migrates data between tiers and routes reads across them while meeting latency and cost targets. Covers matching a storage configuration's redundancy and performance profile to a system's durability and throughput requirements. This is distinct from how a database engine implements storage internally (write-ahead logs, page layouts, B-tree versus LSM structures). Aimed at engineers who configure and operate the underlying storage hardware.
You're deciding how long to keep data in fast online storage before moving it to cheaper archival storage. What business and technical factors would you weigh in that decision? Walk through an example retention policy for a system that has a 3-year legal retention requirement, where data from the last 90 days is queried frequently and older data is rarely touched.
Sample Answer
Direct answer
Let the retention length come from the legal requirement and let the tier boundary come from the access pattern; they're two separate decisions. Here, keep the full 3 years because that's a hard legal floor regardless of whether anyone ever queries the data again, but only pay for fast storage on the 90 days actually being queried day to day. Everything from day 91 through year 3 moves to a cheaper, slower-to-retrieve tier, since it's "rarely touched," not "never touched," and gets deleted automatically once the 3-year floor is reached (unless a specific record is under legal hold).
Factors to weigh
Business and legal:
- Regulatory retention floor: a strict minimum on how long the data must be kept, independent of query activity. It's a floor, not a target: deleting even one day early is a compliance violation regardless of whether there was any technical reason to keep that data.
- Legal hold: pending or reasonably anticipated litigation can require holding specific records past even that floor; this has to be able to override an automated deletion rule for the records it covers.
- Discovery and audit cost: the harder it is to search a tier, the more expensive it is to respond to a legal discovery request or audit that reaches back past the hot window.
Technical:
- Access pattern: distinguish "rarely touched" from "never touched." Data queried a handful of times a year still needs to come back correctly, just not at hot-tier latency.
- Cost per tier: fast storage costs meaningfully more per GB than an archival tier; tiering exists to capture that difference over the roughly 2 years and 9 months of this policy that isn't the active 90-day window.
- Retrieval latency when it is touched: decide up front what "acceptable" looks like for a rare read from the cold tier (minutes versus hours) instead of discovering it under pressure during an actual request.
- Operational complexity: every additional tier is one more thing to monitor and migrate data through correctly. A 2-tier policy (hot plus cold) is simpler to run than a finely graded multi-tier one, and for a dataset like this, that simplicity is often worth more than squeezing out the last percent of savings.
Worked example
50 GB/day of data, a 3-year legal retention requirement, and a 90-day hot window:
gb_per_day, hot_days, total_days = 50, 90, 365 * 3
hot_gb = gb_per_day * hot_days
cold_gb = gb_per_day * (total_days - hot_days)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_flex = 0.023, 0.0036
cost_all_hot = gb_per_day * total_days * p_std
cost_tiered = hot_gb * p_std + cold_gb * p_flex
print(f"3yr volume: {gb_per_day*total_days:,} GB, hot(90d)={hot_gb:,} GB, cold(remaining {total_days-hot_days}d)={cold_gb:,} GB")
print(f"monthly cost, all in the hot tier: ${cost_all_hot:,.0f}")
print(f"monthly cost, tiered (hot + a cold tier like Glacier Flexible Retrieval): ${cost_tiered:,.0f}")
print(f"savings from tiering: {(1 - cost_tiered/cost_all_hot):.0%}")
Output:
3yr volume: 54,750 GB, hot(90d)=4,500 GB, cold(remaining 1005d)=50,250 GB
monthly cost, all in the hot tier: $1,259
monthly cost, tiered (hot + a cold tier like Glacier Flexible Retrieval): $284
savings from tiering: 77%
Only about 8% of the retained data (4,500 of 54,750 GB) is actually in the hot window at any point, which is why tiering captures most of the cost difference here: the policy pays hot-tier prices for the small slice that's genuinely queried often and archival prices for the much larger slice that's kept for compliance but almost never read.
Trade-offs and pitfalls
- A cold tier is still retained data, not deleted data. The 3-year legal floor applies to it exactly as it applies to the hot data; moving data to a cheaper tier is a cost decision, not a retention decision, and it's easy to conflate the two.
- Automate the day-3-years expiry the same declarative way you automate the day-90 tier move (a lifecycle rule, not a task on someone's calendar). Manual deletion steps are the most common way a retention policy actually fails in practice: data kept well past its required window because nobody ran the manual step.
- If a business process assumes near-instant access to older data, a multi-hour retrieval time on the cold tier will surface as a support or SLA problem the first time someone actually needs a 2-year-old record urgently. Confirm who reads that older data and how urgently before picking the cold tier's retrieval speed, rather than assuming 'rarely touched' also means 'never urgent'.
You have six identical disks attached to a Linux server and must create a software RAID configuration that tolerates two simultaneous disk failures while maximizing usable capacity and reasonable rebuild times. Which RAID level would you choose and why? Provide mdadm commands to create the array with a hot spare, configure automatic assembly at boot, and steps to monitor and replace failed drives.
Sample Answer
Direct answer
Use RAID6 (redundant array of independent disks, level 6): it splits each write into fixed-size pieces called chunks (for example 512 KiB each) and lays one chunk on every disk in a repeating round called a stripe, so no single disk holds a whole file and the disks can be read from or written to in parallel. For every stripe it also computes and writes two independent parity blocks (redundant values mathematically derived from the real data in that stripe; if a disk goes missing, the surviving disks' data plus a parity block is enough to recompute exactly what was on the missing disk) instead of one, so the array keeps serving reads and writes correctly through any two simultaneous disk failures, not just some of them. Usable capacity is (n-2) x disk_size (4 of the 6 disks' worth), and rebuild uses the same dual-parity math a rebuild after a single failure would use, just re-run twice if needed. RAID10 (mirrored pairs, striped) is faster to rebuild and has a cheaper write cost, but with 6 disks in 3 mirror pairs it only survives a second failure if that disk isn't the mirror of the first one that failed, a coin-flip-adjacent guarantee, not a deterministic one, so it does not meet "tolerates two simultaneous failures" as stated.
Comparing the candidates
| RAID level | Disks needed | Usable capacity | Simultaneous-failure guarantee | Random-write cost |
|---|---|---|---|---|
| RAID5 (single parity) | n >= 3 | (n-1) x disk | Survives exactly 1 failure; a 2nd is total data loss | 4 I/Os per write (read data, read parity, write data, write parity) |
| RAID6 (dual parity) | n >= 4 | (n-2) x disk | Survives any 2 simultaneous failures, guaranteed | 6 I/Os per write (read data, read parity1, read parity2, write data, write parity1, write parity2) |
| RAID10 (mirror + stripe) | n >= 4, even | n/2 x disk | Survives 2 failures only if they aren't mirror partners; not guaranteed | 2 I/Os per write (mirror write), fastest of the three |
For 6 disks specifically, RAID10 is faster to rebuild and cheaper per write, but it loses on BOTH of the question's requirements: its usable capacity is only 3 disks' worth (n/2) against RAID6's 4 (n-2), and its dual-failure survival is conditional rather than guaranteed. RAID6 is the only one of the three that satisfies both the "maximize usable capacity" ask and the "tolerates two simultaneous failures" ask at once.
Creating the array with a hot spare
Assume the six data disks are /dev/sdb through /dev/sdg and a seventh disk, /dev/sdh, is the hot spare (a disk that sits idle and automatically takes over as soon as a member disk is marked failed, so the rebuild starts without a human swapping hardware first):
# Wipe any old RAID or filesystem signatures so mdadm doesn't warn/refuse
sudo mdadm --zero-superblock /dev/sd{b,c,d,e,f,g,h}
# Create the RAID6 array: 6 active members + 1 spare, metadata format 1.2
# (the modern default: stores the superblock near the start of the device
# instead of at the very end, so it survives partial-disk mistakes better)
sudo mdadm --create /dev/md0 \
--level=6 \
--raid-devices=6 \
--spare-devices=1 \
--metadata=1.2 \
--chunk=512 \
--name=data0 \
/dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf /dev/sdg /dev/sdh
# Watch the initial parity sync (a RAID6 array is usable immediately but
# runs a background full-stripe consistency pass after creation)
watch cat /proc/mdstat
# Put a filesystem on it once sync is running (safe to do; sync continues underneath)
sudo mkfs.xfs /dev/md0
--metadata=1.2, --name=, --chunk=, --raid-devices=, and --spare-devices= are the actual mdadm --create option names (verified by running mdadm --create with these exact flags against a scratch mdadm 4.3 install: the parser accepted every flag and only rejected the dummy missing placeholder devices I fed it, which is the expected outcome for that reason, not an unrecognized-option error).
Automatic assembly at boot
mdadm --create activates the array immediately, but the kernel has no memory of it across a reboot unless you record it:
# Append the array's UUID-based definition to the mdadm config file
sudo mdadm --detail --scan | sudo tee -a /etc/mdadm/mdadm.conf
# Rebuild the initramfs (the small boot-time filesystem that assembles RAID/LVM
# before the real root filesystem is mounted) so it picks up the new array
sudo update-initramfs -u
# Confirm mdadm.conf now has an ARRAY line naming /dev/md0 by UUID, not by
# /dev/sdX letters (those letters can be reassigned on reboot; the UUID can't)
cat /etc/mdadm/mdadm.conf
On the Debian/Ubuntu package (confirmed by actually installing mdadm and reading its generated /etc/mdadm/mdadm.conf), the file ships with the comment "Run update-initramfs -u after updating this file", which is exactly this step. On an RPM-based distro the equivalent is dracut -f instead of update-initramfs -u; the mdadm --detail --scan >> /etc/mdadm/mdadm.conf step is the same on both.
Monitoring the array
# Point-in-time state: which disks are active/spare/failed, sync %, etc.
sudo mdadm --detail /dev/md0
# Live kernel view, updates continuously
cat /proc/mdstat
# Run mdadm as a background daemon that emails on any array event
# (disk failure, degraded state, rebuild complete, spare activated)
sudo mdadm --monitor --scan --daemonise --mail=root --delay=1800
# Schedule periodic full-stripe consistency checks (most distros ship this
# as a cron/systemd-timer job already, e.g. /etc/cron.d/mdadm on Debian)
echo check | sudo tee /sys/block/md0/md/sync_action
Pair this with SMART monitoring on the physical disks themselves (smartctl -a /dev/sdb, via the smartmontools package) since mdadm only sees I/O errors the disk already reported; SMART's predictive attributes (reallocated sector count, pending sector count) often flag a drive as failing before it actually drops out of the array.
Worked example: capacity, rebuild time, and why not RAID5
Six 4 TiB disks (TiB = tebibyte, 2^40 bytes, the binary unit df/lsblk report in):
def raid6_usable_tib(n_disks, disk_tib):
return (n_disks - 2) * disk_tib
def rebuild_hours(disk_gib, assumed_sustained_mib_s):
capacity_mib = disk_gib * 1024
seconds = capacity_mib / assumed_sustained_mib_s
return seconds / 3600
n, disk_tib = 6, 4
usable = raid6_usable_tib(n, disk_tib)
raw = n * disk_tib
print(f"raw={raw} TiB, usable={usable} TiB, parity overhead={raw-usable} TiB ({(raw-usable)/raw:.0%})")
print(f"rebuild at an assumed 150 MiB/s sustained: {rebuild_hours(4096, 150):.1f} hours")
print(f"rebuild throttled to 30 MiB/s (production I/O sharing the disks): {rebuild_hours(4096, 30):.1f} hours")
Output:
raw=24 TiB, usable=16 TiB, parity overhead=8 TiB (33%)
rebuild at an assumed 150 MiB/s sustained: 7.8 hours
rebuild throttled to 30 MiB/s (production I/O sharing the disks): 38.8 hours
(150 MiB/s and 30 MiB/s are stated assumptions, not a measurement of any specific disk; mdadm's real rebuild speed is governed by /proc/sys/dev/raid/speed_limit_min and speed_limit_max, and actual throughput depends on the disks and how much production I/O is competing for them.)
Why RAID5 is the wrong call for an array this size, computed rather than asserted: rebuilding a failed disk in a RAID5 array means reading every block of every surviving disk to reconstruct the missing one. Consumer and nearline SATA drives commonly publish an unrecoverable-read-error (URE) spec of about 1 bit in 10^14 read:
n, disk_tib = 6, 4
surviving = n - 1
bits_to_read = surviving * disk_tib * 8 * 1024**4
ure_rate = 1e-14
p_hit = 1 - (1 - ure_rate) ** bits_to_read
print(f"bits read during a RAID5 rebuild: {bits_to_read:.3e}")
print(f"P(hit >= 1 URE during that single rebuild): {p_hit:.0%}")
Output:
bits read during a RAID5 rebuild: 1.759e+14
P(hit >= 1 URE during that single rebuild): 83%
An 83% chance of hitting an unreadable sector mid-rebuild, on a single-parity array that has zero spare parity left once one disk is already gone, is the concrete reason "6 disks, large capacity, survive 2 failures" rules out RAID5 even before the "two simultaneous failures" requirement is read literally. RAID6's second parity block absorbs exactly that kind of single bad sector during a rebuild.
Replacing a failed drive
# mdadm already marked the failed member 'F' in --detail output and kicked
# the hot spare into a rebuild automatically; confirm that happened:
sudo mdadm --detail /dev/md0 | grep -E "State|faulty|spare|rebuild"
# If the disk hasn't been auto-failed yet (e.g. it's throwing errors but
# hasn't hard-failed), fail and remove it explicitly before pulling it:
sudo mdadm --manage /dev/md0 --fail /dev/sdd
sudo mdadm --manage /dev/md0 --remove /dev/sdd
# Physically swap the drive, then partition/prepare the replacement the
# same way the originals were, and add it back as the new spare:
sudo mdadm --manage /dev/md0 --add /dev/sdd
# Watch the rebuild (or, if a spare already absorbed the first failure,
# this --add just refills the spare pool back to 1):
watch cat /proc/mdstat
--fail, --remove, and --add are real mdadm --manage option names (confirmed against mdadm --manage --help output on the same install).
Trade-offs and pitfalls
- RAID is not backup: it protects against disk hardware failure, not against accidental deletion, ransomware, or a bad
rm -rf, all of which replicate instantly across every mirror/parity member. Keep an actual backup regardless of RAID level. - RAID6's dual-parity write cost (6 I/Os per small random write versus RAID10's 2) makes it a poor fit for a workload that demands high write IOPS (input/output operations per second, the count of individual read or write operations a storage system can complete each second); it earns its keep on capacity-and-durability-first workloads (bulk storage, backups-of-backups, archival-adjacent data), which matches "maximizing usable capacity" in the question.
- The rebuild window is the danger zone: for as long as a rebuild runs, the array is one more failure away from data loss (RAID6 degraded to single-parity mid-rebuild, or fully exposed if a 2nd failure lands before the 1st rebuild finishes). A hot spare shrinks that window by starting the rebuild automatically instead of waiting on a human to notice and swap hardware.
--chunk=512(512 KiB) is a reasonable general-purpose default; a workload dominated by large sequential I/O (video, backups) benefits from a larger chunk, one dominated by small random I/O from a smaller one. Get this from testing your actual workload, not by guessing.- Never trust
/dev/sdXletters as stable identifiers in configuration; they can be reassigned on reboot depending on device enumeration order. That's exactly whymdadm --detail --scanwrites the array definition by UUID.
Describe cloud block storage (for example AWS EBS or Azure Managed Disks). What are its key performance characteristics (IOPS, throughput, latency), and what are some typical use cases, such as scratch space for a compute cluster, a database's backing disk, or a local cache? Compare ephemeral instance-store volumes to persistent attached block volumes, and explain what that difference means for fault tolerance and for any workload that periodically saves its progress to disk.
Sample Answer
Direct answer
Cloud block storage (Amazon EBS, Azure Managed Disks) is a virtual disk your compute instance attaches to over the provider's internal network and addresses like any other block device (it can be partitioned and formatted with a normal filesystem), but the actual bytes live on separate, redundant storage infrastructure rather than on physical disks bolted to the host. Its three headline performance characteristics are IOPS (input/output operations per second: how many individual reads or writes it can service each second), throughput (megabytes per second of sustained data movement), and latency (how long a single operation takes to complete). Ephemeral instance-store volumes flip this trade: they're physically attached to the host, faster and cheaper because there's no network hop and no separate durability layer, but their data dies with the instance.
IOPS, throughput, and latency in practice
Amazon EBS's gp3 volume (the current general-purpose default) gives a concrete feel for the numbers: a baseline of 3,000 IOPS and 125 MiB/s throughput included at no extra cost, provisionable up to a maximum of 80,000 IOPS and 2,000 MiB/s for an additional fee, at single-digit-millisecond latency (verified against AWS's current EBS documentation). For workloads needing tighter, more predictable latency, the Provisioned IOPS io2 Block Express tier goes up to 256,000 IOPS and 4,000 MiB/s, with an average latency under 500 microseconds for 16 KiB operations. Azure's comparable Premium SSD v2 lands on almost the same numbers: 3,000 IOPS / 125 MB/s baseline, up to 80,000 IOPS / 2,000 MB/s provisioned. The pattern to internalize: on network-attached block storage, IOPS and throughput are independently provisionable dials, not fixed properties of "the disk," and you pay more as you turn either dial up.
Typical use cases
| Use case | What matters most | Why |
|---|---|---|
| Scratch space for a compute cluster (shuffle/spill files, intermediate build artifacts) | Raw IOPS and throughput, not durability | The data is recomputable; paying for cross-host replication is wasted cost. Instance store often wins here. |
| A database's backing disk | Consistent low latency and durability that survives the instance disappearing | A database can't tolerate losing committed data because the compute instance was replaced; persistent attached block storage (EBS/Managed Disks) is the right fit, sized for the database's actual IOPS profile. |
| A local cache (e.g. an in-memory-adjacent read cache spilled to disk) | Low latency, moderate durability requirements | Losing the cache on instance replacement is a performance hit, not a correctness problem, so instance store is often an acceptable, cheaper choice. |
Ephemeral instance store versus persistent attached volumes
- Instance store: physically attached to the specific host your instance is running on. Data persists across a plain instance reboot, but does not persist through a stop, hibernate, or terminate, and does not survive the underlying hardware failing, because in all of those cases you either move to different physical hardware or the current hardware is gone (confirmed against AWS's current instance-store documentation). You cannot detach it and reattach it elsewhere; it exists only for the lifetime of that specific instance's association with that specific host.
- Persistent attached block storage (EBS, Managed Disks): the data lives on storage infrastructure separate from any single compute host, replicated by the storage service itself for durability. You can detach the volume and reattach it to a different instance, and the volume survives the original instance being stopped, terminated, or replaced.
What that difference means for fault tolerance and for checkpointing workloads
Fault tolerance here is really a question of what failure domain the data shares. Instance store shares a failure domain with the compute instance: if the instance is replaced (autoscaling churn, a spot/preemptible reclaim, a host hardware fault), the data is gone, full stop. A persistent attached volume's failure domain is the storage service's own redundancy, independent of any one compute instance, so an instance failure doesn't take the volume's data with it.
For a workload that periodically checkpoints its progress, this sets a hard requirement: the checkpoint itself must land on storage that survives independently of the instance that wrote it, which means a persistent attached volume (or object storage) for the checkpoint, even if the working data between checkpoints lives on faster, cheaper instance store. The checkpoint interval then directly bounds your worst-case data loss: a job checkpointing every 10 minutes to EBS can lose at most the last 10 minutes of work if the instance disappears; a job that (mistakenly) checkpoints to instance store loses everything the instant the instance is stopped or replaced, regardless of how recently it checkpointed, because the checkpoint target itself is gone too.
Worked example: sizing a gp3 volume for a database's IOPS profile
A database issuing roughly 5,000 random 16 KiB reads/sec needs both a minimum IOPS and a minimum throughput; compute both instead of guessing at the volume size:
def gp3_min_size_for_iops(iops, ratio_iops_per_gib=500):
return max(1, -(-iops // ratio_iops_per_gib)) # ceiling division
target_iops = 5000
io_size_kib = 16
required_throughput_mib_s = target_iops * io_size_kib / 1024
min_size_gib = gp3_min_size_for_iops(target_iops)
print(f"required throughput for {target_iops} IOPS at {io_size_kib} KiB each: {required_throughput_mib_s:.1f} MiB/s")
print(f"gp3 baseline: 3,000 IOPS / 125 MiB/s included; {target_iops} IOPS needs provisioning above baseline")
print(f"minimum volume size to provision {target_iops} IOPS at gp3's 500 IOPS/GiB ratio: {min_size_gib} GiB")
Output:
required throughput for 5000 IOPS at 16 KiB each: 78.1 MiB/s
gp3 baseline: 3,000 IOPS / 125 MiB/s included; 5000 IOPS needs provisioning above baseline
minimum volume size to provision 5000 IOPS at gp3's 500 IOPS/GiB ratio: 10 GiB
The needed throughput (78.1 MiB/s) stays comfortably under gp3's 125 MiB/s free baseline, so the only dial to turn up is IOPS, and the 500:1 ratio means even a 10 GiB volume can legally carry 5,000 provisioned IOPS; in practice you'd size the volume for actual data capacity first and confirm the IOPS ratio is satisfied, not the other way around.
Trade-offs and pitfalls
- Don't put anything you can't afford to lose on instance store purely because it's faster; verify what "faster" is actually buying you against what "the data disappears on stop/terminate/hardware failure" costs you if you're wrong about the workload being truly recomputable or disposable.
- gp2 (the older general-purpose EBS type) ties IOPS to volume size and relies on burst credits (a banked allowance of extra I/O: the volume accrues credits while running below its baseline IOPS and spends them to briefly exceed that baseline; once the bank runs out, performance drops back down to baseline until credits build back up); gp3 decouples IOPS and throughput from size entirely, which is why gp3 is the current default recommendation, not gp2.
- Network-attached block storage has an inherent latency floor that pure local NVMe (a high-speed local solid-state storage interface) instance store doesn't: even at gp3/io2's best-case latency, there's a network hop involved that local storage skips. For workloads sensitive to the last fraction of a millisecond (not just IOPS), that architectural difference matters more than any provisioned-IOPS number.
- A checkpoint strategy is only as good as its interval: computing the actual maximum tolerable data loss (checkpoint interval times the rate of work) and comparing it against the business requirement is a five-minute calculation that catches a checkpoint-too-infrequent design before it ships, not after an incident.
Given a fixed budget, design a storage tiering system that automatically migrates data between hot, warm, and cold tiers (for example local NVMe, a warm SSD cluster, and an object store) to balance latency and cost. What criteria would you use to decide when a piece of data should move between tiers, how would you architect the automation that carries out those migrations safely, how would reads be routed so a query transparently spans whichever tiers hold the data it needs, and how would you measure and enforce latency and cost SLOs across the whole system?
Sample Answer
Direct answer
Build this as a closed control loop, not a one-time placement decision: a stats collector measures per-partition access frequency and latency, a scoring engine turns those stats into a hot/warm/cold placement decision under an explicit cost budget, a migration executor carries out moves safely (copy first, flip a pointer, never move-then-copy), and a catalog-driven query router lets a single query transparently read partitions that are scattered across all three tiers by pruning to only the tiers and files it actually needs. An SLO (service-level objective: a measurable target for how the system should perform, such as a latency ceiling or a cost cap) monitor watches latency and cost against their targets and feeds back into the scorer, so the loop self-corrects instead of drifting.
Architecture
flowchart LR
NVMe[(Hot: local NVMe)]
SSD[(Warm: SSD cluster)]
Obj[(Cold: object store)]
Stats[Access and latency stats collector]
Scorer[Tier scoring policy engine]
Executor[Migration executor]
Catalog[(Metadata catalog: partition to tier plus stats)]
Router[Query router]
SLO[SLO monitor and budget enforcer]
NVMe --> Stats
SSD --> Stats
Obj --> Stats
Stats --> Scorer
SLO --> Scorer
Scorer --> Executor
Executor --> NVMe
Executor --> SSD
Executor --> Obj
Executor --> Catalog
Router --> Catalog
Catalog --> Router
Router --> NVMe
Router --> SSD
Router --> Obj
Every box is a real, separately deployable component: NVMe (the fastest commonly available local solid-state storage interface, hence the natural choice for the hot tier) sits as one edge of the loop, the object store as the other, and the catalog is the one piece every other component depends on, so it needs to be a small, highly available, always-hot store in its own right (a metadata service, not a file dropped in the object store), the same principle a Hive- or Glue-style metastore (a separate always-on service, used in big-data platforms, that tracks which files make up which table) or an open-table-format manifest such as Iceberg or Delta (a structured metadata file that plays the same role: a durable index of a table's data files) follows.
Tier-placement criteria and scoring
Score each partition on a blend of recency and frequency rather than either alone, since a partition read once an hour ago and a partition read a thousand times an hour ago are not equally "hot":
def tier_score(reads_last_24h, hours_since_last_read, recency_half_life_hours=6):
recency_weight = 0.5 ** (hours_since_last_read / recency_half_life_hours)
return reads_last_24h * recency_weight
# a partition read 200 times in the last day, but not touched in the last 3 hours
print(f"score: {tier_score(200, 3):.1f}")
# a partition read only 10 times, but 1 of those reads was 20 minutes ago
print(f"score: {tier_score(10, 0.33):.1f}")
Output:
score: 141.4
score: 9.6
Rank partitions by this score each cycle, promote the top scorers toward hot (subject to the budget check below) and demote the ones that have fallen furthest since the last cycle. The exact scoring formula matters less than the principle: base it on the same signal (access pattern) the SLOs are measured against, and re-evaluate it on a fixed cadence rather than only reactively.
Migration mechanics (why safety, not just correctness, is the hard part)
The failure mode to design against is a reader hitting a partition mid-move and getting either stale or missing data. The safe sequence:
- Copy the partition's data to the destination tier. The source is untouched and still fully readable throughout.
- Verify the copy (checksum comparison against the source).
- Atomically update the catalog's tier pointer for that partition to the new location. This is the single moment the migration becomes visible to readers; it's a metadata write, not a data write, so it's fast and can be done as one transaction.
- Only after the pointer flip is confirmed, delete the data from the source tier.
Never move-then-copy (delete first, copy second): any failure between those two steps loses the data outright. Never expose a partial write of the copy at the destination: readers must only ever see either the old location (fully intact) or the new one (fully intact and verified), never a half-written destination.
Cross-tier query routing (the mechanism, not just the claim)
A query spans tiers by pruning at the catalog, not by asking every tier "do you have this." The catalog stores, per partition, its current tier plus the same partition statistics (min/max key ranges, row counts, column statistics) a query planner (the part of a database or query engine that decides how to execute a query) would use for predicate pushdown: skipping, that is pruning, whole files or partitions that the query's filter conditions cannot possibly match, using those stored stats, instead of reading every file to check. This is exactly the way a Hive- or Glue-style metastore or an Iceberg/Delta manifest does it today. The query router's job:
- Take the query's predicates (a date range, a key range).
- Look up the catalog to get the list of partitions that match those predicates, along with each matching partition's current tier.
- Group the matching partitions by tier and dispatch a scan to each tier's own read path in parallel (NVMe local reads, SSD-cluster reads, object-store GET/scan requests).
- Merge the results, exactly as a query engine already merges results from multiple files today; spanning tiers is the same fan-out-and-merge pattern, just with three different storage backends instead of one.
This is why the catalog carrying partition statistics matters: without it, "transparently spans tiers" degrades into scanning every tier for every query, defeating the entire cost benefit of the cold tier being cheap because it's rarely touched.
Measuring and enforcing latency and cost SLOs
- Latency SLO: track p95/p99 read latency per tier (95th/99th percentile: the latency value that 95%/99% of reads finish faster than) against a target, e.g. "p99 for hot-tier reads under 5 ms." A tier whose p99 is breaching its target is a signal to the scorer to be more aggressive about keeping genuinely hot data on NVMe rather than a signal to add more NVMe capacity blindly; check whether the SLO breach is a placement problem (wrong data is hot) before treating it as a capacity problem.
- Cost SLO (budget enforcement): the scorer's promotion decisions are budget-checked, not just score-ranked. If promoting everything that scored above threshold would exceed the monthly budget, promote only the highest-scoring subset that fits within remaining headroom, and leave the rest queued for the next cycle:
total_tib, hot_frac, warm_frac = 20, 0.05, 0.25
hot_tib, warm_tib = total_tib*hot_frac, total_tib*warm_frac
cold_tib = total_tib*(1-hot_frac-warm_frac)
# Illustrative unit costs ($/TiB-month), not a live vendor quote
c_nvme, c_ssd, c_obj = 90, 25, 3
budget = 300
cost = hot_tib*c_nvme + warm_tib*c_ssd + cold_tib*c_obj
headroom = budget - cost
requested_promotion_tib = 1.0
marginal_cost_per_tib = c_nvme - c_obj
affordable_tib = headroom / marginal_cost_per_tib
print(f"current cost ${cost:.0f} of ${budget} budget, headroom ${headroom:.0f}")
print(f"scorer flags {requested_promotion_tib:.1f} TiB as hot-worthy; "
f"budget only affords promoting {affordable_tib:.2f} TiB")
Output:
current cost $257 of $300 budget, headroom $43
scorer flags 1.0 TiB as hot-worthy; budget only affords promoting 0.49 TiB
The executor promotes the top-scoring 0.49 TiB and leaves the remaining 0.51 TiB queued rather than silently blowing the budget by promoting the full 1.0 TiB; the next cycle re-evaluates as older hot data ages out and frees headroom. This is the concrete difference between "measuring" an SLO (a dashboard number) and "enforcing" one (a decision the system actually makes because of that number).
Trade-offs and pitfalls
- A scoring cycle that runs too infrequently (say, once a day) means the system reacts to yesterday's access pattern, not today's; too frequently, and the migration churn itself (constant copying) eats the budget it's supposed to protect. Tune the cadence against how quickly your access patterns actually shift, not by habit.
- Budget-constrained promotion queues up demand that never gets served if the workload's hot working set has genuinely outgrown the NVMe budget; that's a signal to revisit the budget or the hot-tier sizing, not a problem the scorer can solve by prioritizing harder.
- Skipping the copy-verify-flip-delete sequence in favor of a faster move-in-place is exactly the shortcut that turns an ordinary migration into a data-loss incident the first time it's interrupted mid-move (a crash, a network partition).
- If the catalog and the query router live in different failure domains, a catalog update that succeeds but doesn't propagate to the router in time produces a stale read plan; treat catalog reads on the query path as needing to be current, not eventually consistent (a system that only promises all readers will see the same data EVENTUALLY, after some unspecified lag, rather than immediately), or the "transparently spans tiers" property silently breaks for a window after every migration.
Design an archival policy that moves older data from a warm or hot storage tier into a cheaper cold or archival tier, without breaking jobs that occasionally still need to read snapshots of that older data. Cover the lifecycle transition rules you would set, how you would catalog or index archived data so it can still be found, the restore workflow and its expected latency, the trade-off between retrieval cost and access speed, and how you would avoid unexpected restore failures or surprise costs when older data actually gets requested back.
Sample Answer
Direct answer
Build the policy around three separate pieces that people usually collapse into one: (1) transition rules that move data between tiers automatically based on age or access pattern, (2) a catalog that always knows which tier and which retrieval class a given piece of data currently lives in, independent of where it physically sits, and (3) a restore path whose retrieval-speed class is chosen up front, at tiering time, to match the worst-case restore-time SLA (service-level agreement: a promised or contractually required target, here the maximum acceptable time to get requested data back) that data might ever need, not improvised after someone requests it back. Verify the whole thing works with real, scheduled restore drills rather than trusting the design on paper.
Lifecycle transition rules
- Trigger type: age-based (simplest: "move to cold after 365 days," configured as a native bucket/table lifecycle policy, e.g. S3 Lifecycle rules or Azure Blob lifecycle management, so the storage service itself enforces it rather than relying on an application job someone eventually forgets to run) or access-frequency-based (more precise: track actual read frequency and demote data that has genuinely gone cold, e.g. S3 Intelligent-Tiering's automatic monitoring). Age-based is the right default; add frequency-based only where access patterns don't correlate cleanly with age.
- Event-based overrides: some data should move immediately on a business event (a case closes, a record is finalized) rather than waiting for an age threshold; model this as an explicit catalog write that fires the transition, not a special case bolted onto the age rule.
- Prefer declarative, storage-service-native lifecycle rules over custom migration code wherever the storage system supports it: a misconfigured lifecycle rule fails loudly (data doesn't move, and you can see that in the catalog's per-tier byte counts); a custom migration job can fail silently and nobody notices until a restore request comes in for data that was never actually archived.
Cataloging and indexing archived data
The moment data leaves the hot, natively-queryable store, you need a durable catalog that maps a logical identifier (a file path, a partition key, a record range) to where it currently lives: tier, storage/retrieval class, the physical object key, its size, a checksum recorded at archive time, and when it becomes eligible for expiry. Without this, "find it" degenerates into enumerating every tier by hand.
- Implement it as a small, always-hot, highly available table (a Postgres table, a DynamoDB table, or, if the archived data is itself tabular, an open-table-format manifest such as Apache Iceberg or Delta Lake: a structured metadata file format used in data-lake/analytics stacks that tracks a table's underlying data files, schema, and partitions), never as something that itself gets archived: the catalog is on the critical path of every restore, so it must always be instantly readable.
- Update the catalog as part of the same operation that performs the tier transition (or reconcile it against the transition job reliably), not as an unrelated best-effort side effect; a catalog that is out of sync with reality is worse than no catalog, because it answers confidently and wrong at exactly the moment (an audit, an incident) when correctness matters most.
Restore workflow and retrieval-speed class
Cold-tier restores aren't uniform: choose among near-instant, expedited, or bulk retrieval classes based on the restore-time SLA the specific data actually needs, decided when the data is tiered, not guessed at restore time.
- Near-instant (e.g. S3 Glacier Instant Retrieval): millisecond-latency GETs, priced closer to a warm tier. Use it for cold data that is rarely read but must never make a caller wait.
- Expedited (e.g. S3 Glacier Flexible Retrieval's Expedited option): typically 1 to 5 minutes, at a premium per-request and per-GB retrieval price. Use it for data bound by a tight, hard restore-time SLA that near-instant pricing can't justify for the whole dataset.
- Bulk: the cheapest per-GB retrieval option, typically hours (and up to about two days on the deepest archival classes). This is the right default for the vast majority of archived data, which is rarely if ever requested back and has no hard restore-time bound.
The mistake this guards against: assuming any cold tier can be rehydrated "fast enough" and only discovering the real restore-time bound during an actual audit or incident, when it's too late to move the data into a faster-retrieval class.
After a restore completes, verify it, don't just trust that the read call returned success: recompute a checksum on the restored object and compare it against the checksum recorded in the catalog at archive time.
Avoiding restore failures and surprise costs
- Restore-drill verification: schedule periodic real test restores, not synthetic health checks. Pick a sample of archived partitions, run the actual restore workflow end to end, verify row counts and checksums against the catalog's recorded values, and log the observed restore duration as evidence the assigned retrieval class is genuinely meeting its SLA target. This is what catches a wrong retrieval-class assignment before a real, time-boxed request does.
- Automated compliance reporting: a recurring report or dashboard that reads the catalog and surfaces per-tier byte counts (so data that silently failed to migrate shows up as an unexpected size anomaly in the wrong tier), upcoming retention expirations (so purges happen exactly on schedule, neither late, which is a compliance finding, nor early, which is evidence destruction), and legal-hold status per record.
- Surprise-cost guardrail: expedited and near-instant retrieval are priced at a premium per GB and per request; a misconfigured or accidental bulk job that restores a whole cold partition at expedited speed instead of bulk can spend a meaningful fraction of a month's storage budget in one run. Enforce a retrieval-class allowlist per caller (only the narrow incident-response path is allowed to request expedited; scheduled batch jobs are pinned to bulk) and alert on retrieval-request volume.
- Legal hold and lifecycle expiry are two independent gates on deletion, and both must block it: a lifecycle rule alone will happily delete data that is under active litigation hold unless the hold is checked first.
Worked example: 7-year audit-log retention with a 4-hour rehydrate requirement
A system ingests 200 GB/day of audit logs, subject to a 7-year regulatory retention requirement, and an audit process that can demand a named partition rehydrated within 4 hours.
gb_per_day, years, hot_days = 200, 7, 30
warm_days = 365 - hot_days
total_gb = gb_per_day * 365 * years
hot_gb = gb_per_day * hot_days
warm_gb = gb_per_day * warm_days
cold_gb = gb_per_day * 365 * (years - 1)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_ia, p_deep, p_flex = 0.023, 0.0125, 0.00099, 0.0036
sla_frac = 0.10 # share of the cold tier subject to the 4h rehydrate SLA
cold_sla_gb = cold_gb * sla_frac
cold_bulk_gb = cold_gb * (1 - sla_frac)
cost_split_cold = cold_sla_gb * p_flex + cold_bulk_gb * p_deep
total_monthly = hot_gb * p_std + warm_gb * p_ia + cost_split_cold
baseline_all_std = total_gb * p_std
print(f"total 7yr volume: {total_gb:,} GB")
print(f"tiers: hot={hot_gb:,} GB, warm={warm_gb:,} GB, cold={cold_gb:,} GB "
f"(of which {cold_sla_gb:,.0f} GB is SLA-bound)")
print(f"monthly cost, tiered with the SLA split: ${total_monthly:,.0f}")
print(f"monthly cost, all-Standard baseline: ${baseline_all_std:,.0f}")
print(f"savings: {(1 - total_monthly/baseline_all_std):.0%}")
Output:
total 7yr volume: 511,000 GB
tiers: hot=6,000 GB, warm=67,000 GB, cold=438,000 GB (of which 43,800 GB is SLA-bound)
monthly cost, tiered with the SLA split: $1,523
monthly cost, all-Standard baseline: $11,753
savings: 87%
The concrete decision the SLA forces: S3 Glacier Deep Archive's retrieval options are Standard (9-12 hours) and Bulk (up to 48 hours); neither meets a 4-hour bound. The 10% of cold data actually subject to that bound has to live in a faster-retrieval class (here, Glacier Flexible Retrieval, whose Expedited option is typically 1-5 minutes) instead, at roughly 3.6x Deep Archive's per-GB price, an explicit, budgeted trade for meeting the SLA rather than an assumption that gets discovered wrong during a real audit. The restore drill for this slice runs monthly: pick one random SLA-bound partition, run an Expedited restore, record the wall-clock time to completion, verify its row count and checksum against the catalog, and post the result to the compliance dashboard, so drift toward missing the 4-hour bound (for instance during a regional capacity crunch) is caught by the drill instead of by an auditor.
Trade-offs and pitfalls
- Defaulting everything to near-instant or expedited retrieval defeats the purpose of tiering: most archived data is never requested back, and the savings tiering exists to capture come from most of it staying in the cheapest, slowest-to-restore class. Reserve the premium retrieval class narrowly, for the specific data actually bound by a hard SLA, the way the worked example does with its 10% split.
- A catalog is only as trustworthy as its consistency with the actual data movement; treat catalog updates as part of the transition operation, not an afterthought.
- "The restore call returned success" and "the restored data is correct" are different claims; skip the checksum/row-count verification and a silently truncated or corrupted restore is indistinguishable from a good one until someone reads the data and notices.
- Don't wait for a real audit to discover a retrieval-class mismatch. A restore drill that only checks "did the API call succeed" without timing it against the SLA target isn't actually validating the thing that matters.
That is every published Storage Systems and Infrastructure question for Systems Engineer so far. Browse the other topics in this category, or practice this one interactively.