Storage Systems and Infrastructure Questions
The physical storage substrate that sits beneath databases and other data-intensive systems: disk and volume management, RAID levels and disk-redundancy trade-offs (capacity vs fault tolerance vs rebuild time), and cloud block and instance-store performance characteristics (IOPS, throughput, latency), including how ephemeral instance-store volumes differ from persistent attached ones. Also covers storage-service tiering: hot, warm, cold, and archival lifecycle policies, the retrieval-cost-versus-latency trade-off, and designing the automation that migrates data between tiers and routes reads across them while meeting latency and cost targets. Covers matching a storage configuration's redundancy and performance profile to a system's durability and throughput requirements. This is distinct from how a database engine implements storage internally (write-ahead logs, page layouts, B-tree versus LSM structures). Aimed at engineers who configure and operate the underlying storage hardware.
You're deciding how long to keep data in fast online storage before moving it to cheaper archival storage. What business and technical factors would you weigh in that decision? Walk through an example retention policy for a system that has a 3-year legal retention requirement, where data from the last 90 days is queried frequently and older data is rarely touched.
Sample Answer
Direct answer
Let the retention length come from the legal requirement and let the tier boundary come from the access pattern; they're two separate decisions. Here, keep the full 3 years because that's a hard legal floor regardless of whether anyone ever queries the data again, but only pay for fast storage on the 90 days actually being queried day to day. Everything from day 91 through year 3 moves to a cheaper, slower-to-retrieve tier, since it's "rarely touched," not "never touched," and gets deleted automatically once the 3-year floor is reached (unless a specific record is under legal hold).
Factors to weigh
Business and legal:
- Regulatory retention floor: a strict minimum on how long the data must be kept, independent of query activity. It's a floor, not a target: deleting even one day early is a compliance violation regardless of whether there was any technical reason to keep that data.
- Legal hold: pending or reasonably anticipated litigation can require holding specific records past even that floor; this has to be able to override an automated deletion rule for the records it covers.
- Discovery and audit cost: the harder it is to search a tier, the more expensive it is to respond to a legal discovery request or audit that reaches back past the hot window.
Technical:
- Access pattern: distinguish "rarely touched" from "never touched." Data queried a handful of times a year still needs to come back correctly, just not at hot-tier latency.
- Cost per tier: fast storage costs meaningfully more per GB than an archival tier; tiering exists to capture that difference over the roughly 2 years and 9 months of this policy that isn't the active 90-day window.
- Retrieval latency when it is touched: decide up front what "acceptable" looks like for a rare read from the cold tier (minutes versus hours) instead of discovering it under pressure during an actual request.
- Operational complexity: every additional tier is one more thing to monitor and migrate data through correctly. A 2-tier policy (hot plus cold) is simpler to run than a finely graded multi-tier one, and for a dataset like this, that simplicity is often worth more than squeezing out the last percent of savings.
Worked example
50 GB/day of data, a 3-year legal retention requirement, and a 90-day hot window:
gb_per_day, hot_days, total_days = 50, 90, 365 * 3
hot_gb = gb_per_day * hot_days
cold_gb = gb_per_day * (total_days - hot_days)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_flex = 0.023, 0.0036
cost_all_hot = gb_per_day * total_days * p_std
cost_tiered = hot_gb * p_std + cold_gb * p_flex
print(f"3yr volume: {gb_per_day*total_days:,} GB, hot(90d)={hot_gb:,} GB, cold(remaining {total_days-hot_days}d)={cold_gb:,} GB")
print(f"monthly cost, all in the hot tier: ${cost_all_hot:,.0f}")
print(f"monthly cost, tiered (hot + a cold tier like Glacier Flexible Retrieval): ${cost_tiered:,.0f}")
print(f"savings from tiering: {(1 - cost_tiered/cost_all_hot):.0%}")
Output:
3yr volume: 54,750 GB, hot(90d)=4,500 GB, cold(remaining 1005d)=50,250 GB
monthly cost, all in the hot tier: $1,259
monthly cost, tiered (hot + a cold tier like Glacier Flexible Retrieval): $284
savings from tiering: 77%
Only about 8% of the retained data (4,500 of 54,750 GB) is actually in the hot window at any point, which is why tiering captures most of the cost difference here: the policy pays hot-tier prices for the small slice that's genuinely queried often and archival prices for the much larger slice that's kept for compliance but almost never read.
Trade-offs and pitfalls
- A cold tier is still retained data, not deleted data. The 3-year legal floor applies to it exactly as it applies to the hot data; moving data to a cheaper tier is a cost decision, not a retention decision, and it's easy to conflate the two.
- Automate the day-3-years expiry the same declarative way you automate the day-90 tier move (a lifecycle rule, not a task on someone's calendar). Manual deletion steps are the most common way a retention policy actually fails in practice: data kept well past its required window because nobody ran the manual step.
- If a business process assumes near-instant access to older data, a multi-hour retrieval time on the cold tier will surface as a support or SLA problem the first time someone actually needs a 2-year-old record urgently. Confirm who reads that older data and how urgently before picking the cold tier's retrieval speed, rather than assuming 'rarely touched' also means 'never urgent'.
Describe cloud block storage (for example AWS EBS or Azure Managed Disks). What are its key performance characteristics (IOPS, throughput, latency), and what are some typical use cases, such as scratch space for a compute cluster, a database's backing disk, or a local cache? Compare ephemeral instance-store volumes to persistent attached block volumes, and explain what that difference means for fault tolerance and for any workload that periodically saves its progress to disk.
Sample Answer
Direct answer
Cloud block storage (Amazon EBS, Azure Managed Disks) is a virtual disk your compute instance attaches to over the provider's internal network and addresses like any other block device (it can be partitioned and formatted with a normal filesystem), but the actual bytes live on separate, redundant storage infrastructure rather than on physical disks bolted to the host. Its three headline performance characteristics are IOPS (input/output operations per second: how many individual reads or writes it can service each second), throughput (megabytes per second of sustained data movement), and latency (how long a single operation takes to complete). Ephemeral instance-store volumes flip this trade: they're physically attached to the host, faster and cheaper because there's no network hop and no separate durability layer, but their data dies with the instance.
IOPS, throughput, and latency in practice
Amazon EBS's gp3 volume (the current general-purpose default) gives a concrete feel for the numbers: a baseline of 3,000 IOPS and 125 MiB/s throughput included at no extra cost, provisionable up to a maximum of 80,000 IOPS and 2,000 MiB/s for an additional fee, at single-digit-millisecond latency (verified against AWS's current EBS documentation). For workloads needing tighter, more predictable latency, the Provisioned IOPS io2 Block Express tier goes up to 256,000 IOPS and 4,000 MiB/s, with an average latency under 500 microseconds for 16 KiB operations. Azure's comparable Premium SSD v2 lands on almost the same numbers: 3,000 IOPS / 125 MB/s baseline, up to 80,000 IOPS / 2,000 MB/s provisioned. The pattern to internalize: on network-attached block storage, IOPS and throughput are independently provisionable dials, not fixed properties of "the disk," and you pay more as you turn either dial up.
Typical use cases
| Use case | What matters most | Why |
|---|---|---|
| Scratch space for a compute cluster (shuffle/spill files, intermediate build artifacts) | Raw IOPS and throughput, not durability | The data is recomputable; paying for cross-host replication is wasted cost. Instance store often wins here. |
| A database's backing disk | Consistent low latency and durability that survives the instance disappearing | A database can't tolerate losing committed data because the compute instance was replaced; persistent attached block storage (EBS/Managed Disks) is the right fit, sized for the database's actual IOPS profile. |
| A local cache (e.g. an in-memory-adjacent read cache spilled to disk) | Low latency, moderate durability requirements | Losing the cache on instance replacement is a performance hit, not a correctness problem, so instance store is often an acceptable, cheaper choice. |
Ephemeral instance store versus persistent attached volumes
- Instance store: physically attached to the specific host your instance is running on. Data persists across a plain instance reboot, but does not persist through a stop, hibernate, or terminate, and does not survive the underlying hardware failing, because in all of those cases you either move to different physical hardware or the current hardware is gone (confirmed against AWS's current instance-store documentation). You cannot detach it and reattach it elsewhere; it exists only for the lifetime of that specific instance's association with that specific host.
- Persistent attached block storage (EBS, Managed Disks): the data lives on storage infrastructure separate from any single compute host, replicated by the storage service itself for durability. You can detach the volume and reattach it to a different instance, and the volume survives the original instance being stopped, terminated, or replaced.
What that difference means for fault tolerance and for checkpointing workloads
Fault tolerance here is really a question of what failure domain the data shares. Instance store shares a failure domain with the compute instance: if the instance is replaced (autoscaling churn, a spot/preemptible reclaim, a host hardware fault), the data is gone, full stop. A persistent attached volume's failure domain is the storage service's own redundancy, independent of any one compute instance, so an instance failure doesn't take the volume's data with it.
For a workload that periodically checkpoints its progress, this sets a hard requirement: the checkpoint itself must land on storage that survives independently of the instance that wrote it, which means a persistent attached volume (or object storage) for the checkpoint, even if the working data between checkpoints lives on faster, cheaper instance store. The checkpoint interval then directly bounds your worst-case data loss: a job checkpointing every 10 minutes to EBS can lose at most the last 10 minutes of work if the instance disappears; a job that (mistakenly) checkpoints to instance store loses everything the instant the instance is stopped or replaced, regardless of how recently it checkpointed, because the checkpoint target itself is gone too.
Worked example: sizing a gp3 volume for a database's IOPS profile
A database issuing roughly 5,000 random 16 KiB reads/sec needs both a minimum IOPS and a minimum throughput; compute both instead of guessing at the volume size:
def gp3_min_size_for_iops(iops, ratio_iops_per_gib=500):
return max(1, -(-iops // ratio_iops_per_gib)) # ceiling division
target_iops = 5000
io_size_kib = 16
required_throughput_mib_s = target_iops * io_size_kib / 1024
min_size_gib = gp3_min_size_for_iops(target_iops)
print(f"required throughput for {target_iops} IOPS at {io_size_kib} KiB each: {required_throughput_mib_s:.1f} MiB/s")
print(f"gp3 baseline: 3,000 IOPS / 125 MiB/s included; {target_iops} IOPS needs provisioning above baseline")
print(f"minimum volume size to provision {target_iops} IOPS at gp3's 500 IOPS/GiB ratio: {min_size_gib} GiB")
Output:
required throughput for 5000 IOPS at 16 KiB each: 78.1 MiB/s
gp3 baseline: 3,000 IOPS / 125 MiB/s included; 5000 IOPS needs provisioning above baseline
minimum volume size to provision 5000 IOPS at gp3's 500 IOPS/GiB ratio: 10 GiB
The needed throughput (78.1 MiB/s) stays comfortably under gp3's 125 MiB/s free baseline, so the only dial to turn up is IOPS, and the 500:1 ratio means even a 10 GiB volume can legally carry 5,000 provisioned IOPS; in practice you'd size the volume for actual data capacity first and confirm the IOPS ratio is satisfied, not the other way around.
Trade-offs and pitfalls
- Don't put anything you can't afford to lose on instance store purely because it's faster; verify what "faster" is actually buying you against what "the data disappears on stop/terminate/hardware failure" costs you if you're wrong about the workload being truly recomputable or disposable.
- gp2 (the older general-purpose EBS type) ties IOPS to volume size and relies on burst credits (a banked allowance of extra I/O: the volume accrues credits while running below its baseline IOPS and spends them to briefly exceed that baseline; once the bank runs out, performance drops back down to baseline until credits build back up); gp3 decouples IOPS and throughput from size entirely, which is why gp3 is the current default recommendation, not gp2.
- Network-attached block storage has an inherent latency floor that pure local NVMe (a high-speed local solid-state storage interface) instance store doesn't: even at gp3/io2's best-case latency, there's a network hop involved that local storage skips. For workloads sensitive to the last fraction of a millisecond (not just IOPS), that architectural difference matters more than any provisioned-IOPS number.
- A checkpoint strategy is only as good as its interval: computing the actual maximum tolerable data loss (checkpoint interval times the rate of work) and comparing it against the business requirement is a five-minute calculation that catches a checkpoint-too-infrequent design before it ships, not after an incident.
Design an archival policy that moves older data from a warm or hot storage tier into a cheaper cold or archival tier, without breaking jobs that occasionally still need to read snapshots of that older data. Cover the lifecycle transition rules you would set, how you would catalog or index archived data so it can still be found, the restore workflow and its expected latency, the trade-off between retrieval cost and access speed, and how you would avoid unexpected restore failures or surprise costs when older data actually gets requested back.
Sample Answer
Direct answer
Build the policy around three separate pieces that people usually collapse into one: (1) transition rules that move data between tiers automatically based on age or access pattern, (2) a catalog that always knows which tier and which retrieval class a given piece of data currently lives in, independent of where it physically sits, and (3) a restore path whose retrieval-speed class is chosen up front, at tiering time, to match the worst-case restore-time SLA (service-level agreement: a promised or contractually required target, here the maximum acceptable time to get requested data back) that data might ever need, not improvised after someone requests it back. Verify the whole thing works with real, scheduled restore drills rather than trusting the design on paper.
Lifecycle transition rules
- Trigger type: age-based (simplest: "move to cold after 365 days," configured as a native bucket/table lifecycle policy, e.g. S3 Lifecycle rules or Azure Blob lifecycle management, so the storage service itself enforces it rather than relying on an application job someone eventually forgets to run) or access-frequency-based (more precise: track actual read frequency and demote data that has genuinely gone cold, e.g. S3 Intelligent-Tiering's automatic monitoring). Age-based is the right default; add frequency-based only where access patterns don't correlate cleanly with age.
- Event-based overrides: some data should move immediately on a business event (a case closes, a record is finalized) rather than waiting for an age threshold; model this as an explicit catalog write that fires the transition, not a special case bolted onto the age rule.
- Prefer declarative, storage-service-native lifecycle rules over custom migration code wherever the storage system supports it: a misconfigured lifecycle rule fails loudly (data doesn't move, and you can see that in the catalog's per-tier byte counts); a custom migration job can fail silently and nobody notices until a restore request comes in for data that was never actually archived.
Cataloging and indexing archived data
The moment data leaves the hot, natively-queryable store, you need a durable catalog that maps a logical identifier (a file path, a partition key, a record range) to where it currently lives: tier, storage/retrieval class, the physical object key, its size, a checksum recorded at archive time, and when it becomes eligible for expiry. Without this, "find it" degenerates into enumerating every tier by hand.
- Implement it as a small, always-hot, highly available table (a Postgres table, a DynamoDB table, or, if the archived data is itself tabular, an open-table-format manifest such as Apache Iceberg or Delta Lake: a structured metadata file format used in data-lake/analytics stacks that tracks a table's underlying data files, schema, and partitions), never as something that itself gets archived: the catalog is on the critical path of every restore, so it must always be instantly readable.
- Update the catalog as part of the same operation that performs the tier transition (or reconcile it against the transition job reliably), not as an unrelated best-effort side effect; a catalog that is out of sync with reality is worse than no catalog, because it answers confidently and wrong at exactly the moment (an audit, an incident) when correctness matters most.
Restore workflow and retrieval-speed class
Cold-tier restores aren't uniform: choose among near-instant, expedited, or bulk retrieval classes based on the restore-time SLA the specific data actually needs, decided when the data is tiered, not guessed at restore time.
- Near-instant (e.g. S3 Glacier Instant Retrieval): millisecond-latency GETs, priced closer to a warm tier. Use it for cold data that is rarely read but must never make a caller wait.
- Expedited (e.g. S3 Glacier Flexible Retrieval's Expedited option): typically 1 to 5 minutes, at a premium per-request and per-GB retrieval price. Use it for data bound by a tight, hard restore-time SLA that near-instant pricing can't justify for the whole dataset.
- Bulk: the cheapest per-GB retrieval option, typically hours (and up to about two days on the deepest archival classes). This is the right default for the vast majority of archived data, which is rarely if ever requested back and has no hard restore-time bound.
The mistake this guards against: assuming any cold tier can be rehydrated "fast enough" and only discovering the real restore-time bound during an actual audit or incident, when it's too late to move the data into a faster-retrieval class.
After a restore completes, verify it, don't just trust that the read call returned success: recompute a checksum on the restored object and compare it against the checksum recorded in the catalog at archive time.
Avoiding restore failures and surprise costs
- Restore-drill verification: schedule periodic real test restores, not synthetic health checks. Pick a sample of archived partitions, run the actual restore workflow end to end, verify row counts and checksums against the catalog's recorded values, and log the observed restore duration as evidence the assigned retrieval class is genuinely meeting its SLA target. This is what catches a wrong retrieval-class assignment before a real, time-boxed request does.
- Automated compliance reporting: a recurring report or dashboard that reads the catalog and surfaces per-tier byte counts (so data that silently failed to migrate shows up as an unexpected size anomaly in the wrong tier), upcoming retention expirations (so purges happen exactly on schedule, neither late, which is a compliance finding, nor early, which is evidence destruction), and legal-hold status per record.
- Surprise-cost guardrail: expedited and near-instant retrieval are priced at a premium per GB and per request; a misconfigured or accidental bulk job that restores a whole cold partition at expedited speed instead of bulk can spend a meaningful fraction of a month's storage budget in one run. Enforce a retrieval-class allowlist per caller (only the narrow incident-response path is allowed to request expedited; scheduled batch jobs are pinned to bulk) and alert on retrieval-request volume.
- Legal hold and lifecycle expiry are two independent gates on deletion, and both must block it: a lifecycle rule alone will happily delete data that is under active litigation hold unless the hold is checked first.
Worked example: 7-year audit-log retention with a 4-hour rehydrate requirement
A system ingests 200 GB/day of audit logs, subject to a 7-year regulatory retention requirement, and an audit process that can demand a named partition rehydrated within 4 hours.
gb_per_day, years, hot_days = 200, 7, 30
warm_days = 365 - hot_days
total_gb = gb_per_day * 365 * years
hot_gb = gb_per_day * hot_days
warm_gb = gb_per_day * warm_days
cold_gb = gb_per_day * 365 * (years - 1)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_ia, p_deep, p_flex = 0.023, 0.0125, 0.00099, 0.0036
sla_frac = 0.10 # share of the cold tier subject to the 4h rehydrate SLA
cold_sla_gb = cold_gb * sla_frac
cold_bulk_gb = cold_gb * (1 - sla_frac)
cost_split_cold = cold_sla_gb * p_flex + cold_bulk_gb * p_deep
total_monthly = hot_gb * p_std + warm_gb * p_ia + cost_split_cold
baseline_all_std = total_gb * p_std
print(f"total 7yr volume: {total_gb:,} GB")
print(f"tiers: hot={hot_gb:,} GB, warm={warm_gb:,} GB, cold={cold_gb:,} GB "
f"(of which {cold_sla_gb:,.0f} GB is SLA-bound)")
print(f"monthly cost, tiered with the SLA split: ${total_monthly:,.0f}")
print(f"monthly cost, all-Standard baseline: ${baseline_all_std:,.0f}")
print(f"savings: {(1 - total_monthly/baseline_all_std):.0%}")
Output:
total 7yr volume: 511,000 GB
tiers: hot=6,000 GB, warm=67,000 GB, cold=438,000 GB (of which 43,800 GB is SLA-bound)
monthly cost, tiered with the SLA split: $1,523
monthly cost, all-Standard baseline: $11,753
savings: 87%
The concrete decision the SLA forces: S3 Glacier Deep Archive's retrieval options are Standard (9-12 hours) and Bulk (up to 48 hours); neither meets a 4-hour bound. The 10% of cold data actually subject to that bound has to live in a faster-retrieval class (here, Glacier Flexible Retrieval, whose Expedited option is typically 1-5 minutes) instead, at roughly 3.6x Deep Archive's per-GB price, an explicit, budgeted trade for meeting the SLA rather than an assumption that gets discovered wrong during a real audit. The restore drill for this slice runs monthly: pick one random SLA-bound partition, run an Expedited restore, record the wall-clock time to completion, verify its row count and checksum against the catalog, and post the result to the compliance dashboard, so drift toward missing the 4-hour bound (for instance during a regional capacity crunch) is caught by the drill instead of by an auditor.
Trade-offs and pitfalls
- Defaulting everything to near-instant or expedited retrieval defeats the purpose of tiering: most archived data is never requested back, and the savings tiering exists to capture come from most of it staying in the cheapest, slowest-to-restore class. Reserve the premium retrieval class narrowly, for the specific data actually bound by a hard SLA, the way the worked example does with its 10% split.
- A catalog is only as trustworthy as its consistency with the actual data movement; treat catalog updates as part of the transition operation, not an afterthought.
- "The restore call returned success" and "the restored data is correct" are different claims; skip the checksum/row-count verification and a silently truncated or corrupted restore is indistinguishable from a good one until someone reads the data and notices.
- Don't wait for a real audit to discover a retrieval-class mismatch. A restore drill that only checks "did the API call succeed" without timing it against the SLA target isn't actually validating the thing that matters.
You have six identical disks attached to a Linux server and must create a software RAID configuration that tolerates two simultaneous disk failures while maximizing usable capacity and reasonable rebuild times. Which RAID level would you choose and why? Provide mdadm commands to create the array with a hot spare, configure automatic assembly at boot, and steps to monitor and replace failed drives.
Sample Answer
Direct answer
Use RAID6 (redundant array of independent disks, level 6): it splits each write into fixed-size pieces called chunks (for example 512 KiB each) and lays one chunk on every disk in a repeating round called a stripe, so no single disk holds a whole file and the disks can be read from or written to in parallel. For every stripe it also computes and writes two independent parity blocks (redundant values mathematically derived from the real data in that stripe; if a disk goes missing, the surviving disks' data plus a parity block is enough to recompute exactly what was on the missing disk) instead of one, so the array keeps serving reads and writes correctly through any two simultaneous disk failures, not just some of them. Usable capacity is (n-2) x disk_size (4 of the 6 disks' worth), and rebuild uses the same dual-parity math a rebuild after a single failure would use, just re-run twice if needed. RAID10 (mirrored pairs, striped) is faster to rebuild and has a cheaper write cost, but with 6 disks in 3 mirror pairs it only survives a second failure if that disk isn't the mirror of the first one that failed, a coin-flip-adjacent guarantee, not a deterministic one, so it does not meet "tolerates two simultaneous failures" as stated.
Comparing the candidates
| RAID level | Disks needed | Usable capacity | Simultaneous-failure guarantee | Random-write cost |
|---|---|---|---|---|
| RAID5 (single parity) | n >= 3 | (n-1) x disk | Survives exactly 1 failure; a 2nd is total data loss | 4 I/Os per write (read data, read parity, write data, write parity) |
| RAID6 (dual parity) | n >= 4 | (n-2) x disk | Survives any 2 simultaneous failures, guaranteed | 6 I/Os per write (read data, read parity1, read parity2, write data, write parity1, write parity2) |
| RAID10 (mirror + stripe) | n >= 4, even | n/2 x disk | Survives 2 failures only if they aren't mirror partners; not guaranteed | 2 I/Os per write (mirror write), fastest of the three |
For 6 disks specifically, RAID10 is faster to rebuild and cheaper per write, but it loses on BOTH of the question's requirements: its usable capacity is only 3 disks' worth (n/2) against RAID6's 4 (n-2), and its dual-failure survival is conditional rather than guaranteed. RAID6 is the only one of the three that satisfies both the "maximize usable capacity" ask and the "tolerates two simultaneous failures" ask at once.
Creating the array with a hot spare
Assume the six data disks are /dev/sdb through /dev/sdg and a seventh disk, /dev/sdh, is the hot spare (a disk that sits idle and automatically takes over as soon as a member disk is marked failed, so the rebuild starts without a human swapping hardware first):
# Wipe any old RAID or filesystem signatures so mdadm doesn't warn/refuse
sudo mdadm --zero-superblock /dev/sd{b,c,d,e,f,g,h}
# Create the RAID6 array: 6 active members + 1 spare, metadata format 1.2
# (the modern default: stores the superblock near the start of the device
# instead of at the very end, so it survives partial-disk mistakes better)
sudo mdadm --create /dev/md0 \
--level=6 \
--raid-devices=6 \
--spare-devices=1 \
--metadata=1.2 \
--chunk=512 \
--name=data0 \
/dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf /dev/sdg /dev/sdh
# Watch the initial parity sync (a RAID6 array is usable immediately but
# runs a background full-stripe consistency pass after creation)
watch cat /proc/mdstat
# Put a filesystem on it once sync is running (safe to do; sync continues underneath)
sudo mkfs.xfs /dev/md0
--metadata=1.2, --name=, --chunk=, --raid-devices=, and --spare-devices= are the actual mdadm --create option names (verified by running mdadm --create with these exact flags against a scratch mdadm 4.3 install: the parser accepted every flag and only rejected the dummy missing placeholder devices I fed it, which is the expected outcome for that reason, not an unrecognized-option error).
Automatic assembly at boot
mdadm --create activates the array immediately, but the kernel has no memory of it across a reboot unless you record it:
# Append the array's UUID-based definition to the mdadm config file
sudo mdadm --detail --scan | sudo tee -a /etc/mdadm/mdadm.conf
# Rebuild the initramfs (the small boot-time filesystem that assembles RAID/LVM
# before the real root filesystem is mounted) so it picks up the new array
sudo update-initramfs -u
# Confirm mdadm.conf now has an ARRAY line naming /dev/md0 by UUID, not by
# /dev/sdX letters (those letters can be reassigned on reboot; the UUID can't)
cat /etc/mdadm/mdadm.conf
On the Debian/Ubuntu package (confirmed by actually installing mdadm and reading its generated /etc/mdadm/mdadm.conf), the file ships with the comment "Run update-initramfs -u after updating this file", which is exactly this step. On an RPM-based distro the equivalent is dracut -f instead of update-initramfs -u; the mdadm --detail --scan >> /etc/mdadm/mdadm.conf step is the same on both.
Monitoring the array
# Point-in-time state: which disks are active/spare/failed, sync %, etc.
sudo mdadm --detail /dev/md0
# Live kernel view, updates continuously
cat /proc/mdstat
# Run mdadm as a background daemon that emails on any array event
# (disk failure, degraded state, rebuild complete, spare activated)
sudo mdadm --monitor --scan --daemonise --mail=root --delay=1800
# Schedule periodic full-stripe consistency checks (most distros ship this
# as a cron/systemd-timer job already, e.g. /etc/cron.d/mdadm on Debian)
echo check | sudo tee /sys/block/md0/md/sync_action
Pair this with SMART monitoring on the physical disks themselves (smartctl -a /dev/sdb, via the smartmontools package) since mdadm only sees I/O errors the disk already reported; SMART's predictive attributes (reallocated sector count, pending sector count) often flag a drive as failing before it actually drops out of the array.
Worked example: capacity, rebuild time, and why not RAID5
Six 4 TiB disks (TiB = tebibyte, 2^40 bytes, the binary unit df/lsblk report in):
def raid6_usable_tib(n_disks, disk_tib):
return (n_disks - 2) * disk_tib
def rebuild_hours(disk_gib, assumed_sustained_mib_s):
capacity_mib = disk_gib * 1024
seconds = capacity_mib / assumed_sustained_mib_s
return seconds / 3600
n, disk_tib = 6, 4
usable = raid6_usable_tib(n, disk_tib)
raw = n * disk_tib
print(f"raw={raw} TiB, usable={usable} TiB, parity overhead={raw-usable} TiB ({(raw-usable)/raw:.0%})")
print(f"rebuild at an assumed 150 MiB/s sustained: {rebuild_hours(4096, 150):.1f} hours")
print(f"rebuild throttled to 30 MiB/s (production I/O sharing the disks): {rebuild_hours(4096, 30):.1f} hours")
Output:
raw=24 TiB, usable=16 TiB, parity overhead=8 TiB (33%)
rebuild at an assumed 150 MiB/s sustained: 7.8 hours
rebuild throttled to 30 MiB/s (production I/O sharing the disks): 38.8 hours
(150 MiB/s and 30 MiB/s are stated assumptions, not a measurement of any specific disk; mdadm's real rebuild speed is governed by /proc/sys/dev/raid/speed_limit_min and speed_limit_max, and actual throughput depends on the disks and how much production I/O is competing for them.)
Why RAID5 is the wrong call for an array this size, computed rather than asserted: rebuilding a failed disk in a RAID5 array means reading every block of every surviving disk to reconstruct the missing one. Consumer and nearline SATA drives commonly publish an unrecoverable-read-error (URE) spec of about 1 bit in 10^14 read:
n, disk_tib = 6, 4
surviving = n - 1
bits_to_read = surviving * disk_tib * 8 * 1024**4
ure_rate = 1e-14
p_hit = 1 - (1 - ure_rate) ** bits_to_read
print(f"bits read during a RAID5 rebuild: {bits_to_read:.3e}")
print(f"P(hit >= 1 URE during that single rebuild): {p_hit:.0%}")
Output:
bits read during a RAID5 rebuild: 1.759e+14
P(hit >= 1 URE during that single rebuild): 83%
An 83% chance of hitting an unreadable sector mid-rebuild, on a single-parity array that has zero spare parity left once one disk is already gone, is the concrete reason "6 disks, large capacity, survive 2 failures" rules out RAID5 even before the "two simultaneous failures" requirement is read literally. RAID6's second parity block absorbs exactly that kind of single bad sector during a rebuild.
Replacing a failed drive
# mdadm already marked the failed member 'F' in --detail output and kicked
# the hot spare into a rebuild automatically; confirm that happened:
sudo mdadm --detail /dev/md0 | grep -E "State|faulty|spare|rebuild"
# If the disk hasn't been auto-failed yet (e.g. it's throwing errors but
# hasn't hard-failed), fail and remove it explicitly before pulling it:
sudo mdadm --manage /dev/md0 --fail /dev/sdd
sudo mdadm --manage /dev/md0 --remove /dev/sdd
# Physically swap the drive, then partition/prepare the replacement the
# same way the originals were, and add it back as the new spare:
sudo mdadm --manage /dev/md0 --add /dev/sdd
# Watch the rebuild (or, if a spare already absorbed the first failure,
# this --add just refills the spare pool back to 1):
watch cat /proc/mdstat
--fail, --remove, and --add are real mdadm --manage option names (confirmed against mdadm --manage --help output on the same install).
Trade-offs and pitfalls
- RAID is not backup: it protects against disk hardware failure, not against accidental deletion, ransomware, or a bad
rm -rf, all of which replicate instantly across every mirror/parity member. Keep an actual backup regardless of RAID level. - RAID6's dual-parity write cost (6 I/Os per small random write versus RAID10's 2) makes it a poor fit for a workload that demands high write IOPS (input/output operations per second, the count of individual read or write operations a storage system can complete each second); it earns its keep on capacity-and-durability-first workloads (bulk storage, backups-of-backups, archival-adjacent data), which matches "maximizing usable capacity" in the question.
- The rebuild window is the danger zone: for as long as a rebuild runs, the array is one more failure away from data loss (RAID6 degraded to single-parity mid-rebuild, or fully exposed if a 2nd failure lands before the 1st rebuild finishes). A hot spare shrinks that window by starting the rebuild automatically instead of waiting on a human to notice and swap hardware.
--chunk=512(512 KiB) is a reasonable general-purpose default; a workload dominated by large sequential I/O (video, backups) benefits from a larger chunk, one dominated by small random I/O from a smaller one. Get this from testing your actual workload, not by guessing.- Never trust
/dev/sdXletters as stable identifiers in configuration; they can be reassigned on reboot depending on device enumeration order. That's exactly whymdadm --detail --scanwrites the array definition by UUID.
That is every published Storage Systems and Infrastructure question for Systems Administrator so far. Browse the other topics in this category, or practice this one interactively.