Storage Systems and Infrastructure Questions

The physical storage substrate that sits beneath databases and other data-intensive systems: disk and volume management, RAID levels and disk-redundancy trade-offs (capacity vs fault tolerance vs rebuild time), and cloud block and instance-store performance characteristics (IOPS, throughput, latency), including how ephemeral instance-store volumes differ from persistent attached ones. Also covers storage-service tiering: hot, warm, cold, and archival lifecycle policies, the retrieval-cost-versus-latency trade-off, and designing the automation that migrates data between tiers and routes reads across them while meeting latency and cost targets. Covers matching a storage configuration's redundancy and performance profile to a system's durability and throughput requirements. This is distinct from how a database engine implements storage internally (write-ahead logs, page layouts, B-tree versus LSM structures). Aimed at engineers who configure and operate the underlying storage hardware.

HardSystem Design
69 practiced

Given a fixed budget, design a storage tiering system that automatically migrates data between hot, warm, and cold tiers (for example local NVMe, a warm SSD cluster, and an object store) to balance latency and cost. What criteria would you use to decide when a piece of data should move between tiers, how would you architect the automation that carries out those migrations safely, how would reads be routed so a query transparently spans whichever tiers hold the data it needs, and how would you measure and enforce latency and cost SLOs across the whole system?

MediumTechnical
84 practiced

Describe cloud block storage (for example AWS EBS or Azure Managed Disks). What are its key performance characteristics (IOPS, throughput, latency), and what are some typical use cases, such as scratch space for a compute cluster, a database's backing disk, or a local cache? Compare ephemeral instance-store volumes to persistent attached block volumes, and explain what that difference means for fault tolerance and for any workload that periodically saves its progress to disk.

MediumTechnical
62 practiced

Design an archival policy that moves older data from a warm or hot storage tier into a cheaper cold or archival tier, without breaking jobs that occasionally still need to read snapshots of that older data. Cover the lifecycle transition rules you would set, how you would catalog or index archived data so it can still be found, the restore workflow and its expected latency, the trade-off between retrieval cost and access speed, and how you would avoid unexpected restore failures or surprise costs when older data actually gets requested back.

EasyTechnical
61 practiced

You're deciding how long to keep data in fast online storage before moving it to cheaper archival storage. What business and technical factors would you weigh in that decision? Walk through an example retention policy for a system that has a 3-year legal retention requirement, where data from the last 90 days is queried frequently and older data is rarely touched.

That is every published Storage Systems and Infrastructure question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.