Infrastructure Scaling, Capacity Planning, and High Availability Questions

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

MediumTechnical
59 practiced

Explain how to translate SLOs into concrete capacity targets. Given a requirement of 99.95% availability and p95 latency under 100ms, explain how you'd set CPU/memory headroom, instance counts, redundancy levels (N+1/N+2), and load balancing choices to meet SLOs including failure scenarios such as instance or AZ loss.

MediumSystem Design
56 practiced

Design an automated node autorepair process for Kubernetes nodes that become NotReady or fail health checks frequently. Include detection logic, steps to cordon/drain, reprovision methods, and how to avoid cascading failures during simultaneous repairs.

EasySystem Design
56 practiced

Design a basic autoscaling policy for a stateless web service running behind a cloud load balancer (for example AWS behind an Application Load Balancer). Decide what metric or metrics should drive your scaling decisions and why, your sampling windows and alarm thresholds for scale-up versus scale-down, cooldown periods, minimum and maximum instance counts, health-check and warm-up considerations, and any safety limits you would add to prevent oscillation.

MediumTechnical
68 practiced

Your manager asks you to prove that a recent capacity or autoscaling change reduced cost (for example, by 20%) while keeping p95 latency and other SLOs intact. Describe a repeatable measurement plan: key metrics (including a cost-efficiency metric such as cost per successful user action), the dashboards you would build, an experiment or rollback plan, statistical tests or confidence intervals, and how you would attribute the cost and latency changes to the capacity change rather than to traffic variance.

HardSystem Design
73 practiced

Design an autoscaling policy for a stateful queue-processing service that scales horizontally based on messages per consumer and must guarantee at-least-once processing with no message loss during scale-in. Walk through exactly what happens to a consumer from the moment it's marked for removal to the moment it exits, and explain how your cooldowns and thresholds would differ from a stateless service's.

Unlock Full Question Bank

Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.