Infrastructure Scaling, Capacity Planning, and High Availability Questions
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
Explain how to translate SLOs into concrete capacity targets. Given a requirement of 99.95% availability and p95 latency under 100ms, explain how you'd set CPU/memory headroom, instance counts, redundancy levels (N+1/N+2), and load balancing choices to meet SLOs including failure scenarios such as instance or AZ loss.
Design an automated node autorepair process for Kubernetes nodes that become NotReady or fail health checks frequently. Include detection logic, steps to cordon/drain, reprovision methods, and how to avoid cascading failures during simultaneous repairs.
Design a basic autoscaling policy for a stateless web service running behind a cloud load balancer (for example AWS behind an Application Load Balancer). Decide what metric or metrics should drive your scaling decisions and why, your sampling windows and alarm thresholds for scale-up versus scale-down, cooldown periods, minimum and maximum instance counts, health-check and warm-up considerations, and any safety limits you would add to prevent oscillation.
Your manager asks you to prove that a recent capacity or autoscaling change reduced cost (for example, by 20%) while keeping p95 latency and other SLOs intact. Describe a repeatable measurement plan: key metrics (including a cost-efficiency metric such as cost per successful user action), the dashboards you would build, an experiment or rollback plan, statistical tests or confidence intervals, and how you would attribute the cost and latency changes to the capacity change rather than to traffic variance.
Design an autoscaling policy for a stateful queue-processing service that scales horizontally based on messages per consumer and must guarantee at-least-once processing with no message loss during scale-in. Walk through exactly what happens to a consumer from the moment it's marked for removal to the moment it exits, and explain how your cooldowns and thresholds would differ from a stateless service's.
Unlock Full Question Bank
Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.