Infrastructure Scaling, Capacity Planning, and High Availability Questions

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

MediumSystem Design
59 practiced

Design session management for a shopping-cart web app that needs strong session affinity for cart updates, while still supporting horizontal autoscaling and high availability. What approach would you take, and how would you weigh the trade-offs around consistency, latency, failover, operational complexity, and cost?

EasyTechnical
73 practiced

Explain the key differences between round-robin, least-connections, and consistent-hashing load balancing algorithms. For each algorithm, describe the scenarios where it performs well and where it performs poorly (consider stateless vs stateful services, highly variable request cost, cache locality and rebalancing cost). Give one concrete example service for each algorithm choice.

MediumTechnical
74 practiced

Given a monolithic service with variable load across features, describe how you would identify components that should scale independently versus those that should remain co-located. Include specific metrics to collect, dependency analysis techniques, and cost vs operational complexity trade-offs.

EasyTechnical
70 practiced

Describe connection draining and graceful termination for services behind load balancers. Explain how to implement this with cloud load balancers, NGINX, and Kubernetes (preStop hook, pod terminationGracePeriod), and how draining should interact with autoscaling events and rolling deployments.

EasyTechnical
69 practiced

What does it mean for a service to be stateless versus stateful, and why does statelessness make load balancing and autoscaling so much easier? Now say you have a component that genuinely has to hold state: how would you scale it horizontally, and what trade-offs come with the approach you'd pick?

Unlock Full Question Bank

Get access to all 11 Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.