Infrastructure Scaling, Capacity Planning, and High Availability Questions

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

HardSystem Design
109 practiced

You are planning a three-region deployment for a sharded database and must guarantee <50ms read latency for most users accessing their nearest region. Describe a data placement and replication strategy, how to route reads and writes, techniques to limit cross-region latency, and how to handle failover and rebalancing when one region becomes unavailable.

HardTechnical
70 practiced

Perform a rigorous postmortem for an incident where autoscaling caused simultaneous scale-in across regions during a lull, eliminating needed headroom; when traffic returned the system experienced packet drops and partial outage. Identify probable root causes, immediate mitigations you would have applied, and propose architectural and process changes to prevent similar incidents in the future.

MediumTechnical
52 practiced

Compare DNS-based load balancing, Anycast routing, and cloud provider global load balancers / CDNs for distributing traffic geographically. For each approach discuss failover time, granularity of routing (geo vs latency), TTL and client caching effects, and operational complexity. Provide real-world examples where each approach is appropriate.

MediumTechnical
60 practiced

Explain how service discovery integrates with load balancers in microservices architectures. Compare DNS-based discovery, client-side discovery with registries (Consul/Eureka), and server-side discovery with gateways or service meshes. For each pattern describe how instances register/deregister, how the LB learns about backends, and how to handle TTLs and stale entries.

HardTechnical
69 practiced

Provide pseudocode or Python-like code for a metric-driven autoscaling algorithm that combines CPU, memory, network throughput, and a custom application metric (such as active sessions) into a single scaling decision. Explain the smoothing, thresholding, and safety mechanisms you build in to avoid flapping and runaway scaling, and walk through the key failure modes your algorithm mitigates.

Unlock Full Question Bank

Get access to all 46 Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.