Infrastructure Scaling, Capacity Planning, and High Availability Questions

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

HardTechnical
55 practiced

A relational database under a write-heavy workload shows rising write latency and increasing lock contention. Walk through how you'd diagnose this using monitoring and database internals: what you'd check, what root causes you'd rule in or out, what you'd mitigate immediately, and what longer-term architecture changes you'd consider to scale writes.

EasyTechnical
61 practiced

Describe a practical load-testing approach you would run before a major release to validate capacity: the test types you'd run and why, your ramp pattern, the key metrics you'd capture, how you'd isolate the test from production traffic, and safe ways to run it without impacting real users.

MediumTechnical
67 practiced

A marketing campaign may create unpredictable short bursts of 5 to 10x normal traffic. How would you prepare the infrastructure so it absorbs the spike without the database buckling or costs spiraling? Walk through what you'd build in ahead of time versus what you'd rely on reacting to in the moment.

HardSystem Design
54 practiced

Design an active-active global architecture for a payments service that must support local low-latency writes while providing robust fraud prevention and reconciliation. Address data partitioning, conflict resolution, capacity headroom for cross-region surges, and operational playbooks for partition heals.

MediumTechnical
75 practiced

Describe how you would instrument application and infrastructure to detect capacity saturation before user impact occurs. Include which metrics to collect at app, infra, and queueing layers; recommended alert thresholds; dashboards; and synthetic checks you would run to validate capacity continuously.

Unlock Full Question Bank

Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.