Infrastructure Scaling, Capacity Planning, and High Availability Questions
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
You are responsible for a 3-5 year capacity roadmap for a data platform that currently handles 10K DAUs and is expected to grow to 10M DAUs. The architecture includes Kafka ingestion, a nearline processing layer (Spark/Flink), an object data lake, and an analytics warehouse. Produce a high-level roadmap that lists growth assumptions, capacity inflection points, when and why to introduce structural changes (e.g., sharded Kafka clusters, hot/cold storage separation, regionalization), staffing and budget implications, and risk mitigation strategies.
Tell me about a time you had to make a capacity-related trade-off that impacted product features or delivery timelines (for example delaying model training to reduce cost or reducing concurrency to meet SLOs). Describe the situation, the options you considered, the decision you made, how you communicated it, and the outcome.
Forecasting problem: current cluster processes 50k RPS with average CPU utilization 60% and p95 latency within SLO. Product expects 30% traffic growth in 6 months and occasional 5x flash traffic spikes. Propose a capacity plan (horizontal vs vertical scaling, autoscaling policies, buffer sizing, and how you would model peak concurrency), how you'd validate the plan with load testing, and how you would present the resulting cost-versus-headroom trade-offs to product and finance stakeholders.
You're presenting a capacity plan that recommends a 20%+ buffer which increases monthly costs. The sales team pushes for a leaner plan to improve profit margins. How do you present the technical tradeoffs, quantify risks, and align stakeholders (sales, finance, SRE) to reach a decision that balances reliability and cost?
You are planning hardware procurement and cloud commitments for the next 3 years. Describe how to translate capacity forecasts into procurement timing, vendor selection criteria, contract types (on-demand, reserved, committed-use), risk mitigation (component lead times, spot shortages), and how to present the plan to procurement and finance.
Unlock Full Question Bank
Get access to all 33 Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.