Infrastructure Scaling, Capacity Planning, and High Availability Questions
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
Design a comprehensive three-to-five-year capacity roadmap for a global SaaS provider with 20% annual user growth, requirements for EU and APAC data residency, and a 99.95% network SLA. Include your recommended network architecture, autoscaling strategies for network functions, circuit procurement model, rough cost estimates, and the quarterly validation experiments you would run to test your assumptions.
Forecasting problem: current cluster processes 50k RPS with average CPU utilization 60% and p95 latency within SLO. Product expects 30% traffic growth in 6 months and occasional 5x flash traffic spikes. Propose a capacity plan (horizontal vs vertical scaling, autoscaling policies, buffer sizing, and how you would model peak concurrency), how you'd validate the plan with load testing, and how you would present the resulting cost-versus-headroom trade-offs to product and finance stakeholders.
Design a 3 to 5 year capacity forecasting methodology for an SRE organization supporting many services while business metrics have high uncertainty. Describe inputs, scenario planning (best, base, worst), probabilistic techniques such as Monte Carlo simulations, SLA-driven capacity choices, stakeholder reporting, and how frequently to update the roadmap.
Given unit costs: VM vCPU $0.02/hr, memory $0.01/GB/hr, storage $0.0001/GB/hr, network egress $0.09/GB, outline a method to estimate monthly cloud cost when scaling from 100 to 10,000 concurrent connections. State the assumptions you would need (per-connection CPU, memory, average session duration), redundancy factor, and how you'd include autoscaling behavior and headroom in the cost estimate.
Given a historical traffic series (for example, 36 months of data) showing 20 percent month-over-month growth and measured per-request CPU and memory usage, outline a three year capacity roadmap. Include how you convert traffic projections into RPS and instance counts, how you add safety margins, how you account for procurement lead times, and what assumptions you must document for finance and exec stakeholders.
Unlock Full Question Bank
Get access to all 16 Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.