Infrastructure Scaling, Capacity Planning, and High Availability Questions
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
You must size a Kubernetes cluster that will run 200 pods, each requesting 0.5 CPU and 1 GiB memory. Design node types and counts, include system overhead, and show calculations including a 20% buffer for bin-packing inefficiency. Discuss use of spot instances and mixed node pools.
How do you plan capacity and estimate resource needs for an analytics cluster used for nightly ETL and daytime dashboards? Include considerations for peak concurrency, buffer for failover, storage IOPS, and cost forecasting. Show how you'd model growth and decide when to add nodes versus optimize jobs.
How do SLOs and SLAs influence capacity planning? Given an SLO stating 99.9% of requests must finish within 200ms, explain how you would translate that into capacity targets and safety margins (including how to account for error budget and traffic variability).
Design a cost-aware provisioning and autoscaling strategy for a service that relies primarily on spot or preemptible instances to control cost, while guaranteeing a minimum reliable capacity and protecting your SLA or SLO. Cover the admission and allocation logic for new nodes, how you detect and replace lost capacity (bidding or fleet strategy, on-demand fallback, checkpointing), how you dynamically weight instance types by price and expected reliability, and the metrics you would monitor to catch a cost-vs-reliability regression.
Explain how you would evaluate and improve the accuracy of capacity forecasts: what metrics would you track, how would you backtest and calibrate the model, and how would you detect a structural break or concept drift in the underlying data before it silently degrades your forecasts?
Unlock Full Question Bank
Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.