Infrastructure Scaling, Capacity Planning, and High Availability Questions
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
Explain how deploying services across multiple availability zones and regions changes network capacity planning. Cover the effect on latency, inter-region bandwidth, and replication costs, and how you'd factor cross-AZ and cross-region transfer fees into your sizing decisions.
You are planning a three-region deployment for a sharded database and must guarantee <50ms read latency for most users accessing their nearest region. Describe a data placement and replication strategy, how to route reads and writes, techniques to limit cross-region latency, and how to handle failover and rebalancing when one region becomes unavailable.
Write a Python script or clear pseudocode that reads a CSV of historic batch jobs (columns: job_id, avg_runtime_seconds, avg_cpu_seconds, avg_concurrent_runs) and outputs a recommended minimum cluster vCPU count to meet a target average job latency. Explain the math and assumptions (parallelism, headroom), and provide sample output for three example jobs.
Describe how you'd design an observability and forecasting system that reliably predicts capacity exhaustion 24 hours ahead. Cover the data you'd feed it, the forecasting approach you'd pick and why, how you'd set alerting thresholds and runbooks for action, and how you'd keep false positives low in a noisy environment.
Explain the key resource utilization metrics you would monitor for capacity planning across compute, memory, storage I/O and network for a cloud service. For each metric (CPU, memory, disk I/O, network throughput) describe a concrete SLI you would define, why that SLI helps capacity planning, and one practical alert threshold or dashboard visualization you would create to track it.
Unlock Full Question Bank
Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.