Infrastructure Scaling, Capacity Planning, and High Availability Questions

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

HardSystem Design
68 practiced

Design an automated right-sizing system that analyzes historical telemetry across compute and storage accounts and recommends instance or container sizes, complete with estimated savings and risk. Walk through the data pipeline, what signals and features you'd engineer, how you'd decide what modeling approach fits each recommendation type, how you'd evaluate whether a recommendation is safe to act on, and how you'd roll changes out with a human still in the loop.

MediumTechnical
52 practiced

Given a constrained budget, outline a repeatable process to right-size the compute instances, containers, and network services you run. Describe your data collection methods, percentile-based sizing analysis, pilot change execution, rollback plans, and how you would quantify validated cost savings while maintaining SLAs.

MediumTechnical
70 practiced

A production cluster suffers from frequent 'pod pending' events because the scheduler cannot find nodes with sufficient CPU or memory. Walk through a prioritized troubleshooting checklist (kubectl commands, metrics, and logs) to identify root cause and recommend short-term and medium-term remediation.

HardSystem Design
59 practiced

Design a multi-region data placement and routing strategy for a globally distributed, read-heavy service that requires low read latency and eventual consistency for writes. Describe your replication topology, routing choices for read locality, estimated network bandwidth for replication, and failure modes that will affect capacity planning and how you would mitigate them.

HardSystem Design
74 practiced

Design an end-to-end experiment to validate a Kubernetes cluster autoscaler's behavior at scale (e.g., supporting 10,000 pods). Include how you would generate load, inject realistic pod start times and custom metrics, simulate node provisioning delays and failures, collect relevant telemetry, and define pass/fail criteria for scaling latency and stability.

Unlock Full Question Bank

Get access to all Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.