Infrastructure Scaling, Capacity Planning, and High Availability Questions
Making infrastructure grow and stay up: horizontal and vertical scaling, autoscaling, load balancing, capacity planning and forecasting, and high-availability and redundancy design. Covers sizing systems for demand, distributing load, and eliminating single points of failure so services remain available as they scale. The reliability-and-growth discipline.
You must perform a non-disruptive rolling upgrade of BGP routers across three geographically separated data centers while maintaining active traffic flows. Provide a detailed step-by-step upgrade procedure, safeguards you would put in place, BGP traffic-engineering techniques (such as local-preference, MED, AS-path prepending) you would use to control traffic during the upgrade, testing/validation checks for each step, and rollback mechanisms.
Design a rigorous load-testing methodology to benchmark the maximum concurrent TCP sessions and connection churn a NAT gateway cluster can sustain. Include traffic generator orchestration, parameter sweeps (connection rate, idle timeout, payload size), statistical analysis to identify capacity limits and failure modes, and how you would extrapolate test results to predict production behavior and required headroom.
Design a basic autoscaling policy for stateless instances behind a cloud load balancer. Specify which metrics you would monitor (CPU, network in/out, requests per second), sampling windows, alarm thresholds for scale up/down, cooldown periods, minimum/maximum instance counts, and any safety limits to prevent oscillation.
Create a load-testing plan to benchmark a global web application that uses regional load balancers and DNS-based global routing. Specify test scenarios (steady-state regional load, sudden global spike, regional failover), traffic distribution by geography, tools to drive traffic and measure results, key metrics to capture (latency, error rates, backend CPU/queueing), and acceptance criteria for performance and stability.
As a network engineer, explain what "capacity planning" means for network infrastructure. Describe the core metrics you would track (for example: bandwidth utilization, throughput, packets-per-second, concurrent connections, device CPU and memory), how you define and choose headroom, and outline a simple 6–12 month process for creating and validating a capacity plan that aligns to SLAs.
Unlock Full Question Bank
Get access to all 35 Infrastructure Scaling, Capacity Planning, and High Availability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.