Scalability Patterns and Techniques Questions

Scaling a system to handle growth in traffic and data: horizontal versus vertical scaling, statelessness, sharding and partitioning strategies, read replicas, and connection pooling. Covers capacity estimation, identifying bottlenecks, and the tradeoffs each scaling axis introduces. The general toolkit for taking a design from thousands to millions of users.

HardTechnical
24 practiced

During a large traffic spike, your cloud autoscaler hit a quota limit and the service breached its SLOs. As the incident commander, what immediate mitigations would you take (manual scaling, throttling), how would you communicate with your cloud provider and stakeholders, and what medium-term fixes (quota monitoring, predictive scaling) would you put in place? How would you update runbooks and alerts to prevent a repeat?

EasyTechnical
36 practiced

Why does connection pooling matter for a service running at scale? Describe best practices for managing both database and HTTP connection pools: pool size, max open connections, idle timeouts, connection lifetime, and behavior under a spike in load. How would you test and tune these settings before production?

EasyTechnical
31 practiced

Explain the difference between horizontal partitioning (sharding) and vertical partitioning for scaling a dataset. What criteria would you use to choose a shard key, how would you plan and execute a resharding or rehashing operation in a cloud environment, and what operational challenges (rebalancing, hotspots, migration windows) should you anticipate?

EasyTechnical
25 practiced

When should you reach for asynchronous processing in a cloud architecture? Give examples of tasks that belong on a queue, the benefits and trade-offs (latency, throughput, complexity), and how you'd choose between a simple message queue, a pub/sub system, and a streaming platform for different workloads.

HardTechnical
36 practiced

Explain the basic queueing-theory concepts behind capacity planning: arrival rate, service rate, utilization, and the M/M/1 queue. Walk through a simple numeric example showing how a small increase in utilization can produce a disproportionate increase in average latency.

Unlock Full Question Bank

Get access to all Scalability Patterns and Techniques interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.