Scalability Patterns and Techniques Questions

Scaling a system to handle growth in traffic and data: horizontal versus vertical scaling, statelessness, sharding and partitioning strategies, read replicas, and connection pooling. Covers capacity estimation, identifying bottlenecks, and the tradeoffs each scaling axis introduces. The general toolkit for taking a design from thousands to millions of users.

EasyTechnical
29 practiced

As an engineering manager, how would you coach a team through deciding where to draw microservice boundaries? Discuss bounded contexts, ownership, data ownership, coupling, how many teams are involved, deployment cadence, and operational cost. Walk through a concrete example of splitting a monolith into two services and your reasoning.

EasyTechnical
29 practiced

Define edge caching and origin caching in plain terms for a cross-functional audience. For a photo-sharing app with 50 million daily active users and highly bursty traffic, which caching layers would you prioritize: CDN, regional caches, or application cache, and why? Briefly describe your invalidation strategy, how much staleness you'd accept, and the performance metrics you'd monitor.

EasyTechnical
53 practiced

As an engineering manager, describe a simple capacity-planning approach for a service expected to grow 3x in traffic over the next 12 months. What inputs would you gather, such as current QPS and P95 CPU/memory per instance? Walk through the key calculations for forecasting instance or shard counts, and how you'd turn that forecast into hiring, infrastructure, or autoscaling decisions.

EasyTechnical
26 practiced

You are a Technical Product Manager for a cloud developer platform. Define horizontal scaling versus vertical scaling in concrete terms, then give two product scenarios (one favoring horizontal, one favoring vertical) and explain the trade-offs in cost, downtime risk, operational complexity, observability, and developer experience. How would you influence engineering's choice, and what metrics would you monitor to validate it?

MediumTechnical
43 practiced

A service uses in-memory caches to reduce database load, but it's experiencing cache thrashing during traffic spikes. As the engineering manager reviewing this with your team, what are the likely causes, and what fixes would you expect the team to propose? Consider eviction-policy tuning, tiered caching (edge, regional, origin), request coalescing, warm caches, and circuit breakers. How would you quantify the expected reduction in database load for a given improvement in hit rate?

Unlock Full Question Bank

Get access to all 9 Scalability Patterns and Techniques interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.