Scalability Patterns and Techniques Questions
Scaling a system to handle growth in traffic and data: horizontal versus vertical scaling, statelessness, sharding and partitioning strategies, read replicas, and connection pooling. Covers capacity estimation, identifying bottlenecks, and the tradeoffs each scaling axis introduces. The general toolkit for taking a design from thousands to millions of users.
Define cache hit ratio, cache miss, and cache warmup. For a typical web service considering an application cache (memcached or Redis), when would you decide it's worth adding one, what hit ratio would justify the cost, and what are three practical ways to improve an existing cache's effectiveness?
Sample Answer
Direct answer
Cache hit ratio is the fraction of lookups a cache can answer without going to the origin (hits divided by total lookups); a cache miss is a lookup where the value isn't present or has expired, forcing a fetch from the origin and usually a write back into the cache; cache warmup is the process of populating a cache, at startup or after a restart, before it can offer a useful hit ratio. Whether adding an application cache (memcached, Redis, or similar) is worth it comes down to a simple cost comparison: what a miss costs you (a database round trip, a slow computation, an external API call) versus what running the cache costs you (infrastructure, and the risk of serving stale data), not a fixed hit-ratio threshold you're supposed to hit.
Structured elaboration
When to add a cache
- The workload is read-heavy with requests that repeat: the same or similar queries recur often enough that a cached answer is reused rather than computed once and discarded.
- The thing being cached is expensive relative to a cache lookup: a slow database query, an external API call, or a CPU-heavy computation are all good candidates; something already fast to compute gains little.
- The data can tolerate the staleness a cache implies, or the system can actively invalidate the cache on writes; if every read must reflect the absolute latest write, caching adds a consistency problem you have to solve, not just a performance win.
What hit ratio would justify the cost
There's no universal number, because "justified" depends on the ratio between the cost avoided per hit and the cost of running the cache, not on the hit ratio in isolation. A commonly cited planning heuristic for latency- and cost-sensitive systems is to aim for roughly 70 to 80% hit ratio; treat that as a starting heuristic to validate against your own workload, not as a target derived from anything specific to your system. A system where each miss is very expensive (an external, rate-limited, or metered API call) can be worth caching even at a 50% hit ratio, while a system where the origin call is already cheap may not justify caching even at 90%.
Three practical ways to improve an existing cache's effectiveness
- Tune time-to-live (TTL, how long a cached value is considered valid before it must be refreshed) per key type rather than using one global value: longer TTLs for stable data, shorter for volatile data, so you're not needlessly re-fetching stable data or serving stale volatile data.
- Improve cache key design: normalize keys (strip user-specific noise that doesn't actually change the result, use consistent prefixes) so semantically identical requests share a cache entry instead of each generating its own miss.
- Warm proactively and avoid a thundering herd on expiry: pre-populate known-hot keys at deploy or restart, and use a read-through or refresh-ahead pattern where a background job refreshes a key just before it expires, rather than letting many concurrent requests all miss on the same expired key at once.
Worked example
Assume, as a planning input rather than a measured fact, that an endpoint currently receives 10,000 requests per minute with no caching, and that each request costs 20 ms of database time. Total database time consumed per minute today:
10,000×20ms=200,000ms=200s of database time per minute
If a cache reaches a 75% hit ratio on this endpoint, the number of requests still reaching the database is:
10,000×(1−0.75)=2,500 requests per minute
2,500×20ms=50,000ms=50s of database time per minute
That's a 75% reduction in database load, directly tracking the hit ratio, from 200 seconds of database time per minute down to 50. Whether that reduction is "worth it" then depends on what those 150 seconds of freed-up database time are worth relative to running the cache, not on the 75% figure by itself.
Trade-offs & pitfalls
- Chasing a higher hit ratio as a goal in itself, rather than as a proxy for reduced backend cost, can lead to caching things that barely help (already-cheap lookups) while a genuinely expensive but less frequent lookup goes uncached.
- A TTL that's too long trades staleness risk for hit ratio; too short and the cache behaves closer to no cache at all under the same access pattern.
- Skewed (heavily concentrated, or Zipfian) access patterns mean a handful of keys drive most hits; eviction policy choice (LRU, least-recently-used, versus LFU, least-frequently-used) matters more under this kind of skew than under uniform access, since LRU can evict a very-frequently-used-but-not-most-recent key that LFU would keep.
- Expose hit and miss counters through your application performance monitoring (APM, application performance monitoring) or metrics stack; without that visibility, a regression in hit ratio after a code change (a key-format change, a new high-cardinality parameter) can go unnoticed until backend load spikes.
You need to decompose a monolithic application into microservices. Walk through a pragmatic approach: how you'd identify service boundaries, ensure data integrity during the migration, avoid distributed-transaction anti-patterns, and choose between orchestration and choreography. When would you reach for the strangler fig pattern?
Sample Answer
Direct answer
A pragmatic decomposition starts from real bounded contexts, not code modules, keeps data integrity intact by translating between the old and new models instead of cutting over all at once, and replaces any temptation toward a distributed two-phase commit with sagas built from idempotent, compensable steps. The strangler fig pattern, routing traffic through a facade so old and new implementations coexist during the transition, is the mechanism that makes all of this incremental and reversible rather than a risky big-bang cutover.
Identifying service boundaries
Use domain-driven design to find bounded contexts along business-capability lines (orders, billing, inventory) rather than along existing code-module lines, which usually reflect implementation history more than the current business domain. Within a candidate boundary, look for a vertical slice: API (application programming interface), business logic, and the data that logic owns, so the resulting service can operate independently without still calling back into the monolith for its own core data. Start with a boundary that is both low-risk (not on the critical write path) and clearly bounded (its data isn't heavily shared with other domains), since that combination validates the migration mechanics without betting the business on the first attempt.
Data integrity during migration
- Anti-corruption layer (ACL): a translation layer between the monolith's data model and the new service's API, so the new service isn't forced to adopt the monolith's legacy shape, and the monolith doesn't need to understand the new service's internals.
- Change data capture (CDC): stream changes from the monolith's database into the new service so reads stay consistent while write ownership is still transitioning.
- Dual writes, done carefully: if the monolith and the new service both need to write during the transition, make every write idempotent and reconcile the two stores regularly; prefer event-driven replication (the monolith publishes a change event that the new service consumes) over synchronous dual-writes, which fail if either side is briefly unavailable.
Avoiding distributed-transaction anti-patterns
Don't reach for two-phase commit (2PC) across the new service boundary. It requires every participant to be available and responsive during the commit window, which directly works against the availability and independent-deployability goals that motivated decomposition in the first place. Instead, model multi-step business flows as sagas: a sequence of local transactions, each with a defined compensating action if a later step fails, so the overall flow degrades gracefully instead of blocking on a distributed lock.
| Orchestration | Choreography | |
|---|---|---|
| Control flow | A central coordinator (a saga orchestrator) explicitly sequences each step and its compensation | Each service reacts to events from the previous step and publishes its own; no central coordinator |
| Best fit | Complex flows with non-trivial compensation logic, where you need one place to see and control the whole sequence | Simple, decoupled flows where steps genuinely don't need central sequencing |
| Observability | Easier: the orchestrator's state machine shows exactly where a given flow is | Harder: the flow is implicit in the pattern of events, so tracing a specific request end to end takes more tooling |
| Coupling | Services stay decoupled from each other but are coupled to the orchestrator's contract | Services are decoupled from any coordinator but implicitly coupled to the event schema and each other's reactions |
A common practical split is to orchestrate the core, high-stakes transaction (say, order to payment to shipment) where compensations matter, and let secondary, lower-stakes integrations (send a confirmation email, update an analytics feed) stay choreographed.
When to reach for the strangler fig pattern
Reach for it whenever you need the old and new implementations to coexist safely during a transition, which in practice is almost always for anything beyond a trivial, low-traffic domain. The mechanics, applied specifically at a monolith's gateway:
- Reverse-proxy routing: put a reverse proxy or API gateway in front of both the monolith and the new service, and route requests to one or the other by path, feature flag, or percentage, so the switch is a routing change rather than a deployment event.
- Traffic shadowing: mirror a copy of live production requests to the new service without using its response, so you can compare its behavior and performance against the monolith's real response under real traffic before the new service is trusted to actually serve anyone.
flowchart LR
C[Client] --> RP[Reverse proxy /<br/>API gateway]
RP -->|serves response| M[Monolith]
RP -.->|mirrored, response discarded| NS[New service]
RP -->|once validated,<br/>cut over by %| NS
Only retire the monolith's code path for a given slice once traffic has been fully cut over and the data behind it has been reconciled, not merely once the new service has been deployed.
Worked example: extracting a payments capability
Define the Payments bounded context and its API contract, then stand up an ACL that maps the monolith's existing payment calls onto that contract. Use CDC to populate the new Payments service's data store from the monolith's database so both stay consistent during the transition. Put a reverse proxy in front of the payments endpoints and start with traffic shadowing: every real payment request is also sent to the new service, its result compared against the monolith's, but only the monolith's response reaches the client. During this shadow period, downstream systems the payments path calls (a fraud check, a ledger write) effectively see close to double their steady-state call volume, since both the monolith and the shadowed new service are calling them, which is worth flagging to whoever owns that downstream capacity before shadowing begins. Once results match consistently, cut real traffic over gradually (a small percentage, then a majority, then all of it) via the reverse proxy, with the monolith's code path kept in place and able to take traffic back at any stage until the migration is fully reconciled.
Trade-offs and pitfalls
- Treating dual writes as sufficient data-integrity strategy on their own, without idempotency keys or a reconciliation job; a single missed or duplicated write during the transition window leaves the two stores silently inconsistent.
- Reaching for orchestration everywhere out of caution; it adds a coordinator every downstream service must trust and depend on, which is unnecessary overhead for flows that don't actually need centralized compensation logic.
- Shadowing traffic without accounting for its load on shared downstream dependencies, which can turn a validation exercise into a self-inflicted capacity incident.
- Declaring the migration done once the new service is receiving 100% of traffic, while the monolith's old code path and data are still live as a fallback; the migration isn't actually finished, and the risk isn't actually retired, until that fallback is deliberately removed.
Compare consistent hashing and range-based partitioning for a large-scale datastore where complex queries and joins across ranges are common. Explain the pros and cons of each for query locality, rebalancing, and ease of scaling, then propose a hybrid partitioning approach that supports complex queries without creating hotspots.
Sample Answer
Direct answer
Consistent hashing gives even load distribution and cheap rebalancing but destroys query locality, forcing every range scan or join into a scatter-gather across many nodes; range-based partitioning gives excellent locality for exactly those queries but creates hotspot risk on skewed keys and makes rebalancing expensive. For a datastore that genuinely needs both complex range queries and elastic scaling, the answer is a hybrid: partition coarsely by range on an attribute that preserves most query locality, then hash within each coarse range to spread load and avoid hotspots inside it.
Structured elaboration
What consistent hashing actually is, before comparing properties. Consistent hashing places both data keys and node identifiers as points on a circular numeric range, called a ring, typically by hashing each one with the same hash function. Each key is then owned by whichever node's point is the first one reached going clockwise from the key's own point on the ring. That single mechanism is what produces both properties compared below: rebalancing is cheap because adding or removing one node only reassigns the keys between it and its ring neighbors, not the whole dataset the way a plain modulo hash (hash(key) mod N) would when N changes; and locality is poor because two keys that are numerically close in the original ID space, like adjacent order IDs, get hashed to essentially random, unrelated points on the ring, so there is no guarantee they land near each other, or on the same node, at all.
Consistent hashing
- Query locality: poor. Related keys are scattered pseudo-randomly across nodes by design, so a range scan or a join across a key range has to fan out to many (potentially all) nodes and merge results, a scatter-gather pattern that is expensive and has a tail-latency cost equal to its slowest participant.
- Rebalancing: cheap. Adding or removing a node (especially with virtual nodes) only moves a small, bounded slice of the keyspace.
- Scaling: straightforward; new capacity absorbs a proportional share of load automatically.
Range-based partitioning
- Query locality: strong. Keys that are close in sort order live on the same node, so range scans and joins over a contiguous key range are single- or few-node operations, and storage engines benefit from sequential I/O and better compression on sorted data.
- Rebalancing: expensive. A skewed key distribution concentrates load on one range, and fixing it means splitting or merging large contiguous chunks of data, a much bigger operation than moving a handful of hash buckets.
- Scaling: harder to automate; adding capacity typically requires a deliberate split-planning step rather than capacity absorbing load on its own.
Hybrid approach
- Choose a coarse partitioning attribute that captures most of the locality your queries need (a time window, a tenant id range, a geographic region) and range-partition on it. This keeps queries that filter or join on that attribute confined to a small number of coarse partitions.
- Within each coarse range, sub-partition by consistent hashing (or a hash of a secondary key) into several sub-shards spread across physical nodes. This is what prevents a single popular coarse range from becoming a hotspot: instead of one range living on one node, it's spread across the sub-shard set.
- Maintain a lightweight metadata/routing layer mapping coarse range to its set of sub-shard endpoints, so the query planner knows which sub-shards a given range maps to without a full scatter.
- Handle skew with adaptive split/merge scoped to the coarse range: if one coarse range's sub-shards are overloaded, split that range's sub-shard count, an operation that only touches that range's data, not the whole dataset.
flowchart TD
Q[Incoming query: range filter] --> P[Router: which coarse ranges match?]
P --> R1[Coarse range 1]
P --> R2[Coarse range 2]
R1 --> H1[Hash sub-shard a]
R1 --> H2[Hash sub-shard b]
R2 --> H3[Hash sub-shard a]
R2 --> H4[Hash sub-shard b]
H1 --> M[Merge results]
H2 --> M
H3 --> M
H4 --> M
This bounds scatter-gather to the coarse ranges a query actually overlaps, times the sub-shard fan-out within each, instead of every shard in the cluster.
Worked example
A metadata/search index over 1 billion objects needs to support both point/range lookups by ingestion time and elastic scaling as volume grows. The team range-partitions coarsely by ingestion week, giving roughly 100 coarse ranges for two years of retained data (a stated design choice, not a derived figure), and hash-partitions each coarse range into 50 sub-shards to spread load evenly within the week, giving 100×50=5,000 total shards across the cluster. A typical query filtering on a 3-week window touches only the 3 matching coarse ranges, fanning out to 3×50=150 sub-shards, or 150/5,000=3% of the cluster, instead of a full scatter-gather across all 5,000 shards that pure consistent hashing on object id would require for the same query. A query that needs a single object by id, with no time filter, still does a targeted hash lookup within whichever coarse range the object's timestamp maps to, so point lookups stay single- or few-shard regardless of the hybrid layout.
Trade-offs & pitfalls
- The hybrid scheme adds a real second layer of metadata and routing logic (which coarse ranges exist, which sub-shards belong to each) that a pure hash or pure range scheme doesn't need; that complexity has to be justified by an actual mixed workload of range queries plus elastic-scaling needs, not adopted by default.
- Choosing the coarse partitioning attribute poorly (one that doesn't align with how queries actually filter) gives you the rebalancing cost of range partitioning without recovering the locality benefit that was the whole point of choosing it.
- A coarse range that itself becomes hot (a recent, actively-written time window, for instance) still needs the adaptive split/merge step; sub-sharding by hash inside it helps distribute load but doesn't eliminate the need to watch for a coarse range outgrowing its allocated sub-shard count.
- Cross-coarse-range joins (joining two objects that fall in different weeks, say) are not free just because the hybrid scheme handles range queries well; that join still fans out across whichever coarse ranges and sub-shards are involved, so it's worth being explicit about which query shapes the hybrid design is actually optimizing for.
Describe stateless versus stateful service designs, and explain why statelessness enables easier horizontal scaling. Include strategies to externalize state (databases, caches, session stores), and discuss scenarios where a stateful service is genuinely necessary, for example leader election or long-lived sticky connections.
Sample Answer
Direct answer
A stateless service keeps no client- or session-specific data in process memory between requests: every request carries (or looks up) everything needed to handle it. A stateful service holds that context in memory across requests, such as an open socket or an in-memory session. Statelessness enables horizontal scaling because any instance can serve any request, so a load balancer can distribute traffic with no per-node reconciliation and instances become disposable: add, remove, or replace them freely.
Structured elaboration
Why statelessness is the enabler, mechanically
- Interchangeability: since no instance holds unique data, a request routed to any instance gets the same result, which is what makes simple, even load distribution possible.
- Painless scaling events: adding capacity means starting new instances with no data migration; removing capacity means stopping instances with no data loss, because nothing lived there that mattered.
- Simpler recovery: a crashed stateless instance is replaced, not repaired; there is no in-memory state to reconstruct.
- Simpler rolling deploys: instances can be cycled one at a time without draining session state first.
Strategies to externalize state
- Durable application state moves to a database (relational or NoSQL) instead of process memory.
- Fast shared access to hot data moves to a distributed cache (an in-memory key-value store shared across instances) rather than a per-instance local cache.
- Session data specifically has two common paths: a centralized session store, or client-held tokens (such as a signed JSON Web Token, JWT) that let the server stay stateless entirely because the client presents its own state on every request.
- Large binary objects move to an object store rather than local disk.
- Asynchronous or queued state (work in flight, not yet durable) moves to a message queue.
- Small amounts of shared coordination state (configuration, leader pointers, locks) move to a dedicated coordination service built for that job.
Where a stateful service is genuinely necessary
| Scenario | Why statelessness does not fit | Typical mitigation |
|---|---|---|
| Leader election | A cluster needs exactly one active decision-maker at a time; that role is inherently shared, coordinated state, not a per-request fact | Use a purpose-built coordination service to hold and arbitrate the leader pointer, and design fast, automated re-election on failure |
| Long-lived connections (WebSocket, gRPC streams) | The connection itself is state; a mid-stream request cannot be freely handed to a different node without breaking the stream | Route stream traffic with connection affinity (a load-balancing mechanic, out of scope here) and design the client to reconnect and resume cleanly on node loss |
| Low-latency in-memory computation | Externalizing every read adds network latency that some workloads cannot absorb | Keep a local cache as a performance optimization layered on top of an externalized source of truth, not as the only copy |
| Stateful stream processing (windowed aggregation) | The computation's correctness depends on an accumulating window that must survive across events | Checkpoint local state to durable storage continuously so a replacement node can resume from the last checkpoint instead of from scratch |
The SRE angle: operating a stateful service you cannot avoid. When a stateful service is genuinely required, the operational burden shifts from "how do I scale it" to "how do I keep it reliable despite being pinned." That means treating the stateful nodes as a smaller, more carefully managed subset of the fleet: automated health checks tied to a fast failover or re-election path, regular drills that exercise that failover before it is needed under real incident pressure, and monitoring that specifically watches for the failure modes unique to statefulness (a stuck leader, a connection that never drains, a checkpoint that falls behind). The mitigation is never "make it stateless anyway"; it is "shrink the stateful surface to the minimum and instrument that minimum heavily."
Worked example
A web application currently stores logged-in session data in each server's memory, so a user's second request must land on the same server (sticky routing) or their session appears to vanish. To make the tier horizontally scalable: move session data out of process memory into a shared session store, or better, switch to a signed token that the client holds and presents on each request so the server does not need to look anything up at all. Either change means any instance can now serve any request, sticky routing is no longer required, and the fleet can be scaled up or down purely on load, with new instances immediately able to serve full traffic with zero data migration.
Trade-offs & pitfalls
- Do not confuse "stateless service" with "no state exists": the state still exists, it has just moved to a system designed to hold it reliably. That external system becomes a new dependency and a new potential bottleneck, so it needs its own scaling plan.
- Client-side tokens remove server lookups but push size and revocation concerns onto the client and the token design; server-side session stores keep revocation simple but add a network hop and a shared-store scaling problem.
- Treating "sticky sessions" as a free fix for statefulness is a common wrong turn: it works, but it silently reintroduces a form of per-node state ownership and undermines the failure-isolation benefit that horizontal scaling was meant to provide.
- For the genuinely stateful cases, skipping the failover-drill discipline is the most common production failure: the coordination logic works in testing and then fails silently the first time it is needed in an incident, because it was never exercised under realistic conditions.
Why does connection pooling matter for a service running at scale? Describe best practices for managing both database and HTTP connection pools: pool size, max open connections, idle timeouts, connection lifetime, and behavior under a spike in load. How would you test and tune these settings before production?
Sample Answer
Direct answer
Connection pooling matters at scale because opening a new database or HTTP connection is expensive relative to a request (TCP handshake, and for a database, authentication and session setup), so reusing a small set of warm connections instead of creating one per request lowers latency and prevents the backend from being overwhelmed by connection churn. The core sizing problem is that a pool is a per-instance setting but the backend has a fleet-wide connection ceiling, so pool size has to be planned across the whole fleet, not tuned in isolation on one instance.
Structured elaboration
Why pooling matters at scale, mechanically
- Connection setup cost: a TCP handshake, TLS negotiation (for HTTP), and for a database, authentication plus session/state initialization, all add latency if paid on every request.
- Backend resource limits: every open connection holds memory and, for a database, often a whole backend process or thread; a backend with a hard maximum connection count can be pushed into refusing connections or degrading badly under connection churn even if query volume itself is modest.
- Reuse turns a per-request cost into a one-time cost amortized across many requests on the same warm connection.
Sizing pools: the fleet-wide constraint
The number one mistake is sizing a pool as if the instance owns the whole backend. It does not; every other instance is drawing from the same ceiling:
pool_size_per_instance≤⌊number_of_app_instancesdb_max_connections⌋If the database allows 500 total connections and the service runs behind 20 instances, each instance's pool must stay at or below ⌊500/20⌋=25 connections, or a fleet at full pool utilization exceeds the database's ceiling and starts getting connection refusals, exactly when load is highest and refusals hurt the most. This constraint must be revisited every time the fleet is resized by autoscaling, which is the part teams most often forget: a pool size tuned for 20 instances silently becomes unsafe the moment autoscaling adds a 21st.
Core pool parameters
| Parameter | What it controls | Tuning guidance |
|---|---|---|
| Pool size (min/max) | How many connections are kept open per instance | Bounded above by the fleet-wide formula above; bounded below by enough to avoid queuing under normal load |
| Max open/concurrent connections | Hard ceiling the pool will not exceed even under burst demand | Set to protect the backend, not just to satisfy the busiest moment; excess demand should queue or fail fast, not force more connections open |
| Idle timeout | How long an unused connection stays open before being closed | Long enough to avoid re-opening connections for normal traffic gaps; short enough to release resources during genuine lulls |
| Max connection lifetime | Forces a connection to be recycled after a set duration regardless of use | Keeps the pool from silently holding stale or half-broken connections open indefinitely; also spreads out reconnections instead of all connections expiring together |
| Acquisition timeout | How long a request will wait for a pooled connection before failing | Should fail fast rather than block indefinitely, so an overload turns into fast, visible errors instead of a pile of hung requests |
Behavior under a load spike: the connection-storm problem
The specific failure mode worth naming: a deploy, a failover, or a sudden traffic spike can cause many instances to simultaneously reconnect or spin up new pooled connections at once, a connection storm, which can itself exceed the database's connection ceiling even though steady-state pool sizing was correct. This has a process/thread-model dimension too: a backend that spawns one OS process or thread per connection (a common relational-database architecture) pays a much higher per-connection memory and context-switch cost under a storm than one built around lightweight connection handling, which changes how conservatively you should size db_max_connections in the first place. Mitigations: stagger reconnects with jitter (small random delays) instead of reconnecting all instances at once, keep pool warm-up gradual rather than instantaneous on instance startup, and prefer acquisition timeouts with backoff over unbounded retry storms.
Testing and tuning before production
- Load-test at realistic peak concurrency and burst shape, not just average throughput, since spikes and connection storms are what actually break pool sizing.
- Vary the number of app instances in the test to confirm the fleet-wide formula holds at the target autoscaling range, not just at today's instance count.
- Watch active/idle/wait-count and wait-time metrics from the pool itself, plus backend-side connection and CPU/IO metrics, and tune size, idle timeout, and lifetime to minimize wait time while keeping the backend under its ceiling.
- Explicitly test the failure path: kill connections mid-flight, simulate a slow backend, and confirm acquisition timeouts and backpressure behave as designed rather than hanging.
Worked example
A service runs 20 instances against a database capped at 500 total connections. Using the formula above, each instance is capped at 25 pooled connections. During a load test that simulates a rolling deploy (all 20 instances restarting within a short window), every instance attempts to rebuild its pool of 25 connections at once: 20×25=500 simultaneous reconnect attempts against a ceiling of exactly 500, with zero margin for any connection still draining from the old instances. Adding jittered reconnect delays and reducing per-instance pool size to 20 (giving 20×20=400, leaving 100 connections of headroom during a rollover) eliminates the connection-storm failures observed in the unthrottled test.
Trade-offs & pitfalls
- Sizing a pool against a single instance's peak load, without dividing by the fleet size, is the most common and most damaging mistake; it works until autoscaling adds instances, then fails exactly under peak traffic.
- A pool with no acquisition timeout turns backend overload into cascading request pile-ups instead of fast, visible failures; for services making many short-lived connections, pairing the pool with a circuit breaker (a resilience pattern that stops sending requests to a struggling dependency, covered under high-availability patterns rather than here) prevents that pile-up from spreading further upstream.
- Idle timeouts set too aggressively cause needless reconnection churn during normal traffic dips; set too loosely, they let leaked or stale connections accumulate unnoticed.
- Connection leaks (code paths that acquire a connection and never release it, often on an error path) are the quiet failure mode: the pool looks correctly sized until leaked connections slowly starve it, and only a saturation metric with alerting catches this before an outage.
Unlock Full Question Bank
Get access to all Scalability Patterns and Techniques interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.