Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
Compare the common load balancing algorithms: round robin, weighted round robin, least connections, and consistent hashing. For each, explain how it behaves under variable request latencies, long-lived connections such as WebSockets, and heterogeneous backend capacity, and give one production use case and one drawback. When would health checks and session affinity change which algorithm you pick?
Sample Answer
Direct answer
Round robin, weighted round robin, least connections, and consistent hashing differ mainly in what signal they use to pick a backend: round robin uses none (pure sequence), weighted round robin uses a static capacity ratio, least connections uses live connection counts, and consistent hashing uses a deterministic key (session id, user id) rather than load at all. That difference in signal is exactly why they behave so differently under variable request latency, long-lived connections, and mixed backend capacity.
Structured elaboration
| Algorithm | Behavior | Variable request latency | Long-lived connections (WebSockets) | Heterogeneous capacity | Production use case | Main drawback |
|---|---|---|---|---|---|---|
| Round robin | Cycles requests through servers in fixed order | Poor: ignores how long a prior request is taking | Poor: a server holding several long connections still gets the next new one | Poor: treats all servers as equal capacity | Homogeneous fleet, short uniform requests (static assets) | Blind to actual load; one slow server keeps getting new work |
| Weighted round robin | Same cycle, but servers get requests proportional to a static weight | Same blind spot as round robin, just scaled by weight | Same as round robin, scaled by weight | Good, if weights reflect real capacity | Mixed instance sizes with known, stable capacity ratios | Weights are static; does not react to a live spike |
| Least connections | New request goes to the server with fewest active connections | Good: naturally favors servers not stuck on slow requests | Good: avoids piling more work on a server already holding many long sessions | Needs weighting; assumes equal cost per connection otherwise | Variable-duration requests, streaming, WebSocket-heavy services | Needs an accurate, ideally shared, connection count; degrades to round robin if counts are only local per LB replica |
| Consistent hashing | Deterministic key (e.g. session id) maps to a server via a hash ring (both servers and keys are hashed onto points on a circle; a key is served by whichever server's point comes first walking clockwise from the key's own point) | Neutral: does not react to load at all | Excellent: same client always lands on the same server without a central session store | Needs virtual nodes weighted by capacity | Caches, CDNs, or any workload that needs strong affinity without a session store | Balances key placement, not load; a single hot key still lands entirely on one node |
When health checks and session affinity change the pick: if backends must hold session state without an external store, consistent hashing (or cookie-based affinity layered on any of the other three) becomes close to mandatory regardless of its load-balancing weaknesses. If health checks are slow to detect a failing node, least connections is more forgiving than round robin because it naturally routes fewer new requests to a node that is already accumulating unanswered connections, buying time until the health check catches up.
Worked example
Two backends, A and B, where A has twice the capacity of B (weight ratio 2:1). A weighted round robin scheduler realizing that ratio with a repeating pattern of period 3, "A, A, B", sent over 9 requests: A, A, B, A, A, B, A, A, B. Counting assignments: A receives requests 1, 2, 4, 5, 7, 8 (6 total), B receives requests 3, 6, 9 (3 total). That is 6:3, which simplifies to 2:1, exactly matching the configured weight, with no need to observe live load at all.
Consistent hashing, traced the same way: using a simplified illustrative hash (real systems use SHA-256 or similar, not this), sum a key's ASCII codes and take the result mod 100 to place it on a 0-99 ring. Three servers sit at fixed ring positions: X = 25, Y = 60, Z = 90 (a key is served by the first server whose position is at or after the key's hash, walking clockwise and wrapping to X if the hash is past the last position).
| Key | ASCII sum | mod 100 | Routed to |
|---|---|---|---|
| "session-42" | 919 | 19 | X (25) |
| "session-77" | 927 | 27 | Y (60) |
| "session-3" | 868 | 68 | Z (90) |
Now remove X from the ring (a node failure or a planned decommission): "session-42" (hash 19) has no server at or after position 19 until Y at 60, so it remaps from X to Y. "session-77" and "session-3" are untouched, since Y and Z were never X's successor for their hashes. This is the traced version of the bounded-remap claim in the table above: removing one of three nodes only disturbs the keys that node actually owned, not the whole ring.
Trade-offs & pitfalls
- Least connections needs an accurate, ideally shared, view of active connections per backend. Behind multiple load balancer replicas each keeping its own local count, the counts are only approximations, and the algorithm's advantage over round robin shrinks.
- Consistent hashing balances key placement, not load. A single popular key still lands entirely on one node no matter how good the hash function is; that is a cache-hot-key problem, not something the algorithm fixes.
- Common wrong turn: reaching for consistent hashing "for scalability" when the actual requirement is even load distribution. It solves affinity and cache locality, not general load balancing; without capacity-weighted virtual nodes it can be less balanced than plain least connections.
Design a monitoring and alerting setup for a load balancer tier so you catch unhealthy backends before users see errors. Which metrics would you include and why, what alert thresholds would you propose, and how would you differentiate severity to avoid alert fatigue?
Sample Answer
Direct answer
Catching unhealthy backends before users notice means monitoring the LB tier along four angles: traffic, latency, and errors as seen from the LB itself (the RED metrics), backend health signals independent of user-facing symptoms, saturation and capacity of the LB tier, and TLS/connection-layer health, then tying alert severity to how much of the user-facing error budget (the amount of failure your reliability target still allows before you've broken your promise to users) is actually being consumed rather than to a single static threshold, so a brief blip and a sustained outage don't page the same way.
The metric set
| Category | Metric | Why it matters |
|---|---|---|
| Traffic | Requests per second, overall and per backend pool | Baseline for interpreting every other metric; a rate change can explain an error-rate change |
| Latency | p50 / p90 / p95 / p99, split into LB-processing time and end-to-end time | Percentiles catch tail degradation that an average hides; splitting LB time from backend time tells you where the slowdown is |
| Errors | 4xx vs 5xx rate, and whether 5xx originated at the LB (no healthy backend) or was passed through from a backend | A 5xx-at-the-LB spike usually means capacity loss, not an application bug |
| Backend health | Per-backend health-check pass/fail rate and consecutive-failure counts | The direct signal for whether a specific node is about to be pulled |
| Connections | Active connections and connection churn (opens/closes per second) | A churn spike often precedes a latency or error spike |
| Capacity | LB-tier CPU, memory, socket/file-descriptor usage | The LB tier itself can become the bottleneck, not just the backends |
| TLS | Handshake failure rate, certificate expiry countdown | Silent failure mode: expired certs or handshake issues look like "backend is fine but nobody can connect" |
| Traffic shaping | Rate-limit rejections, circuit-breaker open/close events | Leading indicators: a circuit breaker tripping on a backend is an early warning before that backend fails health checks outright |
Alert severity: tie it to error-budget burn, not just a raw threshold
A raw threshold ("page if 5xx rate > 1%") treats a 10-minute blip the same as a slow multi-hour decline, a common cause of alert fatigue in either direction: it either pages on noise or is set so loose it misses real problems. The more robust approach expresses the threshold as a multiple of your error-budget burn rate: if you're consuming your monthly error budget faster than the rate at which "consume it evenly over 30 days" would, page proportionally to how much faster.
Worked example
Suppose the LB tier's SLO (service-level objective, the reliability target you've committed to, here 99.9% of requests succeeding) is 99.9% success rate over a 30-day window. The allowed error budget is 0.1% of requests, equivalently, 0.1% of the 30-day window can be fully down and still meet the SLO:
- 30 days in minutes: 30×24×60=43,200 minutes.
- Budget: 43,200×0.001=43.2 minutes of full-downtime-equivalent for the entire month.
Now suppose the current 5xx rate over the last hour is holding at 2%, against an objective of 0.1%. The burn rate is:
burn rate=0.1%2%=20×At 20 times the sustainable rate, the entire month's budget would be exhausted in:
2030 days=1.5 daysA burn rate that would exhaust the monthly budget in 1.5 days is unambiguously page-worthy; a burn rate of, say, 1.2 times sustainable is not, it just needs tracking. That is the difference between a severity-1 page and a low-priority ticket, expressed in the same unit (budget consumption), not two arbitrarily different percentage cutoffs.
Avoiding alert fatigue
- Tier severity by projected budget-exhaustion time (minutes-to-exhaust pages immediately, weeks-to-exhaust becomes a ticket), rather than by a single static percentage.
- Deduplicate and correlate: a health-check-flapping alert and a 5xx-rate alert firing from the same underlying cause should collapse into one incident, not page twice.
- Route by audience: on-call needs the actionable, current-state view (which backend, since when); a weekly review needs the trend (is health-check-induced churn increasing over time).
- Require every page-level alert to link directly to a runbook step; an alert with no clear next action trains people to ignore it.
Trade-offs and pitfalls
Over-indexing on system-level metrics (CPU, connections) without a user-facing correlate risks paging on things users never notice; conversely, only alerting on aggregate error rate can hide a problem affecting one backend pool or one tenant while the aggregate stays healthy. Circuit-breaker and rate-limit rejection counts are easy to skip because they're not "errors" in the traditional sense, but they're often the earliest signal, by the time health checks start failing outright, the circuit breaker has usually already tripped.
What observability signals (metrics, logs, traces) would justify automatically removing an instance from load balancer rotation before users see errors? For each signal, describe roughly what threshold or pattern would trigger removal and one way that signal could produce a false positive.
Sample Answer
Direct Answer
Pull an instance automatically only when a signal shows a sustained departure from its own steady-state baseline, smoothed over a short window, and prefer requiring at least two independent signals to agree before removing an instance, rather than trusting any single metric. The one exception is a hard, unambiguous signal, like a completely unresponsive health-check endpoint, which should short-circuit straight to removal since there's nothing left to corroborate.
Signals, Triggers, and False-Positive Modes
The table below spans all three signal categories the question names, metrics, logs, and traces, not just the metrics-heavy ones that are easiest to instrument first.
| Signal | Example trigger pattern | A way it produces a false positive |
|---|---|---|
| Error rate (5xx / exceptions) | Sustained error rate over 5% for 2 minutes, or a spike over 20% for 30 seconds | A misbehaving upstream client, or a brief warm-up period right after deploy |
| Latency / SLO breach | p95 over the SLO (service-level objective, the reliability target you've committed to) target for 3 consecutive windows | A GC pause during a low-traffic period, or one unusually large request skewing a small sample |
| Failed readiness probes | 3 consecutive failed probes at a 10s interval | A brief network blip between the load balancer and the instance, unrelated to the instance's actual health |
| Resource exhaustion (CPU, memory, file descriptors) | CPU over 90% for 5+ minutes, or FD usage over 95% | A scheduled batch job or legitimate traffic surge, not an actual degradation |
| Connection/thread pool saturation | Pool at max for 2+ minutes with a growing queue | A slow downstream dependency backing up the queue, not a fault in this instance itself |
| Downstream timeout rate (from traces) | Over 10% of traces show upstream-to-downstream timeouts in a 5 minute window | A scheduled maintenance window on the downstream service, not a problem with this instance |
| Structured error-log pattern (logs) | A specific stack-trace signature, or a rate of "connection refused" / "out of memory" log lines, exceeding N occurrences within a 2-minute rolling window | A noisy but harmless dependency that logs every retry attempt as an error line even though the retry itself eventually succeeds, inflating the log-based count without any real user-facing failure |
Smoothing and Windowing
Use a rolling window (not an instantaneous sample) for every signal except hard health-check failure, and require the trigger condition to hold across the whole window rather than at a single point. This is what separates "sustained degradation" from "one bad request." The window length is a direct trade-off against detection speed: shorter windows catch real problems faster but inherit more noise from small sample sizes, especially on low-traffic instances where a 2-minute window might contain only a handful of requests.
Combining Signals to Cut False Positives
Requiring two independent signals to agree before acting cuts the false-positive rate multiplicatively, at the cost of occasionally missing a real problem that only shows up in one signal. Suppose each of three independent signals (error rate, latency, saturation) has an independent 1% chance of a false trigger in any given window (p=0.01). Requiring any 2 of 3 to agree:
P(at least 2 of 3 trigger)=(23)p2(1−p)+(33)p3 =3×0.0001×0.99+0.000001=0.000297+0.000001=0.000298Compared to acting on any single signal alone (p=0.01), that's a reduction of:
0.0002980.01≈33.6×fewer false removals, assuming the signals really are independent. In practice error rate and latency are often correlated (a struggling instance tends to show both), so the real-world reduction is smaller than this idealized number, but the direction and the general logic hold: corroboration is strictly cheaper in false positives than any single signal.
Trade-offs and Pitfalls
- A hard signal (health-check endpoint completely unreachable) should bypass the corroboration requirement entirely; waiting for a second signal to agree that a dead instance is dead only adds latency for no benefit.
- Combining signals reduces false positives but can delay detection of a real problem that manifests in only one signal strongly (say, a memory leak that spikes memory well before latency degrades); the corroboration requirement is a precision/recall trade (catching more real problems, higher recall, versus fewer false alarms, higher precision, pull in opposite directions), not a free lunch.
- Automated removal needs its own safety valve: a cap on how many instances can be removed in a short window, so a correlated event (a bad deploy, a shared dependency outage) that trips the same signal across many instances at once doesn't remove enough capacity to cause the outage it was meant to prevent.
- Removal should default to a graceful drain (stop new traffic, let in-flight requests finish) rather than an immediate hard kill, except when the health signal indicates the instance can no longer serve any traffic at all.
Describe how you would coordinate load balancer connection draining with autoscaling group scale-in so instances finish in-flight work before termination. Cover cooldowns, warm-up periods, and how the load balancer's health-check readiness state should gate when a new instance starts receiving traffic.
Sample Answer
Direct answer
The load balancer must gate two moments in an instance's life: it should not receive traffic until its readiness check passes (not just a process-alive check), and on scale-in it must stop receiving new connections while letting in-flight requests finish, deregistration with connection draining, before the instance is terminated. Both are the LB's job, independent of whatever policy decided to scale in the first place.
Structured elaboration
Scale-out: readiness gates traffic, not liveness. A liveness check ("is the process up") and a readiness check ("is the app actually able to serve correctly") are different questions. The LB should register the instance in rotation, or the orchestrator should mark it Ready, only after readiness passes: connection pools opened, local caches warmed enough to avoid a cold-cache latency cliff, and any required config pulled. A liveness-only gate puts unwarmed instances into rotation and produces a latency or error spike on every scale-out event.
Scale-in: draining before termination, in order:
- Instance is marked for termination (by the autoscaler or an operator).
- The LB immediately stops sending new connections to it (deregister from rotation) but does not kill existing connections.
- A deregistration delay (connection draining window) gives in-flight requests time to complete normally.
- Only after the drain window elapses, or all in-flight requests finish, whichever is defined as the policy, does the instance actually terminate.
Sizing the drain window. It must exceed the longest realistic in-flight request duration for the service, with margin, or draining will forcibly cut live requests, defeating the point.
Health-check readiness during warm-up. The initial health-check grace period (the time before a failed check counts against a brand-new instance) must exceed real startup time, so a slow-booting instance is not falsely marked unhealthy and cycled before it ever gets a chance to serve.
Cooldowns, and where autoscaling policy stops and LB mechanics start. Which metric and threshold trigger a scale event is an autoscaling-policy decision outside this question's scope, but the cooldown itself belongs here since the question names it directly: a cooldown is the minimum time an autoscaler must wait after one scaling action before it is allowed to trigger another, and it exists because the fleet's aggregate metrics take time to reflect the previous action, an autoscaler that reacts to a still-stale signal will overreact, stacking a second scale-out on top of one still warming up, or yanking capacity out from under a scale-in that hasn't finished draining. The cooldown has to be sized against the LB mechanics above, not chosen independently of them: it should exceed both the drain window (a scale-in's capacity isn't actually gone until draining finishes) and the health-check grace period (a scale-out's capacity isn't actually serving, and won't show up in the fleet's metrics, until warm-up and readiness finish), see the worked example below for a concrete number. Whatever triggers a scale-out or scale-in event, the LB's readiness gate and drain sequence are what make that event safe for in-flight traffic; the cooldown is what keeps the autoscaler from firing faster than those mechanics can complete.
Worked example
Say p99 request duration for this service, measured over the last observation period, is 8 seconds, and cold-start warm-up (open DB pool + populate local cache) takes 15 seconds.
Deregistration delay, with margin over the slowest realistic request:
drain window=8s×1.25≈10sHealth-check grace period, with margin over startup time:
grace period=15s+10s (buffer)=25s, rounded up to 30sSo: a new instance is excluded from the health check's failure count for its first 30 seconds while it warms up, and a terminating instance gets up to 10 seconds to finish in-flight requests after being pulled from rotation, before the process is killed.
Cooldown, sized to exceed both of those numbers so the autoscaler never acts on a stale signal:
cooldown=max(drain window,grace period)×1.5≈30s×1.5=45sSo the autoscaler waits at least 45 seconds after any scale-out or scale-in action before it will consider triggering another one, comfortably longer than the 30 second grace period a new instance needs before its readiness, and therefore its true added capacity, is even visible in the fleet's metrics.
Trade-offs and pitfalls
- Too-short drain windows silently drop real user requests mid-response; this is one of the most common causes of a small, hard-to-reproduce error-rate blip during routine scale-in.
- Too-long drain windows slow down scale-in, which matters when scale-in is reacting to a cost or capacity signal; there's a real tension between "never cut a request" and "shed capacity promptly."
- A liveness-only readiness gate causes a recurring latency spike on every scale-out, because unwarmed instances start serving immediately; teams often discover this only when scale-out frequency increases (e.g. spiky traffic) and the spikes become visible in aggregate latency graphs.
- Grace periods that are too short flap healthy-but-slow-booting instances, causing the autoscaler and LB to fight each other: instance boots, fails health check before finishing warm-up, gets cycled, repeats.
- Drain and readiness are load-balancer-owned; they are not a substitute for correct autoscaling trigger design. A perfectly tuned drain window cannot fix an autoscaler that scales in too aggressively in the first place, that is a separate, upstream problem.
Design a Layer 7 load balancer that provides session affinity using consistent hashing on a session cookie. It must support health checks and rebalance sessions gracefully when nodes are added or removed. Discuss hash ring maintenance, virtual nodes, and how you would drain and migrate sessions without dropping in-flight traffic.
Sample Answer
Direct Answer
Hash the session cookie onto a consistent-hash ring of virtual nodes so a session's requests keep landing on the same backend under normal conditions, publish the ring from a small versioned control plane so every proxy agrees on current ownership, pull unhealthy nodes out via active health checks, and when a node is intentionally removed, keep it serving its existing sessions for a bounded drain window while new sessions route elsewhere, only decommissioning it once its active-session count reaches zero or the window expires. The two things that make this safe are versioning the ring (so proxies never disagree mid-transition) and treating removal as a drain, not a hard cut.
Architecture
flowchart LR
Client -->|cookie + ring_version| Proxy[Edge Proxy]
Proxy -->|lookup| RingSvc[Ring Metadata Service]
HealthChecker[Health Checker] -->|mark unhealthy| RingSvc
Proxy --> Backend1[Backend 1]
Proxy --> Backend2[Backend 2]
Proxy -.draining.-> Backend3[Backend 3]
DrainOrch[Drain Orchestrator] -->|coordinate| RingSvc
Backend3 -->|flush session state| Store[(Shared Session Store)]
Backend1 --> Store
Backend2 --> Store
The proxies are stateless: all ring state lives in the ring metadata service and is cached locally with a version number. Health checks and the drain orchestrator only ever mutate the ring through that service, never by having a proxy make a unilateral decision.
Hash Ring and Virtual Nodes
- Use a large hash space (64-bit) with each physical backend assigned many virtual nodes, commonly in the range of 100 to 500, so that any single node's removal spreads its load across many different successors instead of dumping it on one (see the worked example below for why this matters).
- Store the ring as a sorted array in the metadata service; proxies cache a copy and watch for version bumps rather than polling on every request.
- Every ring mutation (join, leave, health-driven removal) increments the ring version atomically, so a proxy can always tell whether its cached copy is current.
Cookie and Version Handling
The session cookie carries three fields: the session id (the hash key), the ring version the session was originally assigned under, and an HMAC (a keyed cryptographic checksum: proof the cookie's contents haven't been altered since a server signed them) to prevent tampering. On a new session, the proxy issues a fresh cookie against the current ring version. On a returning session, the proxy honors the session's recorded ring version for a bounded overlap window even after the ring has moved on, which is what lets an in-flight session keep reaching its original backend during a drain instead of being silently reassigned mid-conversation.
Health Checks
Active HTTP health checks run against each backend on a fixed interval. A backend that fails its threshold is marked unhealthy in the ring metadata service, which bumps the ring version; proxies pick up the change and stop routing new sessions there. Existing sessions already pinned to that backend via cookie should still be judged against the same health signal: if the backend is actually down, in-flight requests will fail regardless of stickiness, so unhealthy removal is immediate, not drained (draining is reserved for planned, intentional removal, covered next).
Graceful Drain and Session Migration
For a planned node removal (scale-down, deploy, decommission):
- Mark the node "draining" in the ring metadata service.
- Remove the node's virtual-node entries from the ring and bump the ring version.
- Proxies pick up the new version for new sessions immediately; sessions whose cookie still carries the old ring version continue routing to the draining node for a bounded overlap window.
- During the overlap window, the draining node's session state is either already externalized (shared store, no action needed) or actively migrated: the drain orchestrator streams in-memory session state to the new owning node or to the shared store.
- Once the draining node's active-session count reaches zero, or the overlap window expires (whichever comes first), the node is decommissioned. Sessions that hadn't finished by then either see a session reset (acceptable for many applications) or, if externalized state was used, resume transparently on the new node.
For node addition, the reverse: add the virtual nodes, bump the version, and let new sessions start flowing to the new node immediately. Existing sessions are unaffected because their arcs on the ring did not move.
Session State Strategy
Prefer an external session store (a clustered cache such as Redis) over in-memory backend state whenever possible: it decouples session survival from any single backend's lifecycle entirely, removing the need for the migration step above. If in-memory sessions are unavoidable (e.g., a stateful protocol upgrade like a long-lived WebSocket), the migration path in step 4 is mandatory, not optional.
Worked Example: How Much Load a Removal Redistributes
Take 20 physical backends, each with V=150 virtual nodes, so T=3000 ring tokens total. Removing one physical backend removes its 150 tokens; each of those tokens' clockwise successors is, in a well-shuffled ring, effectively a random draw among the other 19 physical nodes. So the removed node's 201=5% share of keys spreads across roughly 19 recipients rather than one. Each surviving node's expected new share:
201+20×191=38019+1=191≈5.263%which is exactly the uniform share you'd expect from evenly splitting the whole keyspace across 19 nodes, the same identity that shows up whenever virtual-node count is high enough to approximate rendezvous hashing's (an alternative hashing scheme that scores every node directly against each key and always redistributes evenly on removal, without needing virtual nodes) exact uniform redistribution. Compare that to a single-token ring (no virtual nodes) removing one of 20 nodes: the removed node's entire 5% lands on one successor, whose share jumps from 5% to 10%, a 2x hotspot instead of a 0.26-point bump.
Trade-offs and Pitfalls
- The ring metadata service is a coordination dependency; if it uses a consensus protocol (Raft or similar) for consistency, that adds latency to ring mutations (acceptable, since mutations are rare) but also means the service itself needs its own HA story.
- Overlap window length is a direct trade-off between session continuity and routing complexity: a longer window keeps more in-flight sessions alive during churn but means proxies must correctly honor two ring versions simultaneously for longer.
- Signed ring-version cookies need HMAC key rotation handled carefully: rotating the signing key while old cookies are still in flight requires accepting both old and new keys for a transition period, or sessions get silently invalidated.
- If a CDN or edge cache sits in front of this layer, cache keys for session-specific content must not be shared across ring reassignment; either scope those responses as non-cacheable or ensure the cache key includes a stable shard identifier that survives rebalancing.
- A single very active session ("hot session") can skew load on whichever node currently owns it; virtual nodes fix cluster-wide statistical balance, not a single oversized key, so hot-key handling needs a separate mitigation (e.g., splitting that session's read traffic) if it becomes a real problem.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.