Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
Why are load balancers used in distributed systems? Describe at least four distinct problems they solve, and how they contribute to horizontal scaling, availability, and fault isolation. Where would you typically place load balancers in a three-tier architecture (edge, internal service, service-to-service)?
Sample Answer
Direct answer
Load balancers exist to decouple how much traffic arrives from which single machine has to answer it. They solve at least four distinct problems: spreading load across available capacity so no single server bottlenecks, removing single points of failure through health-based failover, enabling safe rollouts such as canary (releasing a new version to a small slice of traffic first) and blue-green deployment (running old and new versions side by side, then cutting all traffic over at once) with graceful draining, and centralizing cross-cutting concerns like TLS termination and rate limiting. In a typical three-tier architecture, load balancers show up at the edge (public-facing traffic), between internal services (service to service), and inside a service mesh (sidecar to sidecar).
Structured elaboration
Named problems load balancers solve:
- Single-server bottleneck / horizontal scaling. Spreading requests across many instances lets the system add capacity by adding instances instead of a bigger single machine.
- Availability and automatic failover. Health checks remove failing instances from rotation so one bad node does not take down the service.
- Fault isolation and safe rollout. Canary and blue-green routing, plus connection draining during deploys, contain a bad release to a small slice of traffic instead of every user.
- TLS termination and centralized security policy. Certificates and security rules (WAF, Web Application Firewall: a layer that inspects and blocks malicious HTTP traffic before it reaches a backend, and rate limiting) live in one place instead of being duplicated on every backend.
- Rate limiting and overload protection. The balancer can shed or throttle excess traffic before it reaches backend capacity, protecting downstream services.
Placement in a three-tier architecture:
graph LR
Client[Client] --> EdgeLB["Edge LB (L7, TLS termination)"]
EdgeLB --> GwA[Gateway instance A]
EdgeLB --> GwB[Gateway instance B]
GwA --> InternalLB[Internal service LB]
GwB --> InternalLB
InternalLB --> SvcA[Service instance A]
InternalLB --> SvcB[Service instance B]
SvcA --> MeshLB["Sidecar LB (service-to-service)"]
SvcB --> MeshLB
MeshLB --> Downstream[Downstream service]
The edge tier handles public traffic and usually terminates TLS. The internal service tier balances across replicas of a given service, often behind a service discovery layer. The service-to-service tier (frequently a sidecar, a small proxy process deployed alongside each service instance, in a mesh) balances outbound calls one internal service makes to another, applying the same health-check and retry logic without a separate hop through a centralized balancer.
A fast way to spot gaps in an unfamiliar system's traffic layer: check whether all three tiers actually exist (a single edge LB fanning straight out to every backend is a common shortcut that skips the internal tier), whether health checks are active at each tier, and whether TLS termination and rate limiting are centralized or scattered across services.
Worked example
Three backend instances, each rated for 1,000 requests per second, pooled behind a load balancer: combined capacity is 3×1000=3000 requests per second. If health checks detect one instance failing and pull it out of rotation, capacity becomes 2×1000=2000 requests per second, so the system keeps serving at 2000/3000≈66.7% of its original capacity instead of dropping to zero, which is what a single unpooled instance would do on failure.
Trade-offs & pitfalls
- The load balancer itself becomes critical infrastructure. A single-instance balancer just moves the single point of failure up one layer; it needs its own high availability (active-active pair, DNS or anycast failover).
- Every additional tier (edge, internal, mesh) adds a proxy hop's worth of processing and another thing to operate; do not add a tier a system does not need.
- A misconfigured health check at any tier can silently remove that tier's benefit (failover, safe rollout) even though the balancer process is technically running.
Explain the difference between readiness and liveness health checks (and startup checks, where supported). Design the checks for a database-backed web service (for example, covering DB connectivity and queue backlog), and explain how a misconfigured probe can cause cascading restarts or route traffic into a blackhole.
Sample Answer
Direct answer
Liveness answers "should this process be restarted"; readiness answers "should this instance receive traffic right now"; and a startup check, where the platform supports one, answers "has this instance finished initializing," so liveness does not kill a process that is simply still starting up. Keep liveness checks local and cheap, never calling the database or another network dependency, and keep readiness checks inclusive of whatever actually determines if this instance can do its job right now.
Structured elaboration
Designing checks for a database-backed web service:
- Liveness: process and event-loop responsiveness only (for example, can the process handle a trivial in-memory request). No database calls.
- Readiness: a bounded database ping, such as a single
SELECT 1with a tight timeout, so a slow or unreachable database marks the instance not-ready rather than hanging the check itself. - Readiness, if the service consumes a queue: a consumer-lag or backlog check, since a service that is technically up but has fallen far behind its queue is not really ready to take more work either.
- Startup check: if the platform supports one, use it to give the service extra time for one-time initialization (schema checks, cache warm-up, waiting on a migration) before liveness checks start at all, instead of stretching liveness's own thresholds to cover startup.
Pod lifecycle relative to these probes (SIGTERM below is the standard "please shut down" signal the orchestrator sends before forcibly killing a container):
stateDiagram-v2
[*] --> Starting
Starting --> NotReady: startup probe pending
NotReady --> Ready: readiness probe passes
Ready --> NotReady: readiness probe fails
Ready --> Restarting: liveness probe fails
NotReady --> Restarting: liveness probe fails
Restarting --> Starting: container restarted
Ready --> Draining: SIGTERM received
Draining --> [*]: in-flight requests finish
Worked example
A concrete configuration: readiness pings the database with a 100ms timeout and checks that consumer lag stays under a configured threshold; liveness only checks the process. Using the same worst-case reasoning as check-frequency tuning generally, with a check period of 10 seconds:
liveness worst-case restart trigger=period×failureThreshold=10×5=50 (seconds) readiness worst-case removal=period×failureThreshold=10×3=30 (seconds)Readiness is deliberately more sensitive (lower threshold) than liveness here, because pulling an instance out of rotation is cheap and reversible, while restarting a process is disruptive and should require stronger evidence.
How a misconfigured probe causes each failure mode, for this exact service:
- Cascading restarts: if liveness includes the database ping (against the design above) and the database has a brief outage, every replica fails liveness at roughly the same time. The orchestrator restarts the whole fleet concurrently, and the new pods immediately hit the same still-recovering database, failing liveness again, a self-inflicted restart storm that outlives the original database blip.
- Traffic blackhole: if readiness is too strict, or its database timeout is tighter than the database's real worst-case latency under load, a load spike can cause every pod to fail readiness at the same moment. The service is left with zero endpoints and every request gets a connection error, even though each pod was actually capable of serving traffic, just slower than the timeout allowed.
Trade-offs & pitfalls
- Never let liveness depend on an external service; a dependency outage should be a readiness problem (route traffic elsewhere) not a liveness problem (restart everything).
- Keep readiness dependency checks shallow and bounded (a ping, not a full transaction) so the check itself cannot become the slow thing that trips the threshold.
- Use disruption budgets or max-unavailable settings (Kubernetes controls that cap how many replicas can be down at once during a rollout or node drain) so an orchestrator responding to a wave of liveness failures cannot restart the entire fleet at once, which is what turns a brief dependency blip into an extended outage.
A downstream service starts returning intermittent errors, and your clients retry aggressively, overloading other components as a result. What is your immediate mitigation plan to stop the cascade, and what architecture changes would you make afterward to prevent it recurring?
Sample Answer
Direct answer
Stop the retry amplification first: shed load at the edge (cap or disable client retries, open a circuit breaker toward the failing dependency) so the retries stop compounding faster than the mitigation. Only once the bleeding is stopped do you dig into root cause. Afterward, the fix is structural: circuit breakers, bulkheads, and bounded retry budgets so this class of failure cannot cascade again, not just a one-off cleanup.
Structured elaboration
Immediate containment (first minutes):
- Cut the amplification loop, not the symptom. A retry storm is a positive feedback loop: errors trigger retries, retries add load, added load causes more errors. The fastest lever is at the edge (API gateway, service mesh, or client library), where you can globally cap or pause retries without a code deploy.
- Open a circuit breaker toward the failing downstream so callers fail fast (return a fast error or a cached/degraded response) instead of queuing and retrying against a dependency that is already struggling.
- Shed non-critical load: rate-limit or reject low-priority traffic first, protect the critical path.
- Make failure cheap for callers: return
503withRetry-Afterso well-behaved clients back off instead of retrying immediately.
Root-cause-directed long-term changes:
- Circuit breakers (client-side and at the gateway): trip on an error-rate threshold over a rolling window, fail fast while open, probe with limited traffic in half-open state before fully closing.
- Bulkheads: give each downstream dependency its own connection pool, thread pool, or queue so one degraded dependency cannot exhaust resources shared with healthy ones.
- Retry budgets with jittered exponential backoff: bound total retry traffic to a fixed fraction of primary request volume (a token-bucket style budget), require idempotency for anything retried, and always jitter backoff so retries do not resynchronize into a new spike.
- Backpressure (letting a struggling downstream signal or force upstream producers to slow down, rather than silently accepting more work than it can handle) over synchronous retries: for non-latency-critical paths, replace "retry immediately" with a durable queue so producers are decoupled from the downstream's instantaneous capacity.
Worked example
Retries are dangerous because they multiply load, and the multiplier grows non-linearly as the error rate rises. If a downstream fails a request with probability p and a client retries up to r times (each retry only happens after the previous attempt failed), the expected number of attempts per logical request is a finite geometric series:
E[attempts]=i=0∑rpi=1−p1−pr+1At a moderate error rate (p=0.3) with 3 retries:
E[attempts]p=0.3, r=3=1+0.3+0.09+0.027=1.417That is only 42% extra load. But during real distress (p=0.7, the same 3 retries):
E[attempts]p=0.7, r=3=1+0.7+0.49+0.343=2.533Load to the downstream more than doubles from the retries alone, on top of whatever caused the original errors. If retries are not capped at all, the series is unbounded and converges to:
r→∞limi=0∑rpi=1−p1so at p=0.9, an uncapped retrying client sends 10x its normal load into an already-failing dependency, which is exactly how a partial outage becomes a total one. And this compounds across a call chain: if 3 dependent hops each independently retry at p=0.3,r=3 (a 1.417x multiplier per hop), the multiplier reaching the origin is:
1.4173≈2.845This is why bulkheads are per-hop, not just at the edge: an unprotected chain amplifies retries multiplicatively with depth.
flowchart LR
Client -->|retries| GW[API Gateway]
GW --> RL[Rate limiter and retry budget]
RL --> CB[Circuit breaker]
CB -->|open: fail fast| Deg[Return 503 with Retry-After]
CB -->|closed: allow| BHA[Bulkhead: pool A]
CB -->|closed: allow| BHB[Bulkhead: pool B]
BHA --> Down[Downstream service]
BHB --> Other[Other backend]
Trade-offs and pitfalls
- Circuit breaker thresholds are a tuning problem, not a one-time setting. Too sensitive and you trip on normal noise (self-inflicted outage); too lax and it never protects anything. Validate thresholds against real traffic variance before trusting them in production.
- Bulkheads cost resources. Partitioning pools per dependency means some pools sit idle while others are saturated; this is a deliberate isolation-versus-utilization trade, not a free lunch.
- Retries must be idempotency-safe. A retry budget that resends non-idempotent writes turns a latency problem into a data-corruption problem; this has to be solved before any retry policy is trusted.
- Hidden retries are the most common miss. SDKs, ORMs, and infrastructure clients (HTTP clients, gRPC stubs, connection pools) often retry internally by default; a mitigation plan that only touches application-level retry code can still leave an amplification loop running underneath it.
- Jitter is not optional. Synchronized backoff without jitter just delays the storm and re-synchronizes it into a new spike at the same instant.
You observe rising p99 latency on your load balancer while backends show stable p95 latency and healthy CPU. Walk through a troubleshooting checklist covering the network, the LB proxies themselves, TLS handshakes, accept/queue backlogs, kernel limits, and client behavior. What instrumentation would you add to pinpoint the root cause?
Sample Answer
Direct answer
When the load balancer's p99 rises but the backends' p95 and CPU stay flat, the divergence is itself the clue: the extra tail latency is being added somewhere the backend can't see, the network path, the LB's own accept/TLS/queueing layer, or client behavior, not inside request processing. Work through the request path in layers (network, LB proxy internals, TLS, kernel accept/backlog, client) and instrument each layer's own latency contribution separately, so the tail latency is attributed to a stage, not guessed at.
Layered checklist
- Network path (client to LB, LB to backend): interface errors, drops, retransmits (
ip -s link,netstat -s), VPC flow logs for SYN/retransmit spikes, targeted traceroutes from affected client geographies. A tail-latency-only symptom with no backend involvement often starts here. - LB/proxy internals: connection counts and churn per backend, event-loop stalls or worker saturation in the proxy process itself (profiling,
straceon accept), time-to-first-byte versus time-to-last-byte split per backend. A proxy that's CPU-saturated or GC-pausing (if it's a managed-runtime proxy) shows up here, invisible to backend CPU metrics. - TLS handshakes: split full-handshake latency from resumed/session-ticket handshake latency, and track the resumption hit rate; a drop in session cache hit rate (cache eviction, a cache not shared across proxy instances, client churn) turns cheap resumed handshakes into expensive full ones for a subset of requests, exactly the shape of a tail-latency-only regression.
- Accept queue / kernel backlog: listen backlog occupancy (
ss -ltn), SYN_RECV counts (connections stuck mid-handshake, waiting on the final ACK),somaxconnandtcp_max_syn_backloglimits, ephemeral port and TIME_WAIT counts. If the accept queue is intermittently near its limit, new connections queue briefly even though every connection that does get through is processed at normal speed, invisible to backend request-processing metrics by construction. - Client behavior: slow or high-RTT clients, retry storms, or a small subset of very chatty clients; correlate p99 offenders by client IP, geography, or user agent rather than assuming they're uniform across all traffic.
Worked example: why queueing produces a p99 problem with a flat p95
This is a general queueing-theory property, illustrated with a simple M/M/1 model (queueing-theory shorthand: Markovian/memoryless request arrivals, Markovian/memoryless service times, 1 server) and pinned parameters, not a claim about any specific measured system. For a queue with service rate μ and utilization ρ=λ/μ (arrival rate λ), the expected wait time in queue is:
Wq(ρ)=μ(1−ρ)ρFix μ=1000 (an arbitrary service-rate unit, just for the shape of the curve) and evaluate at three utilizations:
ρ=0.80ρ=0.98ρ=0.995:Wq=1000×0.200.80=0.0040⇒4.00 ms:Wq=1000×0.020.98=0.0490⇒49.00 ms:Wq=1000×0.0050.995=0.1990⇒199.00 msA move from ρ=0.80 to ρ=0.995 (a 24% increase in load) inflates queueing wait by roughly 50x. The requests that land during the brief windows where the accept queue or a proxy worker pool is momentarily near saturation (micro-bursts, GC pauses, a slow client holding a worker) are exactly the ones that generate the p99 tail, while the bulk of requests, arriving when the system is comfortably under its knee, still finish fast and keep p95 (and backend CPU, which averages over time) looking healthy. This is why p99 and p95 can diverge sharply even though nothing about the backend's steady-state processing changed: the tail is a queueing phenomenon, not a processing-time phenomenon.
Instrumentation to add
- Break end-to-end latency into stages, TCP connect, TLS handshake, proxy accept/queue wait, backend processing, response write, and emit each as its own histogram (not just the total), tagged by backend and client region.
- Accept-queue depth and backlog saturation as a time series, not just a point-in-time check, so a transient near-saturation event that self-resolves in under a second is still visible.
- TLS session-resumption rate as its own metric, separate from handshake latency, since a resumption-rate drop is a leading indicator of the handshake-latency problem.
- Kernel-level counters (SYN_RECV: mid-handshake connections, TIME_WAIT: recently closed connections still held by the kernel, retransmits) exported alongside application metrics on the same dashboard and timeline, so a kernel-layer cause doesn't require manually correlating two separate tools during an incident.
- Sampled packet captures triggered automatically when p99 crosses a threshold, so there's raw evidence from the actual bad window instead of trying to reproduce it after the fact.
Trade-offs and pitfalls
- Chasing this in backend-only dashboards is the classic dead end, since by construction the backend's own view (CPU, p95, request-processing time) is healthy; the cause lives in a layer that doesn't report through the backend's own telemetry.
- Averages and even p95 hide exactly this kind of problem by design, since it only affects a small fraction of requests, that's what makes it a p99 problem and not a p50 problem; don't let a "p95 looks fine" dashboard close the investigation.
- Fixing the wrong layer (e.g. scaling backend CPU when the problem is TLS resumption cache eviction) burns real money and doesn't move the metric; stage-by-stage instrumentation exists specifically to prevent this kind of misdiagnosis.
- Once queueing is confirmed as the mechanism, the fix is capacity or admission control (a bigger backlog, more proxy workers, load shedding, or reducing ρ by scaling out), not code-level micro-optimization of request handling, since request handling was never where the time went.
Design a header-based routing framework at the edge load balancer that supports rules like 'route requests with header X-Canary=true to the canary pool' and 'route requests with a given X-Tenant-ID to a tenant-specific pool.' Cover rule storage and distribution, evaluation performance at the load balancer, conflict resolution between rules, and safe fallback behavior.
Sample Answer
Direct answer
Store rules as a versioned, prioritized ruleset in a central control plane, compile them into a fast local matcher (a trie, a tree structure where each path from the root spells out a prefix so a lookup finds a match in steps proportional to the key's length rather than scanning every rule, or decision graph, not a naive if-chain) on each edge node, and resolve conflicts with an explicit priority number plus a deterministic tie-break, never "whichever rule happens to match first" in file order. Safe fallback means every rule set has a default action if nothing matches or if validation fails, and that default is always the stable pool, never an unrouted request.
Structured elaboration
- Rule storage and distribution. Rules live in a control plane as immutable, versioned sets: predicate (header match, presence, regex, hash-bucket range), action (route to pool, set header), priority, and a version ID. The control plane pushes signed, compiled deltas to edge nodes (streaming or fast poll); each node keeps the last-known-good version and rolls back to it automatically if a new version fails local validation.
- Evaluation performance. Compile the ruleset into a structure suited to the predicate shape: a trie or hash lookup for exact-match headers like tenant ID, an Aho-Corasick automaton (a data structure that matches many string patterns against input in a single pass, faster than checking each pattern one at a time) or precompiled regex engine for pattern rules, and a short-circuit check for header presence before touching anything more expensive. The goal is that adding rules does not turn evaluation into a linear scan of every predicate on every request.
- Conflict resolution. Every rule carries an explicit priority number (lower evaluates first); ties break on rule creation time, then rule ID, so the outcome is fully deterministic and reproducible from the ruleset alone, not from evaluation order that happened to exist in memory.
- Header trust and spoofing. Only trust routing headers set by a component inside the trust boundary (the edge itself, after authentication), never a client-supplied header with the same name; if a header like X-Tenant-ID must originate from the client, validate it against the authenticated identity before using it for routing, or the routing layer becomes a way to reach another tenant's pool.
- Safe fallback. If no rule matches, or the active ruleset fails validation (schema error, missing referenced pool), route to the stable pool by default. This default must never be "no rule, no route."
graph LR
Author[Rule author: API/UI] --> RCP[Control plane: versioned ruleset]
RCP -->|signed compiled delta| Node1[Edge node: compiled matcher]
Req[Request] --> Node1
Node1 --> Match{Highest-priority match}
Match -->|tenant rule| TenantPool[Tenant-specific pool]
Match -->|canary rule| CanaryPool[Canary pool]
Match -->|no match / validation failed| Stable[Stable pool: default fallback]
Worked example
A simplified illustrative hash (real systems use SHA-256 or similar, not this): sum the ASCII codes of a key's characters and take the result mod 100. A canary rule routes to the canary pool when that value is below 10 (a 10% split).
| Key | ASCII sum | mod 100 | Below 10? | Routed to |
|---|---|---|---|---|
| "d" | 100 | 0 | yes | canary |
| "e" | 101 | 1 | yes | canary |
| "n" | 110 | 10 | no (boundary) | stable |
| "o" | 111 | 11 | no | stable |
The boundary case ("n", exactly 10) shows why the predicate needs an explicit, documented comparison (strictly less than, not less-than-or-equal): the threshold is a rule, not a fuzzy target, and the same key always lands on the same side of it, which is what lets a canary rollout stay stable across LB restarts without a shared session store.
Now layer conflict resolution: a request carries both a tenant header identifying "acme" (matches the tenant rule, priority 10) and a canary flag with a hash landing in the canary bucket (matches the canary rule, priority 20). Because lower priority numbers evaluate first, the tenant rule wins and the request goes to acme's dedicated pool, even though it would otherwise have qualified for canary.
Trade-offs & pitfalls
- Compiling to a trie or precompiled matcher costs build time on every rule update; for a control plane pushing frequent changes, measure compile time against your update frequency, not just steady-state lookup speed.
- Never let a client-controlled header carry routing authority without validating it against something the client cannot forge (an authenticated claim); an unvalidated tenant header is a direct tenant-isolation bypass, not just a routing bug.
- Rolling out a new ruleset version without a dry-run or shadow-evaluation stage means the first time you learn two rules conflict unexpectedly is in production; validate new versions against recent real traffic before activating them.
- Common wrong turn: treating "most specific rule wins" as if it were the same as explicit priority. Specificity is a heuristic humans use when writing rules; the engine needs an explicit, deterministic number, or two engineers' intuitions about "more specific" will eventually disagree.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.