Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
What observability signals (metrics, logs, traces) would justify automatically removing an instance from load balancer rotation before users see errors? For each signal, describe roughly what threshold or pattern would trigger removal and one way that signal could produce a false positive.
Sample Answer
Direct Answer
Pull an instance automatically only when a signal shows a sustained departure from its own steady-state baseline, smoothed over a short window, and prefer requiring at least two independent signals to agree before removing an instance, rather than trusting any single metric. The one exception is a hard, unambiguous signal, like a completely unresponsive health-check endpoint, which should short-circuit straight to removal since there's nothing left to corroborate.
Signals, Triggers, and False-Positive Modes
The table below spans all three signal categories the question names, metrics, logs, and traces, not just the metrics-heavy ones that are easiest to instrument first.
| Signal | Example trigger pattern | A way it produces a false positive |
|---|---|---|
| Error rate (5xx / exceptions) | Sustained error rate over 5% for 2 minutes, or a spike over 20% for 30 seconds | A misbehaving upstream client, or a brief warm-up period right after deploy |
| Latency / SLO breach | p95 over the SLO (service-level objective, the reliability target you've committed to) target for 3 consecutive windows | A GC pause during a low-traffic period, or one unusually large request skewing a small sample |
| Failed readiness probes | 3 consecutive failed probes at a 10s interval | A brief network blip between the load balancer and the instance, unrelated to the instance's actual health |
| Resource exhaustion (CPU, memory, file descriptors) | CPU over 90% for 5+ minutes, or FD usage over 95% | A scheduled batch job or legitimate traffic surge, not an actual degradation |
| Connection/thread pool saturation | Pool at max for 2+ minutes with a growing queue | A slow downstream dependency backing up the queue, not a fault in this instance itself |
| Downstream timeout rate (from traces) | Over 10% of traces show upstream-to-downstream timeouts in a 5 minute window | A scheduled maintenance window on the downstream service, not a problem with this instance |
| Structured error-log pattern (logs) | A specific stack-trace signature, or a rate of "connection refused" / "out of memory" log lines, exceeding N occurrences within a 2-minute rolling window | A noisy but harmless dependency that logs every retry attempt as an error line even though the retry itself eventually succeeds, inflating the log-based count without any real user-facing failure |
Smoothing and Windowing
Use a rolling window (not an instantaneous sample) for every signal except hard health-check failure, and require the trigger condition to hold across the whole window rather than at a single point. This is what separates "sustained degradation" from "one bad request." The window length is a direct trade-off against detection speed: shorter windows catch real problems faster but inherit more noise from small sample sizes, especially on low-traffic instances where a 2-minute window might contain only a handful of requests.
Combining Signals to Cut False Positives
Requiring two independent signals to agree before acting cuts the false-positive rate multiplicatively, at the cost of occasionally missing a real problem that only shows up in one signal. Suppose each of three independent signals (error rate, latency, saturation) has an independent 1% chance of a false trigger in any given window (p=0.01). Requiring any 2 of 3 to agree:
P(at least 2 of 3 trigger)=(23)p2(1−p)+(33)p3 =3×0.0001×0.99+0.000001=0.000297+0.000001=0.000298Compared to acting on any single signal alone (p=0.01), that's a reduction of:
0.0002980.01≈33.6×fewer false removals, assuming the signals really are independent. In practice error rate and latency are often correlated (a struggling instance tends to show both), so the real-world reduction is smaller than this idealized number, but the direction and the general logic hold: corroboration is strictly cheaper in false positives than any single signal.
Trade-offs and Pitfalls
- A hard signal (health-check endpoint completely unreachable) should bypass the corroboration requirement entirely; waiting for a second signal to agree that a dead instance is dead only adds latency for no benefit.
- Combining signals reduces false positives but can delay detection of a real problem that manifests in only one signal strongly (say, a memory leak that spikes memory well before latency degrades); the corroboration requirement is a precision/recall trade (catching more real problems, higher recall, versus fewer false alarms, higher precision, pull in opposite directions), not a free lunch.
- Automated removal needs its own safety valve: a cap on how many instances can be removed in a short window, so a correlated event (a bad deploy, a shared dependency outage) that trips the same signal across many instances at once doesn't remove enough capacity to cause the outage it was meant to prevent.
- Removal should default to a graceful drain (stop new traffic, let in-flight requests finish) rather than an immediate hard kill, except when the health signal indicates the instance can no longer serve any traffic at all.
Design an automated control loop that shifts a percentage of production traffic to a canary over a fixed window and rolls back automatically on an SLO breach. Cover how you would smooth the weight changes, what guardrails you'd set (minimum observation windows, maximum error thresholds), and how the rollback itself executes quickly and safely.
Sample Answer
Direct answer
An automated canary control loop is a periodic process that increases a canary's traffic weight in small, guarded steps, checks health metrics against the baseline after every step using a minimum-observation window, and holds or immediately rolls back to zero on an SLO breach (a violation of the service-level objective, the reliability target you've committed to, such as 99.9% success rate). The design problem is really about the guardrails and the rollback path, not the ramp itself.
Structured elaboration
Components:
- A traffic router (Envoy, an ingress controller, or a cloud LB) that supports weighted routing and an instant weight-update API.
- A telemetry pipeline producing short and long rolling windows of error rate, latency, and request volume, split by canary vs. baseline.
- A stateless controller that runs the step/check/decide loop and calls the router's weight API.
- A rollback path that can zero out canary weight immediately, independent of the step loop.
Guardrails (the part that actually matters):
| Guardrail | Purpose | Example setting |
|---|---|---|
| Minimum observation window | Prevents decisions on too little data | require >= 100 canary requests in the last 60s before advancing |
| Short window (fast signal) | Catches sharp regressions quickly | 60s rolling window |
| Long window (stability) | Filters transient noise, confirms sustained breach | 300s rolling window |
| Max error delta | Hard stop on quality regression | canary error rate > baseline + 0.5 pp |
| Max latency ratio | Hard stop on tail-latency regression | canary p99 > baseline p99 x 1.2 |
| Monotonic ramp | Never sneak weight back up after a hold | only increase, decrease only via rollback |
| Cooldown after rollback | Prevents a flapping loop from re-triggering immediately | 30 minute freeze + incident ticket |
Smoothing the weight changes. Compute a fixed step size from the target ramp, then apply the step in small sub-increments at the router level (e.g. over a few seconds) so a single jump doesn't itself cause a connection-pool or cache-warming spike.
Worked example
Target: shift 10% of production traffic to canary over a 10 minute (600s) window, checking every 30s.
steps=30s600s=20 step size=2010%=0.5% per stepSo the loop runs 20 times, adding 0.5 percentage points of canary weight each pass, provided the guardrails above all pass. If step 12 (weight = 6%) shows canary p99 = 340ms against a baseline p99 of 260ms:
260340≈1.31>1.2That exceeds the 1.2x latency-ratio guardrail, so the controller does not advance to step 13. It instead sets weight back to 0% immediately (not a gradual decrease), opens an incident, and freezes further attempts for the cooldown period.
# runs every 30s
current = get_current_weight()
target = min(10, current + 0.5)
m = fetch_metrics(short=60, long=300)
if m.canary_requests_short < 100:
return # not enough signal yet, hold at current weight
if m.error_rate_canary > m.error_rate_baseline + 0.005 or \
m.p99_canary_long > m.p99_baseline_long * 1.2:
set_weight(0) # immediate rollback, not gradual
open_incident(m)
start_cooldown(minutes=30)
else:
set_weight(target) # router ramps this smoothly over a few seconds
Extensions: downstream capacity and CI/CD integration
Downstream-capacity-constrained ramp. The guardrails above are all canary-side (error rate, latency). A canary can look perfectly healthy on its own metrics while the traffic it forwards saturates a downstream dependency, a connection pool, a queue, or a rate-limited third-party API, before that dependency's own health signal even reflects the problem. Add a downstream capacity ceiling as its own guardrail: before each step, check the downstream dependency's current utilization (connection-pool saturation, queue depth, or a published capacity headroom number) and cap the canary weight so its incremental load stays under that ceiling, independent of what the canary's own error and latency metrics say. In practice the step size becomes the smaller of the scheduled step and the downstream headroom divided by requests-per-weight-point, and downstream saturation becomes its own rollback trigger alongside the error and latency guardrails, since the canary's own SLOs will not catch a downstream problem until it is already failing.
CI/CD pipeline integration. The loop is usually not a standalone daemon, it is a stage in the deploy pipeline. A build that passes its test stage triggers the deploy, which starts the canary at 0% weight and hands control to this loop; the pipeline blocks promotion to the next stage (wider rollout, or the next environment) until the loop reports a verdict, not a fixed timer. On success the loop reports pass and the pipeline proceeds to full rollout. On a guardrail breach, the loop's immediate zero-weight rollback from the worked example above is reported back as a failed pipeline stage: the deploy is marked failed, the previous version stays serving at 100%, and the pipeline stops rather than continuing to the next stage. That makes the human's only manual step reviewing the failure, not remembering to check on the canary.
Trade-offs and pitfalls
- Short windows are noisy, long windows are slow. A short-only window trips on transient blips; a long-only window lets real regressions run for minutes before anyone notices. Using both, and requiring the long window to confirm, is the standard resolution.
- Rollback must never itself be a slow ramp. The instinct to "gradually" bring weight back down defeats the purpose; on breach, cut to zero immediately and investigate after.
- Minimum-request guardrails matter more for low-traffic services. A service doing 50 RPS total needs a longer observation window or a lower canary weight cap just to get statistically meaningful samples per step.
- Business metrics can lag system metrics. A checkout canary can look perfectly healthy on latency and error rate while conversion silently drops; if the metric that matters is business-level, it needs its own window and threshold, not just infra SLOs.
- This is inherently single-service. Running the same loop concurrently across many dependent services without coordination risks compounding partial failures across a call graph; large orgs typically centralize this into a shared canary-analysis service rather than one loop per team.
You need to tune health checks across thousands of instances so that rolling deployments don't cause flapping (healthy instances briefly failing checks and being pulled from rotation) while real failures are still caught quickly. Walk through your recommended probe interval, timeout, consecutive-failure threshold, and jitter strategy, and how you would validate and roll out these settings safely.
Sample Answer
Direct Answer
Pick the probe interval and consecutive-failure threshold together, since detection latency is just their product, then push the false-positive rate down by requiring multiple consecutive failures rather than reacting to one, and add jitter so thousands of instances don't probe (or fail) in lockstep during a rolling deploy. Keep per-instance removal decisions fast and deterministic, but gate any cluster-level alerting or autoscaling reaction on an aggregated, percentage-based signal, so the normal wave of transient blips a rolling deploy produces never triggers a page or a scaling event.
Recommended Parameters
| Parameter | Recommended range | Rationale |
|---|---|---|
| Probe interval | 8-15s | Short enough to catch real failures in under a minute; long enough that jittered probes don't cluster |
| Timeout | 2-4s (well under the interval) | Fails a hung/blocked instance fast without waiting out the full interval |
| Consecutive-failure threshold (unhealthy) | 2-4 | Requires sustained failure before removal; see the worked example for why this matters more than the interval itself |
| Consecutive-success threshold (healthy again) | 1-2 | Prevents an instance from flapping in and out of rotation on borderline recovery |
Jitter and Staggering
Two different problems need two different fixes:
- Probe timing jitter: randomize each instance's probe interval by roughly ±10-25% (
interval = base * (1 + rand(-0.15, 0.15))) so thousands of probes don't all fire in the same tick. - Deploy-wave staggering: stagger each instance's initial probe phase deterministically, e.g.
start_offset = hash(instance_id) mod base_interval, so a rolling deploy that restarts instances in batches doesn't produce synchronized probe failures across the whole batch at once.
Startup vs Steady-State Probing
A common misconfiguration is using one probe, tuned for steady state, for both traffic-removal and process-restart decisions. A slow-starting instance under a strict steady-state threshold gets killed and restarted before it ever becomes ready, which can cascade into a thundering-restart loop under load (each restart is slower than the last because the fleet is smaller and more loaded). The fix is to separate concerns the way most orchestrators do: a startup-phase probe with a longer grace period and higher failure tolerance that only governs the initial ready-or-not decision, a readiness probe (steady state) that governs load-balancer inclusion, and a liveness probe, decoupled from readiness, that governs whether the process gets restarted at all. Traffic removal should never be driven by the same signal that triggers a restart.
Per-Instance vs Aggregated Logic
- Per-instance (LB inclusion): strict, deterministic, uses the consecutive-failure rule above. One bad instance should not need cluster-wide context to get pulled.
- Cluster-level (alerting, autoscaling): rate-based, not count-based. Example policy: only alert if more than 5% of instances are unhealthy for more than 2 minutes, or if the absolute unhealthy count exceeds a fixed floor and is still rising. This absorbs the expected noise of a rolling deploy (where some fraction of the fleet is always transiently cycling) without needing per-deploy tuning.
Worked Example
Detection latency. With interval =10s and a consecutive-failure threshold k=3:
detection latency=interval×k=10s×3=30sFalse-positive suppression from raising the threshold. Suppose a transient, independent failure (a coincidental GC pause or a single dropped packet) affects any given probe with probability p=0.02 (2%). Requiring k consecutive failures before declaring unhealthy, assuming failures are independent across probes, makes the false-trip probability:
P(false trip)=pkFor k=1 (react on the first failure): P=0.02=2% per check.
For k=3: P=0.023=8×10−6=0.0008%.
That is a reduction factor of:
p3p=p21=0.0221=2500×fewer false trips, at the cost of 3x the detection latency (10s to 30s). This independence assumption is the important caveat: a sustained GC pause or a real network partition produces correlated consecutive failures, where raising k buys much less protection, because the failures aren't independent draws anymore. The threshold you pick is really a statement about how much you trust that your transient-failure sources are actually transient (single-probe-length) rather than sustained.
Validation and Safe Rollout
- Bench test in staging: simulate probe latency and intermittent failures against a synthetic fleet to measure the observed false-positive rate and detection latency before touching production defaults.
- Canary the setting itself: apply the new thresholds to a small slice of instances (roughly 1-2%) first, watching LB churn rate, 4xx/5xx trends, and alert volume against the rest of the fleet as a baseline.
- Ramp: expand to larger slices (10%, 25%, 50%) only if churn and alert-noise metrics stay flat relative to baseline at each stage.
- Perturbation testing: inject latency, packet loss, or CPU pressure in staging to confirm the new thresholds don't cause mass removal under conditions you consider acceptable to tolerate.
- Rollback path: keep the previous thresholds one config change away, and define the specific churn or alert-noise threshold that triggers an automatic revert rather than waiting for a human to notice.
Trade-offs and Pitfalls
- Lower interval and lower threshold means faster detection but more false positives; there is no setting that improves both simultaneously, only a documented trade-off appropriate to the service's error budget.
- Conflating liveness and readiness (see above) is one of the most common causes of deploy-induced outages: it turns a slow-starting instance into a restart loop instead of just a delayed rotation entry.
- Aggregated, percentage-based cluster alerts are necessary specifically because per-instance thresholds, tuned tightly enough to catch real failures fast, will always produce some background noise during any large rolling deploy; treating that noise as page-worthy trains the on-call team to ignore alerts.
- The independence assumption behind the pk false-trip math is an approximation; validate it against real observed failure correlation in staging (step 1) rather than trusting the formula on defaults alone.
Explain the difference between readiness and liveness health checks (and startup checks, where supported). Design the checks for a database-backed web service (for example, covering DB connectivity and queue backlog), and explain how a misconfigured probe can cause cascading restarts or route traffic into a blackhole.
Sample Answer
Direct answer
Liveness answers "should this process be restarted"; readiness answers "should this instance receive traffic right now"; and a startup check, where the platform supports one, answers "has this instance finished initializing," so liveness does not kill a process that is simply still starting up. Keep liveness checks local and cheap, never calling the database or another network dependency, and keep readiness checks inclusive of whatever actually determines if this instance can do its job right now.
Structured elaboration
Designing checks for a database-backed web service:
- Liveness: process and event-loop responsiveness only (for example, can the process handle a trivial in-memory request). No database calls.
- Readiness: a bounded database ping, such as a single
SELECT 1with a tight timeout, so a slow or unreachable database marks the instance not-ready rather than hanging the check itself. - Readiness, if the service consumes a queue: a consumer-lag or backlog check, since a service that is technically up but has fallen far behind its queue is not really ready to take more work either.
- Startup check: if the platform supports one, use it to give the service extra time for one-time initialization (schema checks, cache warm-up, waiting on a migration) before liveness checks start at all, instead of stretching liveness's own thresholds to cover startup.
Pod lifecycle relative to these probes (SIGTERM below is the standard "please shut down" signal the orchestrator sends before forcibly killing a container):
stateDiagram-v2
[*] --> Starting
Starting --> NotReady: startup probe pending
NotReady --> Ready: readiness probe passes
Ready --> NotReady: readiness probe fails
Ready --> Restarting: liveness probe fails
NotReady --> Restarting: liveness probe fails
Restarting --> Starting: container restarted
Ready --> Draining: SIGTERM received
Draining --> [*]: in-flight requests finish
Worked example
A concrete configuration: readiness pings the database with a 100ms timeout and checks that consumer lag stays under a configured threshold; liveness only checks the process. Using the same worst-case reasoning as check-frequency tuning generally, with a check period of 10 seconds:
liveness worst-case restart trigger=period×failureThreshold=10×5=50 (seconds) readiness worst-case removal=period×failureThreshold=10×3=30 (seconds)Readiness is deliberately more sensitive (lower threshold) than liveness here, because pulling an instance out of rotation is cheap and reversible, while restarting a process is disruptive and should require stronger evidence.
How a misconfigured probe causes each failure mode, for this exact service:
- Cascading restarts: if liveness includes the database ping (against the design above) and the database has a brief outage, every replica fails liveness at roughly the same time. The orchestrator restarts the whole fleet concurrently, and the new pods immediately hit the same still-recovering database, failing liveness again, a self-inflicted restart storm that outlives the original database blip.
- Traffic blackhole: if readiness is too strict, or its database timeout is tighter than the database's real worst-case latency under load, a load spike can cause every pod to fail readiness at the same moment. The service is left with zero endpoints and every request gets a connection error, even though each pod was actually capable of serving traffic, just slower than the timeout allowed.
Trade-offs & pitfalls
- Never let liveness depend on an external service; a dependency outage should be a readiness problem (route traffic elsewhere) not a liveness problem (restart everything).
- Keep readiness dependency checks shallow and bounded (a ping, not a full transaction) so the check itself cannot become the slow thing that trips the threshold.
- Use disruption budgets or max-unavailable settings (Kubernetes controls that cap how many replicas can be down at once during a rollout or node drain) so an orchestrator responding to a wave of liveness failures cannot restart the entire fleet at once, which is what turns a brief dependency blip into an extended outage.
Compare connection management for HTTP/2 and gRPC traffic behind a Layer 7 load balancer: long-lived multiplexed connections versus ephemeral short-lived ones. How does connection pooling and multiplexing change throughput and resource usage compared to HTTP/1.1 keep-alive, what per-connection limits would you tune, and how should the load balancer measure load and apply backpressure to avoid head-of-line effects?
Sample Answer
Direct answer
HTTP/1.1 keep-alive needs roughly one TCP connection per concurrent in-flight request, so an L7 balancer scales by managing connection count. HTTP/2 and gRPC multiplex many concurrent streams over a small number of long-lived connections, so the balancer has to scale by managing stream concurrency within a connection instead, and load signals that were adequate for HTTP/1.1 (active connection count) become nearly useless. The practical consequence: fewer sockets and less TLS/TCP handshake overhead, but a new failure mode where one connection getting stalled or overloaded can degrade every stream multiplexed on it (head-of-line effects), which the balancer has to actively guard against.
Connection model comparison
| Dimension | HTTP/1.1 keep-alive | HTTP/2 | gRPC (HTTP/2 framing) |
|---|---|---|---|
| Connection lifetime | Reused per client, but 1 request in flight per connection (no multiplexing) | Long-lived, multiplexes many streams | Long-lived, same as HTTP/2 plus persistent bidirectional streaming RPCs |
| Concurrency unit the LB should track | Connection count | Streams per connection | Streams per connection, plus per-RPC deadlines |
| Typical tuning knob | max connections per backend, idle timeout, pool size | max_concurrent_streams per connection, connection pool size (few per backend), flow-control window | Same as HTTP/2, plus keepalive ping interval for long-idle streaming RPCs |
| Resource cost per unit of throughput | Higher: 1 TCP handshake + TLS handshake per request burst, more open sockets | Lower: handshake cost amortized across many streams, fewer sockets | Lower, same amortization; adds framing/serialization overhead per message |
| Head-of-line risk | None at the LB (each request has its own connection) | Yes: a stalled connection (e.g. TCP loss) blocks every multiplexed stream on it until retransmit | Yes, same mechanism, worse impact if a stream is a long streaming RPC holding the connection open |
Per-connection limits to tune
- HTTP/1.1: max connections per backend (bounds concurrency directly), idle keep-alive timeout (frees sockets from clients that went quiet), and pool size on the LB's upstream side.
- HTTP/2 / gRPC:
max_concurrent_streamsper connection (how many in-flight requests one connection may carry, commonly capped well below the protocol's theoretical maximum to bound blast radius), a small upstream connection pool per backend (a handful, not one) so a single stalled connection doesn't take out all traffic to that backend, per-stream flow-control window size (the amount of unacknowledged data a stream may have in flight before the sender must pause), and a keepalive ping interval to detect a half-open connection before streams queue behind a dead peer.
Load measurement and backpressure
Connection count stops being a useful load signal once multiplexing is in play: a backend with 4 connections and 400 streams looks identical to one with 4 connections and 4 streams if you only count sockets. Instead:
- Measure in-flight streams per backend (or
max_concurrent_streams - current_streamsas available capacity) and use that as the weighting signal for load-aware routing. - Track per-stream and per-connection latency percentiles separately; a rising per-connection tail with stable per-stream counts points at connection-level contention (CPU, flow-control), not request volume.
- Apply admission control at the stream level: reject new streams past a concurrency threshold with a fast, typed backpressure signal (gRPC
RESOURCE_EXHAUSTED, HTTP 429) rather than accepting and queuing, which just moves the head-of-line problem later. - Detect stalled streams (no progress against their flow-control window) and reset them individually instead of tearing down the whole connection, so one bad stream doesn't punish every other stream sharing it.
Worked example
Suppose the LB maintains a pool of 4 HTTP/2 connections to a backend, each configured with max_concurrent_streams = 100:
That backend can serve 400 concurrent RPCs using 4 sockets. Reaching the same 400 concurrent in-flight requests under HTTP/1.1 keep-alive would require roughly 400 separate TCP+TLS-established connections (one per in-flight request), which is exactly the socket and handshake overhead multiplexing removes. The other side of that number: if one of those 4 connections stalls, up to 100 of the 400 in-flight requests (25%) can be head-of-line blocked simultaneously, which is why the pool size (not just 1 connection) and per-stream stall detection both matter.
Trade-offs and pitfalls
- Fewer, fatter connections are more efficient but concentrate risk: a single TCP-level packet loss stalls every stream on that connection at the transport layer, even though the streams are logically independent at the application layer. This TCP-level head-of-line blocking is a known limitation of running multiplexing over TCP (it's the reason HTTP/3/QUIC exists), but that's depth beyond what most interviews expect; the interview-relevant point is just that a small connection pool, not a single connection, bounds the blast radius.
- A common mistake is reusing HTTP/1.1-era LB health/load metrics (connection count, connections-per-second) unchanged for HTTP/2 backends; they will systematically under-detect overload because a backend can look "quiet" on connections while being saturated on streams.
- Setting
max_concurrent_streamstoo high trades efficiency for blast radius; setting it too low defeats the purpose of multiplexing and pushes you back toward connection-count scaling.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.