Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
Explain the difference between readiness and liveness health checks (and startup checks, where supported). Design the checks for a database-backed web service (for example, covering DB connectivity and queue backlog), and explain how a misconfigured probe can cause cascading restarts or route traffic into a blackhole.
Sample Answer
Direct answer
Liveness answers "should this process be restarted"; readiness answers "should this instance receive traffic right now"; and a startup check, where the platform supports one, answers "has this instance finished initializing," so liveness does not kill a process that is simply still starting up. Keep liveness checks local and cheap, never calling the database or another network dependency, and keep readiness checks inclusive of whatever actually determines if this instance can do its job right now.
Structured elaboration
Designing checks for a database-backed web service:
- Liveness: process and event-loop responsiveness only (for example, can the process handle a trivial in-memory request). No database calls.
- Readiness: a bounded database ping, such as a single
SELECT 1with a tight timeout, so a slow or unreachable database marks the instance not-ready rather than hanging the check itself. - Readiness, if the service consumes a queue: a consumer-lag or backlog check, since a service that is technically up but has fallen far behind its queue is not really ready to take more work either.
- Startup check: if the platform supports one, use it to give the service extra time for one-time initialization (schema checks, cache warm-up, waiting on a migration) before liveness checks start at all, instead of stretching liveness's own thresholds to cover startup.
Pod lifecycle relative to these probes (SIGTERM below is the standard "please shut down" signal the orchestrator sends before forcibly killing a container):
stateDiagram-v2
[*] --> Starting
Starting --> NotReady: startup probe pending
NotReady --> Ready: readiness probe passes
Ready --> NotReady: readiness probe fails
Ready --> Restarting: liveness probe fails
NotReady --> Restarting: liveness probe fails
Restarting --> Starting: container restarted
Ready --> Draining: SIGTERM received
Draining --> [*]: in-flight requests finish
Worked example
A concrete configuration: readiness pings the database with a 100ms timeout and checks that consumer lag stays under a configured threshold; liveness only checks the process. Using the same worst-case reasoning as check-frequency tuning generally, with a check period of 10 seconds:
liveness worst-case restart trigger=period×failureThreshold=10×5=50 (seconds) readiness worst-case removal=period×failureThreshold=10×3=30 (seconds)Readiness is deliberately more sensitive (lower threshold) than liveness here, because pulling an instance out of rotation is cheap and reversible, while restarting a process is disruptive and should require stronger evidence.
How a misconfigured probe causes each failure mode, for this exact service:
- Cascading restarts: if liveness includes the database ping (against the design above) and the database has a brief outage, every replica fails liveness at roughly the same time. The orchestrator restarts the whole fleet concurrently, and the new pods immediately hit the same still-recovering database, failing liveness again, a self-inflicted restart storm that outlives the original database blip.
- Traffic blackhole: if readiness is too strict, or its database timeout is tighter than the database's real worst-case latency under load, a load spike can cause every pod to fail readiness at the same moment. The service is left with zero endpoints and every request gets a connection error, even though each pod was actually capable of serving traffic, just slower than the timeout allowed.
Trade-offs & pitfalls
- Never let liveness depend on an external service; a dependency outage should be a readiness problem (route traffic elsewhere) not a liveness problem (restart everything).
- Keep readiness dependency checks shallow and bounded (a ping, not a full transaction) so the check itself cannot become the slow thing that trips the threshold.
- Use disruption budgets or max-unavailable settings (Kubernetes controls that cap how many replicas can be down at once during a rollout or node drain) so an orchestrator responding to a wave of liveness failures cannot restart the entire fleet at once, which is what turns a brief dependency blip into an extended outage.
Describe how you would coordinate load balancer connection draining with autoscaling group scale-in so instances finish in-flight work before termination. Cover cooldowns, warm-up periods, and how the load balancer's health-check readiness state should gate when a new instance starts receiving traffic.
Sample Answer
Direct answer
The load balancer must gate two moments in an instance's life: it should not receive traffic until its readiness check passes (not just a process-alive check), and on scale-in it must stop receiving new connections while letting in-flight requests finish, deregistration with connection draining, before the instance is terminated. Both are the LB's job, independent of whatever policy decided to scale in the first place.
Structured elaboration
Scale-out: readiness gates traffic, not liveness. A liveness check ("is the process up") and a readiness check ("is the app actually able to serve correctly") are different questions. The LB should register the instance in rotation, or the orchestrator should mark it Ready, only after readiness passes: connection pools opened, local caches warmed enough to avoid a cold-cache latency cliff, and any required config pulled. A liveness-only gate puts unwarmed instances into rotation and produces a latency or error spike on every scale-out event.
Scale-in: draining before termination, in order:
- Instance is marked for termination (by the autoscaler or an operator).
- The LB immediately stops sending new connections to it (deregister from rotation) but does not kill existing connections.
- A deregistration delay (connection draining window) gives in-flight requests time to complete normally.
- Only after the drain window elapses, or all in-flight requests finish, whichever is defined as the policy, does the instance actually terminate.
Sizing the drain window. It must exceed the longest realistic in-flight request duration for the service, with margin, or draining will forcibly cut live requests, defeating the point.
Health-check readiness during warm-up. The initial health-check grace period (the time before a failed check counts against a brand-new instance) must exceed real startup time, so a slow-booting instance is not falsely marked unhealthy and cycled before it ever gets a chance to serve.
Cooldowns, and where autoscaling policy stops and LB mechanics start. Which metric and threshold trigger a scale event is an autoscaling-policy decision outside this question's scope, but the cooldown itself belongs here since the question names it directly: a cooldown is the minimum time an autoscaler must wait after one scaling action before it is allowed to trigger another, and it exists because the fleet's aggregate metrics take time to reflect the previous action, an autoscaler that reacts to a still-stale signal will overreact, stacking a second scale-out on top of one still warming up, or yanking capacity out from under a scale-in that hasn't finished draining. The cooldown has to be sized against the LB mechanics above, not chosen independently of them: it should exceed both the drain window (a scale-in's capacity isn't actually gone until draining finishes) and the health-check grace period (a scale-out's capacity isn't actually serving, and won't show up in the fleet's metrics, until warm-up and readiness finish), see the worked example below for a concrete number. Whatever triggers a scale-out or scale-in event, the LB's readiness gate and drain sequence are what make that event safe for in-flight traffic; the cooldown is what keeps the autoscaler from firing faster than those mechanics can complete.
Worked example
Say p99 request duration for this service, measured over the last observation period, is 8 seconds, and cold-start warm-up (open DB pool + populate local cache) takes 15 seconds.
Deregistration delay, with margin over the slowest realistic request:
drain window=8s×1.25≈10sHealth-check grace period, with margin over startup time:
grace period=15s+10s (buffer)=25s, rounded up to 30sSo: a new instance is excluded from the health check's failure count for its first 30 seconds while it warms up, and a terminating instance gets up to 10 seconds to finish in-flight requests after being pulled from rotation, before the process is killed.
Cooldown, sized to exceed both of those numbers so the autoscaler never acts on a stale signal:
cooldown=max(drain window,grace period)×1.5≈30s×1.5=45sSo the autoscaler waits at least 45 seconds after any scale-out or scale-in action before it will consider triggering another one, comfortably longer than the 30 second grace period a new instance needs before its readiness, and therefore its true added capacity, is even visible in the fleet's metrics.
Trade-offs and pitfalls
- Too-short drain windows silently drop real user requests mid-response; this is one of the most common causes of a small, hard-to-reproduce error-rate blip during routine scale-in.
- Too-long drain windows slow down scale-in, which matters when scale-in is reacting to a cost or capacity signal; there's a real tension between "never cut a request" and "shed capacity promptly."
- A liveness-only readiness gate causes a recurring latency spike on every scale-out, because unwarmed instances start serving immediately; teams often discover this only when scale-out frequency increases (e.g. spiky traffic) and the spikes become visible in aggregate latency graphs.
- Grace periods that are too short flap healthy-but-slow-booting instances, causing the autoscaler and LB to fight each other: instance boots, fails health check before finishing warm-up, gets cycled, repeats.
- Drain and readiness are load-balancer-owned; they are not a substitute for correct autoscaling trigger design. A perfectly tuned drain window cannot fix an autoscaler that scales in too aggressively in the first place, that is a separate, upstream problem.
You manage a pool of stateless web servers behind a single load balancer that sees bursty traffic and heterogeneous server capacity. Compare round robin, least connections, and weighted load balancing for this situation: how does each pick a server, and which would you choose?
Sample Answer
Direct answer
Round robin, least connections, and weighted load balancing differ in what information they use to pick a server. Round robin uses none, it just cycles through the pool. Least connections uses current load (active connection count) but assumes every server has equal capacity. Weighted load balancing (typically combined with one of the other two, giving weighted round robin or weighted least connections) uses declared server capacity. For a pool with bursty traffic and heterogeneous capacity, weighted least connections is the right default, because it is the only one of the three that reacts to both current load and known capacity differences.
How each one picks
| Algorithm | Selection rule | Reacts to load? | Reacts to capacity? |
|---|---|---|---|
| Round robin | Next server in a fixed rotation | No | No |
| Least connections | Server with fewest active connections | Yes | No (treats all servers as equal) |
| Weighted (least connections) | Server with the lowest active-connections-to-weight ratio | Yes | Yes |
Why plain round robin and plain least connections both fall short here
Round robin sends the same share of traffic to every server regardless of what it can handle, so a small server gets hit exactly as often as a large one and becomes the bottleneck under burst. Least connections is an improvement, since it avoids piling more work onto an already-busy server, but it still treats a 2-vCPU box and an 8-vCPU box as equally capable: an equal number of active connections is not an equal amount of work when the servers are different sizes.
Worked example: weighted least connections in action
Say the pool has three servers with declared capacity weights 4, 2, and 1 (roughly matching, say, an 8-vCPU, 4-vCPU, and 2-vCPU box). Weighted least connections picks, for each new request, whichever server minimizes:
scorei=wiciwhere ci is server i's current active connection count and wi is its weight. Ties are broken toward the higher-weight server. Tracing 7 back-to-back requests (a burst, so nothing completes mid-trace):
| Request | Scores (S1, S2, S3) | Picked | Active after |
|---|---|---|---|
| 1 | 0, 0, 0 (tie) | S1 | S1=1, S2=0, S3=0 |
| 2 | 0.25, 0, 0 (tie) | S2 | S1=1, S2=1, S3=0 |
| 3 | 0.25, 0.5, 0 | S3 | S1=1, S2=1, S3=1 |
| 4 | 0.25, 0.5, 1.0 | S1 | S1=2, S2=1, S3=1 |
| 5 | 0.5, 0.5 (tie), 1.0 | S1 | S1=3, S2=1, S3=1 |
| 6 | 0.75, 0.5, 1.0 | S2 | S1=3, S2=2, S3=1 |
| 7 | 0.75, 1.0, 1.0 | S1 | S1=4, S2=2, S3=1 |
After 7 requests the split is S1=4, S2=2, S3=1, exactly proportional to the declared weights of 4:2:1, even though every request arrived back-to-back with no completions in between. Plain round robin would have sent requests 1-7 as S1,S2,S3,S1,S2,S3,S1 (a 3:2:2 split), overloading S3 relative to its actual capacity.
Trade-offs and pitfalls
Weighted least connections needs two things plain least connections does not: accurate weights (usually derived from instance size or a load test, not guessed) and a mechanism to keep them current as capacity changes (autoscaling, degraded instances). It is also more stateful, the balancer has to track active connections per server accurately, which matters less for round robin. If request cost varies wildly and isn't reflected by connection count (a handful of expensive long-running requests versus many cheap ones), even weighted least connections can misjudge load, at which point request-cost-aware balancing or queuing becomes worth considering.
Design a high-throughput traffic mirroring pipeline that copies a sample of production requests to a staging cluster without adding latency to the production path. Cover asynchronous delivery, sampling strategy, masking of sensitive data, and how you would prevent a mirrored write from causing a real side effect in the staging system.
Sample Answer
Direct answer
Make the mirroring decision and enqueue non-blocking at the edge so production latency is unaffected, ship the sampled copy asynchronously through a durable stream to staging, and prevent side effects by treating the staging path as read-only by construction: either point it at a storage layer that discards or isolates writes, or convert write methods before they reach a real backend. Sensitive data gets masked before it leaves the edge, not after it lands in staging, since "after" means it already left the trust boundary.
Structured elaboration
- Decision point stays fast. At the edge or sidecar, evaluate sampling rules against cheap, already-available data (route, a few headers) and enqueue a copy into a local, bounded, in-memory buffer. The enqueue is fire-and-forget: on a full buffer, drop and increment a counter, never block the production request.
- Asynchronous transport. A background worker drains the local buffer into a durable, high-throughput stream (for example Kafka or Kinesis), batching and compressing for efficiency. This decouples production's request rate from staging's actual processing rate.
- Masking at the edge, not in staging. Redact or tokenize PII fields and strip auth secrets in the same process that decides to mirror, before the copy is serialized onto the stream. Waiting until staging to redact means an unredacted copy already crossed a trust boundary and sat in a durable log.
- Preventing write side effects. The staging ingress recognizes mirrored traffic (a shadow-mode header) and routes it through an adapter that either discards writes, redirects them to an isolated namespace or database, or serves reads from a replica while writes are stubbed. Add an idempotency key (a unique id attached to a request so the receiving system can recognize and discard a duplicate instead of applying it twice) to every mirrored request so retries in the pipeline itself cannot double-apply a write that did slip through.
- Fidelity validation. A comparator service consumes both the mirrored request and the staging response (correlated by an original trace ID carried in a header) and checks response-shape agreement (status code family, latency distribution, schema) against production, without needing a human to eyeball individual requests. This is what tells you the shadow environment is representative, not just quiet.
graph LR
Prod[Production request] --> Edge[Edge: sample + mask + async enqueue]
Edge --> Serve[Continue serving production, no wait]
Edge -.-> Buffer[Bounded local buffer]
Buffer --> Stream[Durable stream: Kafka/Kinesis]
Stream --> Consumer[Mirror consumer]
Consumer --> Ingress[Staging ingress: shadow-mode adapter]
Ingress --> Staging[Staging services, writes isolated]
Consumer --> Comparator[Comparator: prod vs staging fidelity]
Worked example
Production ingress runs at 50,000 rps. A 2% sampling target sends 50000×0.02=1000 rps into the mirroring pipeline. Staging is provisioned with 20% headroom over that mirrored rate: 1000×1.2=1200 rps of capacity.
If a single stream partition sustains 5,000 rps of small messages, one partition is technically enough for 1,000 rps (⌈1000/5000⌉=1), but the pipeline is provisioned with 4 partitions so 4 consumers can process in parallel; each partition then carries an average of 1000/4=250 rps, well under its 5,000 rps ceiling, leaving headroom for an uneven key distribution across partitions without any single consumer falling behind.
Trade-offs & pitfalls
- Fail-open is the only safe default: if the stream backs up or the buffer fills, drop mirrored traffic and keep counting drops, never apply backpressure (forcing an upstream sender to slow down or block because a downstream consumer can't keep up) to the production request path itself.
- Converting writes at the staging ingress (stubbing them out) keeps staging's data pristine but can make staging diverge functionally from production over time, since write-triggered side effects never happen there; using a read-only replica plus stubbed writes downstream keeps closer functional parity at the cost of extra infrastructure.
- Redacting at staging instead of at the edge is the single most common mistake in shadow-traffic designs: an unmasked copy sitting in a durable stream, even briefly, is a real data-exposure surface, not a theoretical one.
- Sampling and mirroring add real infrastructure cost (stream, consumers, staging capacity) proportional to sampled volume; treat the sampling rate as a cost dial, not just a load-control dial.
Design a monitoring and alerting setup for a load balancer tier so you catch unhealthy backends before users see errors. Which metrics would you include and why, what alert thresholds would you propose, and how would you differentiate severity to avoid alert fatigue?
Sample Answer
Direct answer
Catching unhealthy backends before users notice means monitoring the LB tier along four angles: traffic, latency, and errors as seen from the LB itself (the RED metrics), backend health signals independent of user-facing symptoms, saturation and capacity of the LB tier, and TLS/connection-layer health, then tying alert severity to how much of the user-facing error budget (the amount of failure your reliability target still allows before you've broken your promise to users) is actually being consumed rather than to a single static threshold, so a brief blip and a sustained outage don't page the same way.
The metric set
| Category | Metric | Why it matters |
|---|---|---|
| Traffic | Requests per second, overall and per backend pool | Baseline for interpreting every other metric; a rate change can explain an error-rate change |
| Latency | p50 / p90 / p95 / p99, split into LB-processing time and end-to-end time | Percentiles catch tail degradation that an average hides; splitting LB time from backend time tells you where the slowdown is |
| Errors | 4xx vs 5xx rate, and whether 5xx originated at the LB (no healthy backend) or was passed through from a backend | A 5xx-at-the-LB spike usually means capacity loss, not an application bug |
| Backend health | Per-backend health-check pass/fail rate and consecutive-failure counts | The direct signal for whether a specific node is about to be pulled |
| Connections | Active connections and connection churn (opens/closes per second) | A churn spike often precedes a latency or error spike |
| Capacity | LB-tier CPU, memory, socket/file-descriptor usage | The LB tier itself can become the bottleneck, not just the backends |
| TLS | Handshake failure rate, certificate expiry countdown | Silent failure mode: expired certs or handshake issues look like "backend is fine but nobody can connect" |
| Traffic shaping | Rate-limit rejections, circuit-breaker open/close events | Leading indicators: a circuit breaker tripping on a backend is an early warning before that backend fails health checks outright |
Alert severity: tie it to error-budget burn, not just a raw threshold
A raw threshold ("page if 5xx rate > 1%") treats a 10-minute blip the same as a slow multi-hour decline, a common cause of alert fatigue in either direction: it either pages on noise or is set so loose it misses real problems. The more robust approach expresses the threshold as a multiple of your error-budget burn rate: if you're consuming your monthly error budget faster than the rate at which "consume it evenly over 30 days" would, page proportionally to how much faster.
Worked example
Suppose the LB tier's SLO (service-level objective, the reliability target you've committed to, here 99.9% of requests succeeding) is 99.9% success rate over a 30-day window. The allowed error budget is 0.1% of requests, equivalently, 0.1% of the 30-day window can be fully down and still meet the SLO:
- 30 days in minutes: 30×24×60=43,200 minutes.
- Budget: 43,200×0.001=43.2 minutes of full-downtime-equivalent for the entire month.
Now suppose the current 5xx rate over the last hour is holding at 2%, against an objective of 0.1%. The burn rate is:
burn rate=0.1%2%=20×At 20 times the sustainable rate, the entire month's budget would be exhausted in:
2030 days=1.5 daysA burn rate that would exhaust the monthly budget in 1.5 days is unambiguously page-worthy; a burn rate of, say, 1.2 times sustainable is not, it just needs tracking. That is the difference between a severity-1 page and a low-priority ticket, expressed in the same unit (budget consumption), not two arbitrarily different percentage cutoffs.
Avoiding alert fatigue
- Tier severity by projected budget-exhaustion time (minutes-to-exhaust pages immediately, weeks-to-exhaust becomes a ticket), rather than by a single static percentage.
- Deduplicate and correlate: a health-check-flapping alert and a 5xx-rate alert firing from the same underlying cause should collapse into one incident, not page twice.
- Route by audience: on-call needs the actionable, current-state view (which backend, since when); a weekly review needs the trend (is health-check-induced churn increasing over time).
- Require every page-level alert to link directly to a runbook step; an alert with no clear next action trains people to ignore it.
Trade-offs and pitfalls
Over-indexing on system-level metrics (CPU, connections) without a user-facing correlate risks paging on things users never notice; conversely, only alerting on aggregate error rate can hide a problem affecting one backend pool or one tenant while the aggregate stays healthy. Circuit-breaker and rate-limit rejection counts are easy to skip because they're not "errors" in the traditional sense, but they're often the earliest signal, by the time health checks start failing outright, the circuit breaker has usually already tripped.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.