Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
Design a Layer 7 load balancer that provides session affinity using consistent hashing on a session cookie. It must support health checks and rebalance sessions gracefully when nodes are added or removed. Discuss hash ring maintenance, virtual nodes, and how you would drain and migrate sessions without dropping in-flight traffic.
Sample Answer
Direct Answer
Hash the session cookie onto a consistent-hash ring of virtual nodes so a session's requests keep landing on the same backend under normal conditions, publish the ring from a small versioned control plane so every proxy agrees on current ownership, pull unhealthy nodes out via active health checks, and when a node is intentionally removed, keep it serving its existing sessions for a bounded drain window while new sessions route elsewhere, only decommissioning it once its active-session count reaches zero or the window expires. The two things that make this safe are versioning the ring (so proxies never disagree mid-transition) and treating removal as a drain, not a hard cut.
Architecture
flowchart LR
Client -->|cookie + ring_version| Proxy[Edge Proxy]
Proxy -->|lookup| RingSvc[Ring Metadata Service]
HealthChecker[Health Checker] -->|mark unhealthy| RingSvc
Proxy --> Backend1[Backend 1]
Proxy --> Backend2[Backend 2]
Proxy -.draining.-> Backend3[Backend 3]
DrainOrch[Drain Orchestrator] -->|coordinate| RingSvc
Backend3 -->|flush session state| Store[(Shared Session Store)]
Backend1 --> Store
Backend2 --> Store
The proxies are stateless: all ring state lives in the ring metadata service and is cached locally with a version number. Health checks and the drain orchestrator only ever mutate the ring through that service, never by having a proxy make a unilateral decision.
Hash Ring and Virtual Nodes
- Use a large hash space (64-bit) with each physical backend assigned many virtual nodes, commonly in the range of 100 to 500, so that any single node's removal spreads its load across many different successors instead of dumping it on one (see the worked example below for why this matters).
- Store the ring as a sorted array in the metadata service; proxies cache a copy and watch for version bumps rather than polling on every request.
- Every ring mutation (join, leave, health-driven removal) increments the ring version atomically, so a proxy can always tell whether its cached copy is current.
Cookie and Version Handling
The session cookie carries three fields: the session id (the hash key), the ring version the session was originally assigned under, and an HMAC (a keyed cryptographic checksum: proof the cookie's contents haven't been altered since a server signed them) to prevent tampering. On a new session, the proxy issues a fresh cookie against the current ring version. On a returning session, the proxy honors the session's recorded ring version for a bounded overlap window even after the ring has moved on, which is what lets an in-flight session keep reaching its original backend during a drain instead of being silently reassigned mid-conversation.
Health Checks
Active HTTP health checks run against each backend on a fixed interval. A backend that fails its threshold is marked unhealthy in the ring metadata service, which bumps the ring version; proxies pick up the change and stop routing new sessions there. Existing sessions already pinned to that backend via cookie should still be judged against the same health signal: if the backend is actually down, in-flight requests will fail regardless of stickiness, so unhealthy removal is immediate, not drained (draining is reserved for planned, intentional removal, covered next).
Graceful Drain and Session Migration
For a planned node removal (scale-down, deploy, decommission):
- Mark the node "draining" in the ring metadata service.
- Remove the node's virtual-node entries from the ring and bump the ring version.
- Proxies pick up the new version for new sessions immediately; sessions whose cookie still carries the old ring version continue routing to the draining node for a bounded overlap window.
- During the overlap window, the draining node's session state is either already externalized (shared store, no action needed) or actively migrated: the drain orchestrator streams in-memory session state to the new owning node or to the shared store.
- Once the draining node's active-session count reaches zero, or the overlap window expires (whichever comes first), the node is decommissioned. Sessions that hadn't finished by then either see a session reset (acceptable for many applications) or, if externalized state was used, resume transparently on the new node.
For node addition, the reverse: add the virtual nodes, bump the version, and let new sessions start flowing to the new node immediately. Existing sessions are unaffected because their arcs on the ring did not move.
Session State Strategy
Prefer an external session store (a clustered cache such as Redis) over in-memory backend state whenever possible: it decouples session survival from any single backend's lifecycle entirely, removing the need for the migration step above. If in-memory sessions are unavoidable (e.g., a stateful protocol upgrade like a long-lived WebSocket), the migration path in step 4 is mandatory, not optional.
Worked Example: How Much Load a Removal Redistributes
Take 20 physical backends, each with V=150 virtual nodes, so T=3000 ring tokens total. Removing one physical backend removes its 150 tokens; each of those tokens' clockwise successors is, in a well-shuffled ring, effectively a random draw among the other 19 physical nodes. So the removed node's 201=5% share of keys spreads across roughly 19 recipients rather than one. Each surviving node's expected new share:
201+20×191=38019+1=191≈5.263%which is exactly the uniform share you'd expect from evenly splitting the whole keyspace across 19 nodes, the same identity that shows up whenever virtual-node count is high enough to approximate rendezvous hashing's (an alternative hashing scheme that scores every node directly against each key and always redistributes evenly on removal, without needing virtual nodes) exact uniform redistribution. Compare that to a single-token ring (no virtual nodes) removing one of 20 nodes: the removed node's entire 5% lands on one successor, whose share jumps from 5% to 10%, a 2x hotspot instead of a 0.26-point bump.
Trade-offs and Pitfalls
- The ring metadata service is a coordination dependency; if it uses a consensus protocol (Raft or similar) for consistency, that adds latency to ring mutations (acceptable, since mutations are rare) but also means the service itself needs its own HA story.
- Overlap window length is a direct trade-off between session continuity and routing complexity: a longer window keeps more in-flight sessions alive during churn but means proxies must correctly honor two ring versions simultaneously for longer.
- Signed ring-version cookies need HMAC key rotation handled carefully: rotating the signing key while old cookies are still in flight requires accepting both old and new keys for a transition period, or sessions get silently invalidated.
- If a CDN or edge cache sits in front of this layer, cache keys for session-specific content must not be shared across ring reassignment; either scope those responses as non-cacheable or ensure the cache key includes a stable shard identifier that survives rebalancing.
- A single very active session ("hot session") can skew load on whichever node currently owns it; virtual nodes fix cluster-wide statistical balance, not a single oversized key, so hot-key handling needs a separate mitigation (e.g., splitting that session's read traffic) if it becomes a real problem.
Design a monitoring and alerting setup for a load balancer tier so you catch unhealthy backends before users see errors. Which metrics would you include and why, what alert thresholds would you propose, and how would you differentiate severity to avoid alert fatigue?
Sample Answer
Direct answer
Catching unhealthy backends before users notice means monitoring the LB tier along four angles: traffic, latency, and errors as seen from the LB itself (the RED metrics), backend health signals independent of user-facing symptoms, saturation and capacity of the LB tier, and TLS/connection-layer health, then tying alert severity to how much of the user-facing error budget (the amount of failure your reliability target still allows before you've broken your promise to users) is actually being consumed rather than to a single static threshold, so a brief blip and a sustained outage don't page the same way.
The metric set
| Category | Metric | Why it matters |
|---|---|---|
| Traffic | Requests per second, overall and per backend pool | Baseline for interpreting every other metric; a rate change can explain an error-rate change |
| Latency | p50 / p90 / p95 / p99, split into LB-processing time and end-to-end time | Percentiles catch tail degradation that an average hides; splitting LB time from backend time tells you where the slowdown is |
| Errors | 4xx vs 5xx rate, and whether 5xx originated at the LB (no healthy backend) or was passed through from a backend | A 5xx-at-the-LB spike usually means capacity loss, not an application bug |
| Backend health | Per-backend health-check pass/fail rate and consecutive-failure counts | The direct signal for whether a specific node is about to be pulled |
| Connections | Active connections and connection churn (opens/closes per second) | A churn spike often precedes a latency or error spike |
| Capacity | LB-tier CPU, memory, socket/file-descriptor usage | The LB tier itself can become the bottleneck, not just the backends |
| TLS | Handshake failure rate, certificate expiry countdown | Silent failure mode: expired certs or handshake issues look like "backend is fine but nobody can connect" |
| Traffic shaping | Rate-limit rejections, circuit-breaker open/close events | Leading indicators: a circuit breaker tripping on a backend is an early warning before that backend fails health checks outright |
Alert severity: tie it to error-budget burn, not just a raw threshold
A raw threshold ("page if 5xx rate > 1%") treats a 10-minute blip the same as a slow multi-hour decline, a common cause of alert fatigue in either direction: it either pages on noise or is set so loose it misses real problems. The more robust approach expresses the threshold as a multiple of your error-budget burn rate: if you're consuming your monthly error budget faster than the rate at which "consume it evenly over 30 days" would, page proportionally to how much faster.
Worked example
Suppose the LB tier's SLO (service-level objective, the reliability target you've committed to, here 99.9% of requests succeeding) is 99.9% success rate over a 30-day window. The allowed error budget is 0.1% of requests, equivalently, 0.1% of the 30-day window can be fully down and still meet the SLO:
- 30 days in minutes: 30×24×60=43,200 minutes.
- Budget: 43,200×0.001=43.2 minutes of full-downtime-equivalent for the entire month.
Now suppose the current 5xx rate over the last hour is holding at 2%, against an objective of 0.1%. The burn rate is:
burn rate=0.1%2%=20×At 20 times the sustainable rate, the entire month's budget would be exhausted in:
2030 days=1.5 daysA burn rate that would exhaust the monthly budget in 1.5 days is unambiguously page-worthy; a burn rate of, say, 1.2 times sustainable is not, it just needs tracking. That is the difference between a severity-1 page and a low-priority ticket, expressed in the same unit (budget consumption), not two arbitrarily different percentage cutoffs.
Avoiding alert fatigue
- Tier severity by projected budget-exhaustion time (minutes-to-exhaust pages immediately, weeks-to-exhaust becomes a ticket), rather than by a single static percentage.
- Deduplicate and correlate: a health-check-flapping alert and a 5xx-rate alert firing from the same underlying cause should collapse into one incident, not page twice.
- Route by audience: on-call needs the actionable, current-state view (which backend, since when); a weekly review needs the trend (is health-check-induced churn increasing over time).
- Require every page-level alert to link directly to a runbook step; an alert with no clear next action trains people to ignore it.
Trade-offs and pitfalls
Over-indexing on system-level metrics (CPU, connections) without a user-facing correlate risks paging on things users never notice; conversely, only alerting on aggregate error rate can hide a problem affecting one backend pool or one tenant while the aggregate stays healthy. Circuit-breaker and rate-limit rejection counts are easy to skip because they're not "errors" in the traditional sense, but they're often the earliest signal, by the time health checks start failing outright, the circuit breaker has usually already tripped.
What is traffic mirroring (shadowing) at the load balancer level? Explain a safe approach to mirror production traffic to a staging cluster for testing, how you would sample it to limit load, and the risks, such as side effects in downstream systems or data contamination.
Sample Answer
Direct answer
Traffic mirroring (also called shadowing) copies live production requests and sends a duplicate to a second target, usually a staging or canary cluster, without ever returning that target's response to the real client. It lets you exercise new code with real traffic shapes and volumes before it serves anyone, with zero risk to the production response path.
Structured elaboration
Mechanism. The load balancer or a sidecar proxy (a small proxy process deployed alongside each service instance, rather than centralized in the LB tier: Envoy's request_mirror_policy, NGINX mirror, or an application-level fan-out) forwards a copy of the request asynchronously, fire-and-forget. The production request path never waits on the mirror's response, so mirror latency or errors cannot affect real users.
Safe rollout approach:
- Isolate the mirror target. It must not share a production database or write path. Point it at a replica or a sandboxed data store so any writes it performs cannot corrupt live data.
- Sanitize before forwarding. Strip or mask PII, auth tokens, and session cookies, or route through a sanitizing proxy layer, since the mirror environment is held to a lower trust bar than production.
- Suppress side effects. Anything that is not idempotent, payments, emails, third-party webhooks, external API calls, must be stubbed or short-circuited in the mirror environment. This is the single most important safety control.
- Tag mirrored traffic. Add a header or trace attribute (e.g.
x-mirrored: true) so logs, metrics, and alerting pipelines can separate mirror noise from real signal. - Sample instead of mirroring 100%. Use deterministic hashing on a stable key (user id or request id) so the same entities are consistently included or excluded, which keeps A/B-style comparisons stable across requests from the same user.
Worked example
Assume production runs at 50,000 requests per second (RPS) and you want to validate a new service version without over-provisioning staging.
Sampling at 2% (hash of user id, keep if hash(user_id) mod 100 < 2):
To avoid the mirror cluster itself becoming a bottleneck, provision it with headroom, say 1.5x the expected mirrored load:
staging capacity=1,000×1.5=1,500 RPSThat is a cluster sized for 1,500 RPS instead of one sized for the full 50,000 RPS production load, a large cost saving while still exercising a statistically meaningful, consistently-sampled slice of real traffic.
Trade-offs and pitfalls
- Async mirroring is the default for a reason. Synchronous (blocking) mirroring couples production latency to the mirror's health, which defeats the point; always fire-and-forget.
- Data contamination is the top real-world failure mode. A mirrored write that reaches a shared resource (a shared cache, a shared queue, a third-party API) is the most common way mirroring causes a production incident.
- Mirroring is not free. It roughly doubles egress and backend load for the sampled slice, so uncapped 100% mirroring at scale is rarely justified; sample deliberately.
- Mirror-cluster failures should never alert on-call for production. Route mirror telemetry to its own dashboards and alert channels, or a slow-to-fail staging service will generate noisy pages for a system nobody's users are touching.
You observe rising p99 latency on your load balancer while backends show stable p95 latency and healthy CPU. Walk through a troubleshooting checklist covering the network, the LB proxies themselves, TLS handshakes, accept/queue backlogs, kernel limits, and client behavior. What instrumentation would you add to pinpoint the root cause?
Sample Answer
Direct answer
When the load balancer's p99 rises but the backends' p95 and CPU stay flat, the divergence is itself the clue: the extra tail latency is being added somewhere the backend can't see, the network path, the LB's own accept/TLS/queueing layer, or client behavior, not inside request processing. Work through the request path in layers (network, LB proxy internals, TLS, kernel accept/backlog, client) and instrument each layer's own latency contribution separately, so the tail latency is attributed to a stage, not guessed at.
Layered checklist
- Network path (client to LB, LB to backend): interface errors, drops, retransmits (
ip -s link,netstat -s), VPC flow logs for SYN/retransmit spikes, targeted traceroutes from affected client geographies. A tail-latency-only symptom with no backend involvement often starts here. - LB/proxy internals: connection counts and churn per backend, event-loop stalls or worker saturation in the proxy process itself (profiling,
straceon accept), time-to-first-byte versus time-to-last-byte split per backend. A proxy that's CPU-saturated or GC-pausing (if it's a managed-runtime proxy) shows up here, invisible to backend CPU metrics. - TLS handshakes: split full-handshake latency from resumed/session-ticket handshake latency, and track the resumption hit rate; a drop in session cache hit rate (cache eviction, a cache not shared across proxy instances, client churn) turns cheap resumed handshakes into expensive full ones for a subset of requests, exactly the shape of a tail-latency-only regression.
- Accept queue / kernel backlog: listen backlog occupancy (
ss -ltn), SYN_RECV counts (connections stuck mid-handshake, waiting on the final ACK),somaxconnandtcp_max_syn_backloglimits, ephemeral port and TIME_WAIT counts. If the accept queue is intermittently near its limit, new connections queue briefly even though every connection that does get through is processed at normal speed, invisible to backend request-processing metrics by construction. - Client behavior: slow or high-RTT clients, retry storms, or a small subset of very chatty clients; correlate p99 offenders by client IP, geography, or user agent rather than assuming they're uniform across all traffic.
Worked example: why queueing produces a p99 problem with a flat p95
This is a general queueing-theory property, illustrated with a simple M/M/1 model (queueing-theory shorthand: Markovian/memoryless request arrivals, Markovian/memoryless service times, 1 server) and pinned parameters, not a claim about any specific measured system. For a queue with service rate μ and utilization ρ=λ/μ (arrival rate λ), the expected wait time in queue is:
Wq(ρ)=μ(1−ρ)ρFix μ=1000 (an arbitrary service-rate unit, just for the shape of the curve) and evaluate at three utilizations:
ρ=0.80ρ=0.98ρ=0.995:Wq=1000×0.200.80=0.0040⇒4.00 ms:Wq=1000×0.020.98=0.0490⇒49.00 ms:Wq=1000×0.0050.995=0.1990⇒199.00 msA move from ρ=0.80 to ρ=0.995 (a 24% increase in load) inflates queueing wait by roughly 50x. The requests that land during the brief windows where the accept queue or a proxy worker pool is momentarily near saturation (micro-bursts, GC pauses, a slow client holding a worker) are exactly the ones that generate the p99 tail, while the bulk of requests, arriving when the system is comfortably under its knee, still finish fast and keep p95 (and backend CPU, which averages over time) looking healthy. This is why p99 and p95 can diverge sharply even though nothing about the backend's steady-state processing changed: the tail is a queueing phenomenon, not a processing-time phenomenon.
Instrumentation to add
- Break end-to-end latency into stages, TCP connect, TLS handshake, proxy accept/queue wait, backend processing, response write, and emit each as its own histogram (not just the total), tagged by backend and client region.
- Accept-queue depth and backlog saturation as a time series, not just a point-in-time check, so a transient near-saturation event that self-resolves in under a second is still visible.
- TLS session-resumption rate as its own metric, separate from handshake latency, since a resumption-rate drop is a leading indicator of the handshake-latency problem.
- Kernel-level counters (SYN_RECV: mid-handshake connections, TIME_WAIT: recently closed connections still held by the kernel, retransmits) exported alongside application metrics on the same dashboard and timeline, so a kernel-layer cause doesn't require manually correlating two separate tools during an incident.
- Sampled packet captures triggered automatically when p99 crosses a threshold, so there's raw evidence from the actual bad window instead of trying to reproduce it after the fact.
Trade-offs and pitfalls
- Chasing this in backend-only dashboards is the classic dead end, since by construction the backend's own view (CPU, p95, request-processing time) is healthy; the cause lives in a layer that doesn't report through the backend's own telemetry.
- Averages and even p95 hide exactly this kind of problem by design, since it only affects a small fraction of requests, that's what makes it a p99 problem and not a p50 problem; don't let a "p95 looks fine" dashboard close the investigation.
- Fixing the wrong layer (e.g. scaling backend CPU when the problem is TLS resumption cache eviction) burns real money and doesn't move the metric; stage-by-stage instrumentation exists specifically to prevent this kind of misdiagnosis.
- Once queueing is confirmed as the mechanism, the fix is capacity or admission control (a bigger backlog, more proxy workers, load shedding, or reducing ρ by scaling out), not code-level micro-optimization of request handling, since request handling was never where the time went.
Design the load balancing layer for a service with no sticky-session requirement and autoscaling backends. Include L4-vs-L7 placement, service discovery integration, TLS termination, health checks, algorithm choice, and how you would validate the design with load testing. Then explain what changes if the service instead needs to be internet-facing across three regions with a 99.99% uptime target: edge proxies, regional failover and routing, and connection draining during rolling deploys.
Sample Answer
Direct answer
For a no-sticky-session, autoscaling internal service, put a lightweight L7 proxy close to the caller (sidecar or node-local), back it with a service registry that autoscaling updates directly, and pick a load-aware algorithm like least-request since request cost varies; making the same service internet-facing across three regions at 99.99% uptime then layers an edge/global tier on top for TLS termination, GeoDNS or Anycast entry, and cross-region failover, the internal design doesn't get thrown away, it becomes the inside of each region.
Part 1: internal load-balancing layer
graph TD
Client[Client] --> Proxy[Node-local Envoy Proxy]
Proxy --> Registry[Service Registry]
Registry --> CP[Control Plane xDS]
CP --> Proxy
Proxy --> BE1[Backend Instance 1]
Proxy --> BE2[Backend Instance 2]
Proxy --> BE3[Backend Instance N]
HC[Active Health Checks] --> Registry
L4 vs. L7 placement. Choose L7 (HTTP-aware) here specifically because there's no sticky-session requirement to preserve and because request cost is heterogeneous, L7 lets the proxy make a per-request decision (least-request, header-based routing) rather than pinning a whole TCP connection to one backend the way an L4 balancer would. L4 is the right call when connections are long-lived and uniform (e.g. raw TCP streams); that's not this case.
Service discovery integration. Backends register on startup and deregister on shutdown with a registry (Kubernetes Endpoints, Consul, or a cloud-native equivalent). The registry is the single source of truth the proxy layer subscribes to, autoscaling doesn't need its own LB-integration logic, it just needs to correctly start and stop instances; registration is what makes new instances visible. The control plane (labeled CP in the diagram above) is what pushes that registry state out to every proxy, commonly via xDS, Envoy's config-discovery protocol for streaming routing and endpoint updates to proxies in near real time.
TLS termination. For internal-only traffic, terminate TLS at the node-local proxy (mTLS between proxies) rather than at a far-away central point, this keeps the encrypted hop as short as possible while still giving every call in-transit encryption.
Health checks. Active checks (HTTP/gRPC readiness probe) plus passive checks (mark an endpoint unhealthy after consecutive 5xx or timeouts) feeding the registry, consistent with the readiness-gating principle: an instance only receives traffic after readiness passes, not merely liveness.
Algorithm choice. Least-request (or a dynamically-weighted round robin) over plain round-robin, because round-robin assumes uniform request cost, which explicitly isn't true here; least-request adapts to backends that are momentarily slower without needing an external signal.
Validating with load testing. Define success criteria before running anything:
- Baseline: ramp to expected peak RPS over several minutes, record p50/p95/p99 and error rate.
- Failure injection: kill a fraction of backend instances mid-test, confirm the registry and proxy converge to the healthy set and error rate recovers within a bounded number of health-check intervals.
- Autoscale test: sustain load past the scale-out threshold, confirm new instances only receive traffic after readiness passes, never merely because they are alive (this is the direct test of the readiness-gating principle stated above: an instance only receives traffic once its readiness probe passes, not merely once it is alive).
Part 2: what changes when internet-facing, 3 regions, 99.99% uptime
99.99% uptime⇒allowed downtime=(1−0.9999)×365×24×60≈52.6 minutes/yearThat budget is what forces every change below, roughly an hour a year, split across every deploy, every regional blip, and every DNS propagation delay.
| Aspect | Single-region internal (Part 1) | Internet-facing, 3 regions, 99.99% |
|---|---|---|
| Entry point | Node-local proxy, no public exposure | Edge proxy / managed LB per region, fronted by GeoDNS or Anycast (announcing the same IP address from multiple regions and letting network routing send each client to the topologically nearest one, entering the LB layer without a DNS lookup per region) for entry routing |
| TLS termination | At the node-local proxy (mTLS internally) | At the regional edge, public cert management, HSTS (HTTP Strict Transport Security, a header that forces browsers to only use HTTPS for the site), WAF (web application firewall, which filters malicious HTTP requests) |
| Failure domain | Single region; instance-level failures only | Must tolerate a whole region failing; requires cross-region failover, not just intra-region rerouting |
| Routing decision | Local proxy picks a backend instance | Two-tier: global layer picks a region, regional layer picks a backend instance within it |
| Connection draining | Drain on instance termination during rolling deploys, within one region | Same mechanism, but must also cover draining an entire region during a regional rolling deploy without breaching the uptime budget |
| Uptime math that matters | Not the binding constraint at this stage | 52.6 min/year budget directly bounds how much reliance on DNS propagation delay the design can afford: at a 60s TTL, worst-case failover staleness is about 60s; halving it to 30s tightens that bound but roughly doubles steady-state DNS query volume against the authoritative servers, since twice as many cached entries expire and get re-resolved per unit time, so shortening TTL for faster failover has a direct infrastructure-load cost, not a free dial |
Edge proxies. A regional edge (managed cloud LB or a dedicated edge proxy tier) now owns TLS termination, WAF, and the public entry point per region; the internal L7 layer from Part 1 sits behind it unchanged.
Regional failover and routing. GeoDNS or a global server load balancer selects a healthy region; within the selected region, traffic hits the same registry-backed L7 layer from Part 1. This is a strict superposition, Part 1's design becomes the inside of each region rather than being replaced.
Connection draining during rolling deploys. The mechanism doesn't change (deregister, then drain in-flight requests before terminating), what changes is scope: a regional rolling deploy must drain each instance without ever dropping a region's effective capacity below what the other two regions can absorb, otherwise a routine deploy risks tripping the same failure mode as a real regional outage.
Trade-offs and pitfalls
- Don't put a heavy central LB in front of every internal call at this scale. A single centralized LB at high internal RPS becomes both a latency tax on every call and an operational single point of failure; node-local/sidecar placement avoids both by keeping the hop local.
- The 52.6 minute/year budget makes DNS-only failover for the internet-facing tier a real risk, not just a design nicety. At a 60s TTL, worst-case failover staleness (about 60s) alone consumes over a third of the entire annual downtime budget in a single failover event; dropping TTL to 30s roughly halves that staleness bound but roughly doubles steady-state DNS query volume against the authoritative servers, so treat TTL as a concrete, numbered trade-off, not something any value is automatically "fast enough" for once the real availability math is on the table.
- Retries and hedging need limits. Unbounded retries at the internal L7 layer can turn one slow backend into cascading load on the rest of the fleet; cap retries, use jittered backoff, and reserve hedging (duplicate a request to a second backend) for genuinely latency-sensitive paths only.
- Load testing that only exercises the happy path won't validate the uptime target. The tests that actually matter for a 99.99% target are the failure-injection and rolling-deploy drain tests, not just a peak-RPS ramp.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.