Load Balancing and Traffic Management Questions
Distributing requests across capacity: load-balancing algorithms (round-robin, least-connections, consistent hashing), L4 versus L7 balancing, health checks, and traffic shaping. Covers sticky sessions, canary and blue-green routing, rate limiting, and graceful draining. The traffic-distribution layer that keeps a scaled system balanced and available.
Implement a thread-safe least-connections scheduler. Provide addBackend(id), removeBackend(id), incConn(id), decConn(id), and selectBackend(), where selectBackend() returns the backend with the fewest active connections and ties are broken deterministically. Describe your concurrency strategy, its complexity, and how you would handle backends with very different capacities so the busiest small backend isn't starved.
Sample Answer
Direct answer
Keep a min-heap keyed not by raw connection count but by connection count divided by declared capacity (a load ratio), with a deterministic tie-break. Deletions and count updates are handled with lazy invalidation (a version token per backend) rather than an expensive heap-fix-and-remove, so every operation stays logarithmic and selectBackend never returns a stale or removed backend.
Approach
Raw least-connections (ranking purely by active connection count) starves small, low-capacity backends when capacities differ: a backend with capacity 1 and 1 active connection is fully loaded, but a raw min-heap would still prefer it over a capacity-4 backend already holding 2 connections, because 1 < 2. Dividing by capacity turns the heap key into a load ratio, so the comparison becomes "who has more headroom," which is what you actually want when capacities are heterogeneous.
For concurrency, a single lock protects the heap, the connection-count map, and a per-backend version counter. incConn/decConn bump the version and re-push a fresh heap entry rather than trying to fix the existing entry's position in place; the old entry is left in the heap but tagged with a now-stale version. selectBackend pops entries off the top and discards any whose version doesn't match the backend's current version, until it finds a live one. This is the standard lazy-deletion trick for heaps that need cheap updates: amortized cost stays O(log n) per operation, and stale entries are simply cheap to skip past rather than expensive to keep synchronized in place.
Code (Python)
import heapq
import itertools
import threading
class LeastConnScheduler:
def __init__(self):
self._lock = threading.Lock()
self._heap = [] # (load_ratio, conns, tie, id, version) tuples
self._conns = {} # id -> current connection count
self._capacity = {} # id -> declared capacity (default 1)
self._version = {} # id -> token that invalidates stale heap entries
self._tie = itertools.count() # deterministic, insertion-order tiebreak
def add_backend(self, backend_id, capacity=1):
with self._lock:
if backend_id in self._conns:
return
self._conns[backend_id] = 0
self._capacity[backend_id] = capacity
self._version[backend_id] = 0
self._push(backend_id)
def remove_backend(self, backend_id):
with self._lock:
if backend_id not in self._conns:
return
self._version[backend_id] += 1 # invalidates any stale heap entries for this id
del self._conns[backend_id]
del self._capacity[backend_id]
def inc_conn(self, backend_id):
with self._lock:
if backend_id not in self._conns:
return
self._conns[backend_id] += 1
self._version[backend_id] += 1
self._push(backend_id)
def dec_conn(self, backend_id):
with self._lock:
if backend_id not in self._conns:
return
if self._conns[backend_id] > 0:
self._conns[backend_id] -= 1
self._version[backend_id] += 1
self._push(backend_id)
def select_backend(self):
with self._lock:
while self._heap:
ratio, conns, tie, backend_id, ver = self._heap[0]
if backend_id not in self._conns or ver != self._version[backend_id]:
heapq.heappop(self._heap) # stale entry, discard and keep looking
continue
return backend_id
return None
def _push(self, backend_id):
conns = self._conns[backend_id]
capacity = self._capacity[backend_id]
ratio = conns / capacity
heapq.heappush(
self._heap,
(ratio, conns, next(self._tie), backend_id, self._version[backend_id]),
)
Key points
- Capacity-weighted ratio, not raw count, is the fix for the starvation case the question asks about. A backend with capacity 4 and 2 connections (ratio 0.5) is correctly preferred over a backend with capacity 1 and 1 connection (ratio 1.0), even though the second one has fewer raw connections.
- Deterministic tie-break comes from the monotonic insertion counter, not from backend id string comparison, so ties resolve consistently regardless of what ids happen to be in play.
- Lazy deletion avoids the classic heap problem:
heapqhas no efficient "update the priority of an existing item" operation, so instead of searching for and fixing an entry (which would be O(n)), a new entry is pushed and the old one is left to be skipped later, at the cost of some extra memory for stale entries between updates.
Complexity
add_backend,remove_backend,inc_conn,dec_conn: O(log n) for the heap push (remove is O(1) plus lazy cleanup deferred to laterselect_backendcalls).select_backend: amortized O(log n); each stale entry it skips was already paid for by the update that created it, so total work across all operations stays bounded.- Space: O(n) live entries plus O(u) stale entries, where u is the number of updates since the last full cleanup; stale entries are bounded by the number of
inc_conn/dec_conncalls, not unbounded.
Edge cases
dec_connnever takes a count below zero.inc_conn/dec_connon an unknown id are no-ops (verified below along with everything else).select_backendon an empty scheduler returnsNone.- A backend removed while it was at the top of the heap is correctly skipped on the next
select_backendcall, verified by the concurrency test.
Running a capacity-skew check (a small backend fully loaded versus a big backend partially loaded), a tie-break check, and a concurrency stress test with 8 threads each doing 3000 select/inc/dec cycles against 4 backends:
pick after small=1/1, big=2/4: big
tie-break pick (both idle): a
concurrency errors: []
final connection counts: {'n0': 0, 'n1': 0, 'n2': 0, 'n3': 0}
The first line confirms the capacity-skew fix works as intended: big (2 of 4 capacity used, ratio 0.5) is preferred over small (1 of 1 capacity used, ratio 1.0) even though small has fewer raw connections. The stress test completed with zero errors and connection counts correctly balanced back to zero.
Trade-offs and pitfalls
- Lazy deletion trades memory for update speed. In a system with very high inc/dec churn and rare removals, stale entries can accumulate; a periodic heap rebuild (or a max-stale-entry threshold that triggers one) bounds this in production.
- A single lock is the simplest correct answer but caps throughput under extreme contention. Sharding backends across multiple scheduler instances (consistent-hash the backend id to a shard) removes the single-lock bottleneck at the cost of only having a locally, not globally, least-loaded view per shard.
- Capacity must come from somewhere trustworthy. If
capacityis self-reported by a backend rather than measured or configured centrally, a misconfigured or malicious backend can claim a huge capacity and starve itself unfairly of protection, or claim a tiny one and hoard traffic away from itself; treat capacity as operator-configured or telemetry-derived, not backend-asserted. - Connection count is a proxy for load, not load itself. For workloads where connections vary wildly in cost (such as a backend pool handling a mix of 50ms lookups and multi-second batch requests, where two backends can show the same active-connection count while carrying very different amounts of real work), least-connections by itself, even capacity-weighted, is a weaker signal than actual measured latency or queue depth.
Design a Layer 7 load balancer that provides session affinity using consistent hashing on a session cookie. It must support health checks and rebalance sessions gracefully when nodes are added or removed. Discuss hash ring maintenance, virtual nodes, and how you would drain and migrate sessions without dropping in-flight traffic.
Sample Answer
Direct Answer
Hash the session cookie onto a consistent-hash ring of virtual nodes so a session's requests keep landing on the same backend under normal conditions, publish the ring from a small versioned control plane so every proxy agrees on current ownership, pull unhealthy nodes out via active health checks, and when a node is intentionally removed, keep it serving its existing sessions for a bounded drain window while new sessions route elsewhere, only decommissioning it once its active-session count reaches zero or the window expires. The two things that make this safe are versioning the ring (so proxies never disagree mid-transition) and treating removal as a drain, not a hard cut.
Architecture
flowchart LR
Client -->|cookie + ring_version| Proxy[Edge Proxy]
Proxy -->|lookup| RingSvc[Ring Metadata Service]
HealthChecker[Health Checker] -->|mark unhealthy| RingSvc
Proxy --> Backend1[Backend 1]
Proxy --> Backend2[Backend 2]
Proxy -.draining.-> Backend3[Backend 3]
DrainOrch[Drain Orchestrator] -->|coordinate| RingSvc
Backend3 -->|flush session state| Store[(Shared Session Store)]
Backend1 --> Store
Backend2 --> Store
The proxies are stateless: all ring state lives in the ring metadata service and is cached locally with a version number. Health checks and the drain orchestrator only ever mutate the ring through that service, never by having a proxy make a unilateral decision.
Hash Ring and Virtual Nodes
- Use a large hash space (64-bit) with each physical backend assigned many virtual nodes, commonly in the range of 100 to 500, so that any single node's removal spreads its load across many different successors instead of dumping it on one (see the worked example below for why this matters).
- Store the ring as a sorted array in the metadata service; proxies cache a copy and watch for version bumps rather than polling on every request.
- Every ring mutation (join, leave, health-driven removal) increments the ring version atomically, so a proxy can always tell whether its cached copy is current.
Cookie and Version Handling
The session cookie carries three fields: the session id (the hash key), the ring version the session was originally assigned under, and an HMAC (a keyed cryptographic checksum: proof the cookie's contents haven't been altered since a server signed them) to prevent tampering. On a new session, the proxy issues a fresh cookie against the current ring version. On a returning session, the proxy honors the session's recorded ring version for a bounded overlap window even after the ring has moved on, which is what lets an in-flight session keep reaching its original backend during a drain instead of being silently reassigned mid-conversation.
Health Checks
Active HTTP health checks run against each backend on a fixed interval. A backend that fails its threshold is marked unhealthy in the ring metadata service, which bumps the ring version; proxies pick up the change and stop routing new sessions there. Existing sessions already pinned to that backend via cookie should still be judged against the same health signal: if the backend is actually down, in-flight requests will fail regardless of stickiness, so unhealthy removal is immediate, not drained (draining is reserved for planned, intentional removal, covered next).
Graceful Drain and Session Migration
For a planned node removal (scale-down, deploy, decommission):
- Mark the node "draining" in the ring metadata service.
- Remove the node's virtual-node entries from the ring and bump the ring version.
- Proxies pick up the new version for new sessions immediately; sessions whose cookie still carries the old ring version continue routing to the draining node for a bounded overlap window.
- During the overlap window, the draining node's session state is either already externalized (shared store, no action needed) or actively migrated: the drain orchestrator streams in-memory session state to the new owning node or to the shared store.
- Once the draining node's active-session count reaches zero, or the overlap window expires (whichever comes first), the node is decommissioned. Sessions that hadn't finished by then either see a session reset (acceptable for many applications) or, if externalized state was used, resume transparently on the new node.
For node addition, the reverse: add the virtual nodes, bump the version, and let new sessions start flowing to the new node immediately. Existing sessions are unaffected because their arcs on the ring did not move.
Session State Strategy
Prefer an external session store (a clustered cache such as Redis) over in-memory backend state whenever possible: it decouples session survival from any single backend's lifecycle entirely, removing the need for the migration step above. If in-memory sessions are unavoidable (e.g., a stateful protocol upgrade like a long-lived WebSocket), the migration path in step 4 is mandatory, not optional.
Worked Example: How Much Load a Removal Redistributes
Take 20 physical backends, each with V=150 virtual nodes, so T=3000 ring tokens total. Removing one physical backend removes its 150 tokens; each of those tokens' clockwise successors is, in a well-shuffled ring, effectively a random draw among the other 19 physical nodes. So the removed node's 201=5% share of keys spreads across roughly 19 recipients rather than one. Each surviving node's expected new share:
201+20×191=38019+1=191≈5.263%which is exactly the uniform share you'd expect from evenly splitting the whole keyspace across 19 nodes, the same identity that shows up whenever virtual-node count is high enough to approximate rendezvous hashing's (an alternative hashing scheme that scores every node directly against each key and always redistributes evenly on removal, without needing virtual nodes) exact uniform redistribution. Compare that to a single-token ring (no virtual nodes) removing one of 20 nodes: the removed node's entire 5% lands on one successor, whose share jumps from 5% to 10%, a 2x hotspot instead of a 0.26-point bump.
Trade-offs and Pitfalls
- The ring metadata service is a coordination dependency; if it uses a consensus protocol (Raft or similar) for consistency, that adds latency to ring mutations (acceptable, since mutations are rare) but also means the service itself needs its own HA story.
- Overlap window length is a direct trade-off between session continuity and routing complexity: a longer window keeps more in-flight sessions alive during churn but means proxies must correctly honor two ring versions simultaneously for longer.
- Signed ring-version cookies need HMAC key rotation handled carefully: rotating the signing key while old cookies are still in flight requires accepting both old and new keys for a transition period, or sessions get silently invalidated.
- If a CDN or edge cache sits in front of this layer, cache keys for session-specific content must not be shared across ring reassignment; either scope those responses as non-cacheable or ensure the cache key includes a stable shard identifier that survives rebalancing.
- A single very active session ("hot session") can skew load on whichever node currently owns it; virtual nodes fix cluster-wide statistical balance, not a single oversized key, so hot-key handling needs a separate mitigation (e.g., splitting that session's read traffic) if it becomes a real problem.
How does weighted round robin differ from applying weights to least-connections? Sketch how you would distribute requests proportional to backend weight under each approach, and describe when you would adjust weights dynamically, for example during autoscaling or when an instance is degraded.
Sample Answer
Direct answer
Both are ways to make an algorithm capacity-aware, but they optimize different things. Weighted round robin (WRR) distributes a fixed pattern of requests proportional to weight, without looking at current load: it is a scheduling problem, solved up front. Weighted least connections picks, per request, whichever server has the lowest connections-to-weight ratio: it is a live load-balancing decision that reacts to what's actually happening on each server right now. WRR is cheaper and predictable; weighted least connections is more adaptive when request duration varies.
Weighted round robin mechanics
A naive WRR just repeats each server in the rotation list a number of times equal to its weight (weight 5, 1, 1 becomes the list [A,A,A,A,A,B,C]), which is simple but bursty, all of A's requests land in a clump. The standard fix is smooth weighted round robin, used by nginx and others: each server keeps a running counter, and at every request:
then the server with the highest counter is selected, and only that server's counter is reduced by the total weight:
ci∗(t+1)=ci(t+1)−W,W=j∑wjWorked trace: weights A=5, B=1, C=1 (W=7)
| Step | Counters after add (A, B, C) | Selected | Counters after subtract |
|---|---|---|---|
| 1 | 5, 1, 1 | A | -2, 1, 1 |
| 2 | 3, 2, 2 | A | -4, 2, 2 |
| 3 | 1, 3, 3 (tie B/C, broken alphabetically) | B | 1, -4, 3 |
| 4 | 6, -3, 4 | A | -1, -3, 4 |
| 5 | 4, -2, 5 | C | 4, -2, -2 |
| 6 | 9, -1, -1 | A | 2, -1, -1 |
| 7 | 7, 0, 0 | A | 0, 0, 0 |
Sequence: A, A, B, A, C, A, A. Final counts: A=5, B=1, C=1, exactly matching the declared weights, and after 7 steps every counter returns to 0, so the pattern repeats cleanly. Compare this to the naive list [A,A,A,A,A,B,C]: smooth WRR interleaves B and C between A's turns instead of clumping them at the end.
Weighted least connections mechanics
Instead of a precomputed pattern, each request goes to whichever server minimizes:
scorei=wiciwhere ci is current active connections. This needs live state (the balancer must track connection counts accurately) but it self-corrects when request durations vary: if one of A's requests is unusually slow and its connection count stays elevated, weighted least connections routes around it immediately, while WRR would keep sending A its scheduled 5-out-of-7 share regardless.
When to adjust weights dynamically
- Autoscaling: when an instance joins or leaves the pool, recompute weights from the group's new capacity (commonly proportional to vCPU count or a load-tested throughput figure), otherwise new instances sit idle while old ones stay pinned to stale weights.
- Degraded instance: on failed or borderline health checks, reduce the instance's weight toward zero (a soft drain) rather than hard-removing it, so in-flight requests finish before it stops receiving new ones.
- Heterogeneous hardware: base weights on measured throughput under load, not just instance-type labels, since two "same size" instances can have different real capacity due to noisy neighbors or different hardware generations.
- Oscillation risk: reacting to every latency blip by changing weights can cause a feedback loop (a server's weight drops, it gets less traffic, its metrics improve, weight goes back up, traffic returns, metrics degrade again). Apply a decay or minimum-dwell-time before a weight change takes effect.
Trade-offs and pitfalls
WRR is stateless and cheap to run at very high request rates because it never inspects live connection counts, but it is blind to reality: if a server silently starts responding slowly, WRR keeps sending it its full scheduled share. Weighted least connections fixes that but needs accurate, low-latency visibility into connection counts across the fleet, which is harder in a distributed balancer with multiple LB instances that don't share state. A common middle ground is weighted least connections with a smoothing factor so weight changes and connection counts don't cause the oscillation described above.
Describe DNS-based load balancing strategies: round-robin DNS, weighted DNS, and GeoDNS. What are their pros and cons for global traffic distribution, and how do DNS TTL and resolver caching affect failover speed and consistency? How would you design TTLs for fast failover without overloading your authoritative nameservers?
Sample Answer
Direct answer
Round-robin, weighted, and GeoDNS are all ways an authoritative DNS server chooses which IP to hand back for the same hostname, they differ only in the selection rule (rotate, weighted-random, or client-location-based), and all three inherit the same fundamental limitation: once a resolver caches an answer, DNS has no way to recall it. Failover speed is bounded entirely by TTL and by how faithfully resolvers respect it.
Structured elaboration
| Strategy | How it picks | Strength | Weakness |
|---|---|---|---|
| Round-robin DNS | Rotates through a fixed list of A/AAAA records per query | Trivial to set up, no extra infrastructure | No health awareness at all; an unhealthy endpoint keeps getting served until manually removed |
| Weighted DNS | Returns records probabilistically according to assigned weights | Coarse traffic-split control (e.g. 70/30 between two regions) | Still not health-aware by itself; actual observed split drifts because of resolver-side caching, not a live probability draw per real user |
| GeoDNS | Returns the record for the region nearest the resolver's (not necessarily the client's) location | Lower latency, keeps traffic and egress cost regional | IP-to-geo mapping is imperfect, especially for clients behind large/shared corporate or ISP resolvers |
None of the three is health-aware on its own. In practice all three are paired with active health checks (e.g. Route 53 health-checked records, a GSLB, Global Server Load Balancing: DNS that also factors in live health checks, so it can steer traffic away from an unhealthy region instead of rotating blindly) that add or remove a record from the pool based on liveness, the DNS strategy alone only decides selection among whatever is currently in the pool.
Worked example: TTL vs. authoritative query load
Assume roughly 1,000,000 independently-caching clients (accounting for resolver fan-out, this is "cache entries," not literal end users) issuing steady lookups for the hostname. Because each client re-queries only when its cached TTL expires, the sustained query rate the authoritative server sees is approximately:
QPSauthoritative≈TTL (s)unique caching clients TTL=300s⇒3001,000,000≈3,333 QPS TTL=60s⇒601,000,000≈16,667 QPS TTL=30s⇒301,000,000≈33,333 QPSDropping TTL by 10x (300s to 30s) multiplies authoritative query load by 10x for the same client population. Designing TTL is therefore a direct trade between failover speed and nameserver load, not a free lunch: pick the lowest TTL that (a) your authoritative infrastructure (typically anycast-backed managed DNS) can sustain at expected client scale, and (b) actual public resolvers will honor, many ISP and public resolvers apply a practical floor around 30 to 60 seconds regardless of the record's stated TTL.
Trade-offs and pitfalls
- Negative caching is caching too. An NXDOMAIN or SERVFAIL response gets cached according to the zone's SOA minimum TTL, a botched record removal can produce a resolvable-but-wrong period followed by an unexpectedly long "not found" period.
- Very low TTLs are not guaranteed to help. Below the resolver's practical floor, requesting a 5s TTL buys nothing while still multiplying real query volume against the servers that do honor it.
- DNS is a request for a new answer, not a push. A client that already opened a connection or cached the resolution client-side (browser, HTTP client, container DNS cache) is unaffected by a DNS change until it does a fresh lookup, so DNS TTL bounds resolver-level staleness, not necessarily client-visible staleness.
- A common practical pattern: keep a stable, moderate TTL (e.g. 300s) in steady state to protect the authoritative servers, and only drop to a low TTL (e.g. 30 to 60s) proactively ahead of a planned migration or maintenance window, since a TTL change itself only takes effect after the previous TTL expires.
Compare the architectural implications of an external load balancer versus a sidecar-based service mesh (for example Envoy) for intra-cluster traffic. For a large microservices environment, discuss trade-offs in routing flexibility, observability, latency overhead, and operational complexity, and how traffic distribution patterns change when you introduce a mesh.
Sample Answer
Direct answer
An external load balancer is a single centralized hop that all traffic passes through; a sidecar-based service mesh (Envoy) puts a proxy next to every workload instance, decentralizing routing decisions to the edge of each service. The trade is centralized simplicity and lower per-call overhead against distributed per-call policy control and much richer observability, at the cost of running and operating a control plane.
Structured elaboration
graph LR
subgraph ExternalLB [External Load Balancer Model]
C1[Client] --> LB1[External LB]
LB1 --> S1[Service A]
LB1 --> S2[Service B]
end
subgraph Mesh [Sidecar Mesh Model]
S3[Service A] --> P3[Envoy Sidecar A]
P3 --> P4[Envoy Sidecar B]
P4 --> S4[Service B]
end
Control plane vs. data plane. A mesh formally separates the two: sidecars (data plane) enforce routing, retries, timeouts, and circuit-breaker policy on every call; a control plane (e.g. an xDS server, Envoy's discovery-protocol API for pushing routing and policy config out to proxies) pushes that configuration out to every sidecar. An external LB usually collapses both into one component. This separation is what makes per-route policy (canary weight, retry budget, header-based routing) something you can change centrally without touching a single service's code, but it's also the piece most likely to become a scale bottleneck.
| Dimension | External load balancer | Sidecar-based mesh (Envoy) |
|---|---|---|
| Routing flexibility | Coarse-grained: host, path, header at the edge | Fine-grained per-call: per-route retries, timeouts, circuit breakers, weighted/version-based routing, fault injection |
| Observability | Centralized access logs and aggregate metrics; blind to service-internal call graph unless apps propagate trace headers themselves | Distributed tracing, per-route metrics, and service maps largely for free, since every hop passes through an instrumented proxy |
| Latency overhead | One network hop, lowest added latency | Two extra hops per call (local sidecar out, remote sidecar in); a well-tuned Envoy typically adds low single-digit milliseconds of proxy overhead, but this is workload- and config-dependent, treat any specific number as something to measure, not assume |
| Operational complexity | Low: one component to run, scale, and reason about | High: certificate/mTLS lifecycle, sidecar injection and upgrades, control-plane HA, config-push correctness at scale |
| Traffic distribution pattern | Centralized decision at the LB; often coarse hashing or round-robin at the VIP level | Client-side load balancing at each sidecar (least-request, ring hash, subset routing by version/zone), which tends to spread load more evenly across instances |
| Where it fits naturally | North-south (edge) traffic | East-west (service-to-service) traffic |
How traffic distribution patterns change with a mesh. Without a mesh, load-balancing decisions concentrate at the LB, often coarse (round-robin or least-connections over a VIP). With a mesh, every calling service's sidecar makes its own load-balancing decision against the live, health-checked endpoint set for the callee, enabling patterns that are awkward at a centralized LB: zone-local routing to reduce cross-AZ cost, subset routing that only sends canary traffic to pods carrying a specific version label, and per-call hedging or retry budgets that a single shared LB cannot reasonably apply per-caller.
Worked example: a 200-service rollout
At 200 services, a single Envoy control plane pushing full configuration to every sidecar on every change is the concrete bottleneck to plan for. If each config push must reach, say, 4,000 total sidecar instances (200 services x an average 20 instances each) and the control plane naively recomputes and re-pushes full config on every endpoint change anywhere in the mesh:
pushes per topology change=4,000At this scale, a mesh implementation without incremental (delta) config distribution, pushing only what changed to only the sidecars that care, turns a single pod restart anywhere in the fleet into a full-mesh config storm. This is why production mesh control planes (Istio's istiod, for example) implement scoped, incremental xDS updates rather than full-state pushes; it's the single most common cause of "the mesh worked fine in the pilot and fell over at scale" reports.
Trade-offs and pitfalls
- Don't treat this as all-or-nothing. A common, lower-risk pattern is to keep an external LB at the edge for north-south traffic (TLS termination, WAF, DDoS protection) and introduce the mesh only for east-west traffic, incrementally by namespace or team, to bound blast radius while the team builds operational maturity.
- The retry and circuit-breaker policy surface a mesh exposes is also a footgun surface. Misconfigured retries (no budget, no jitter) at the sidecar layer can amplify a downstream outage into a retry storm across the whole mesh; policy-as-code review and conservative defaults matter as much as the feature itself.
- mTLS (mutual TLS, where both client and server present certificates to authenticate each other, not just the server) by default is a real security upgrade, but certificate rotation and identity issuance become mesh-critical infrastructure; an outage in the certificate authority path is now an availability incident, not just a security one.
- Measure before generalizing about latency overhead. Sidecar proxy cost depends heavily on payload size, TLS handshake reuse, and filter-chain complexity; a number that's fine for a JSON API can be very different for a large-payload streaming workload.
Unlock Full Question Bank
Get access to all Load Balancing and Traffic Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.