Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
Describe how a service mesh such as Istio or Linkerd integrates with service discovery and dynamic configuration. Explain how sidecars discover backends, how the control plane distributes routing/policy config, and trade-offs around performance, operational complexity, and discoverability when adding a mesh to an existing platform.
Sample Answer
Direct answer
A service mesh moves discovery and traffic configuration out of application code and into proxies. A control plane watches the platform's own source of truth (in Kubernetes: Services, EndpointSlices and pod labels, plus the mesh's routing resources) and translates it into configuration that it pushes to a proxy next to every workload. The application calls a plain hostname; the local proxy already knows the healthy backend IPs, the routing rules, retry and timeout policy, and the certificates for mTLS. You gain uniform routing, security and telemetry for every language; you pay in per-pod CPU and memory, an extra network hop per side, a control plane whose config push can itself become a scaling problem, and a system that is harder to debug because the real routing decision now lives in proxy state you cannot see in your code.
Terms
- Service mesh: a layer of proxies that handle service-to-service traffic, plus a control plane that configures them. Istio and Linkerd are the two most common on Kubernetes.
- Sidecar: a proxy container injected into each application pod. Traffic in and out of the pod is redirected through it (typically using iptables, the Linux kernel's packet-filtering and routing rules, configured at pod start to send traffic through the proxy first).
- Data plane vs control plane: the proxies that carry requests are the data plane; the component that tells them what to do is the control plane (
istiodin Istio; thedestination,identityand proxy-injector components in Linkerd). - mTLS (mutual TLS): both sides of a connection present certificates, so the server knows which workload is calling, not just that the traffic is encrypted.
- xDS: the family of discovery APIs Envoy proxies use to receive configuration. Envoy is the specific proxy Istio injects as the sidecar; xDS covers LDS (listeners), RDS (routes), CDS (clusters, meaning upstream services) and EDS (endpoints, meaning the IPs behind a cluster).
- Kubernetes Service, EndpointSlice, pod labels: the platform's own primitives the mesh builds on. A Service is a stable name for a group of pods; an EndpointSlice is the live, auto-updated list of that group's actual pod IPs and their readiness; pod labels are the key/value tags (
app: reviews) Services and mesh rules select pods by. - gRPC: a binary remote-procedure-call protocol; the control plane uses it as the transport for pushing config and endpoints to every sidecar.
How sidecars discover backends
flowchart LR
K8s[Kubernetes API: Services, EndpointSlices, mesh custom resources] -->|watch| CP[Control plane]
CP -->|push config and endpoints over gRPC| PA[Sidecar in pod A]
CP -->|push| PB[Sidecar in pod B]
AppA[App A] -->|calls reviews:9080| PA
PA -->|mTLS, load balanced| PB
PB --> AppB[App B, a reviews pod]
Istio (Envoy sidecars):
istiodwatches Services and EndpointSlices, so it knows every pod IP behind every Service and which are ready.- It turns that into xDS: a cluster for each service, endpoints (EDS) listing healthy pod IPs, and routes for any VirtualService rules.
- Each Envoy sidecar holds a long-lived gRPC stream to
istiod. When a pod becomes ready or is deleted,istiodpushes an endpoint update and the sidecar's load balancer changes within seconds, without DNS TTLs and without the app doing anything. - The app still resolves
reviewsthrough normal cluster DNS, but the sidecar intercepts the connection and picks the actual backend from its own endpoint list.
Linkerd (its own Rust micro-proxy, linkerd2-proxy):
The proxy asks the control plane's destination service, over gRPC, for the endpoints and policy of each destination the application actually talks to, and keeps a watch open for that destination. That on-demand model means a proxy only holds state for the services its workload uses, which is a real difference from Istio's default behaviour.
How the control plane distributes routing and policy config
The two resources to actually know first are VirtualService and DestinationRule: together they cover almost every day-to-day routing decision. The rest of this list matters once you also need identity-based security policy or aren't yet on Istio's newer routing APIs.
- Routing: in Istio, VirtualService (match rules, weighted splits such as "5% to v2", retries, timeouts, fault injection: deliberately injecting errors or delay into some requests to test how callers handle it) and DestinationRule (load-balancing algorithm, connection pools, outlier detection which ejects endpoints that keep failing, and subsets such as
v1/v2by label). In both meshes, the Kubernetes Gateway API's HTTPRoute can express traffic splits too. - Security policy: PeerAuthentication (require mTLS) and AuthorizationPolicy (which identities may call which paths) in Istio; Server and AuthorizationPolicy resources in Linkerd. Workload certificates are issued and rotated automatically by the control plane.
- Distribution: operators apply these as Kubernetes resources; the control plane validates and translates them and pushes the result. This is dynamic configuration delivery: a traffic shift from 5% to 50% is an API write, effective fleet-wide in seconds, with no redeploy.
Worked example: why config scope matters at scale
A cluster has 500 services and 5,000 pods, and each workload really calls about 10 services with about 10 endpoints each.
- Default Istio behaviour sends each sidecar configuration for the whole mesh. Endpoint entries held across the fleet: 5,000 sidecars × 5,000 endpoints = 25,000,000. Every endpoint change also triggers a push to every sidecar.
- Scoped (Istio's
Sidecarresource restricting each workload to the hosts it uses, ordiscoverySelectorslimiting which namespaces are visible): 5,000 × (10 × 10) = 500,000 entries, 50 times less proxy memory and push work. - Linkerd's on-demand lookups give roughly the scoped behaviour by default.
The lesson: with a push-everything control plane, proxy memory and control-plane CPU grow roughly with (number of proxies) × (size of the mesh). Both of those factors grow with the same underlying number, how many pods are in the cluster, so doubling the fleet does not double the cost, it roughly quadruples it: that is what "quadratic as the platform grows" means here. Scoping is not an optimisation to add later; it belongs in the first rollout.
Trade-offs when adding a mesh to an existing platform
| Dimension | Cost | Mitigation |
|---|---|---|
| Performance | Two extra proxy hops per call (client sidecar and server sidecar), plus TLS handshakes; per-pod CPU and memory for every sidecar | Measure p99 (99th percentile) latency on your own traffic before and after on a pilot namespace; do not trust vendor benchmarks. Consider Istio's ambient mode (generally available since Istio 1.24), which replaces per-pod sidecars with a per-node ztunnel (a shared node-level proxy handling basic encrypted routing for every pod on that node) for L4 (TCP-level mTLS and policy) and optional per-namespace waypoint proxies (added only where a namespace needs HTTP-level routing decisions) only where L7 (HTTP-level routing) is needed |
| Operational complexity | A new critical control plane to upgrade and monitor; injection webhooks (the Kubernetes mechanism that automatically adds the sidecar container to a pod as it is created, so nobody edits every deployment by hand); sidecar start-up ordering (the app starts before its proxy is ready, or the proxy exits before the app finishes draining, meaning finishing in-flight requests and closing connections cleanly) | Pilot one namespace; use native sidecar support (a Kubernetes feature that starts and stops sidecar containers in the right order automatically, instead of relying on timing tricks) in recent Kubernetes versions for start/stop ordering; keep the mesh version upgrade as a rehearsed runbook |
| Discoverability (debugging) | The routing decision is no longer in code or DNS; a request can fail because of a DestinationRule nobody on the app team knows about | Standard tooling: istioctl proxy-config endpoints <pod> and istioctl proxy-status (is this sidecar in sync with istiod?), linkerd viz for per-route metrics; mesh resources owned in the same repo as the service |
| Two discovery systems | During migration, meshed pods use proxy-pushed endpoints while unmeshed callers still use DNS or a client-side library; their views of "healthy" can differ | Migrate by call path (callee first, then callers) and keep outlier detection consistent with existing health checks |
| Control-plane failure | If istiod is down, sidecars keep their last config and traffic keeps flowing, but no new endpoints or policy arrive, so scaling events go unseen | Run the control plane highly available; alert on proxy config staleness |
Recommendation
For an existing platform, adopt a mesh when you need at least two of: uniform mTLS with identity-based policy, traffic shifting for progressive delivery (rolling a change out to a growing slice of traffic while watching for regressions, rather than to everyone at once), and consistent retries, timeouts and telemetry across several languages. If you only need discovery and load balancing, Kubernetes Services plus a good client library are cheaper. If you do adopt: start with one namespace, scope config from day one, and prefer the lighter footprint (Linkerd, or Istio ambient) unless you specifically need Envoy's L7 feature set everywhere.
You're setting up alerting for a Consul- or etcd-backed service registry that several teams' services depend on for discovery and config. What would you monitor to catch degradation before it causes an outage, and what would make you page someone versus just log it?
Sample Answer
Direct answer
Monitor the registry at three layers: the consensus layer (does the cluster have a leader and enough healthy members to keep it), the storage layer (is the disk fast enough and is the database nowhere near its size limit), and the consumer layer (can client services actually resolve names and read fresh config). Page a human only for conditions that are already hurting consumers or that leave the cluster one failure away from losing quorum; everything that is merely unusual goes to logs, dashboards or a next-business-day ticket.
Terms first
- Service registry: the database (here Consul or etcd) where running instances announce "I am
paymentsat 10.0.3.14:8443", and where other services look those addresses up. Many teams also keep dynamic configuration in the same store. - Raft, leader, quorum: Consul servers and etcd members agree on every write using the Raft consensus protocol. One member is the leader and orders all writes. A write only commits once a quorum (a strict majority,
floor(n/2) + 1members) has stored it. A 3-member cluster needs 2 alive; a 5-member cluster needs 3. - Leader election: when followers stop hearing from the leader they hold an election. While there is no leader, no writes succeed: registrations, deregistrations and config changes all stall.
- fsync: forcing a write to physical disk. Raft cannot acknowledge a write until its log entry is fsynced, so a slow disk makes the whole cluster slow.
- Watch / blocking query: a long-lived request that returns when data changes. Clients use them to learn about new instances without polling hard.
- SLO (service-level objective): the reliability target you promise, for example "99.9% of name lookups succeed".
- Error budget / burn rate: the error budget is the failure a 99.9% SLO allows you, up to 0.1% of requests; the burn rate is how fast you are spending it. "Burning error budget fast" means failures are happening quickly enough to exhaust the whole budget before the SLO's period ends, which is why it pages before the SLO is technically breached.
What to monitor
1. Consensus health (the cluster itself)
| Signal | etcd | Consul |
|---|---|---|
| Is there a leader? | etcd_server_has_leader (0 or 1 per member) | consul.raft.state.leader increments; consul.raft.leader.lastContact shows how long followers go without hearing from the leader |
| Election churn | etcd_server_leader_changes_seen_total | consul.raft.state.candidate (a server started an election) |
| Headroom before quorum loss | count of healthy members vs quorum | consul.autopilot.failure_tolerance (voting servers you can still lose) and consul.autopilot.healthy |
| Failed writes | etcd_server_proposals_failed_total | consul.client.rpc.failed, consul.client.rpc.exceeded (rate-limited) |
Failure tolerance is the single most useful number here. It turns "one member is down" from a vague worry into a precise statement: a 5-server cluster with one down still tolerates one more failure; a 3-server cluster with one down tolerates zero.
2. Storage health (the usual root cause)
- WAL fsync latency (WAL: write-ahead log, the append-only file etcd writes every change to before applying it, so a crash can replay from it; fsync forces that write out to physical disk instead of leaving it buffered) (
etcd_disk_wal_fsync_duration_seconds) and backend commit latency (how long etcd's underlying storage engine takes to durably commit a batch of writes) (etcd_disk_backend_commit_duration_seconds). etcd's own guidance is that the p99 (99th percentile) should stay under 10 ms and 25 ms respectively. When fsync is slow, heartbeats miss their deadline and followers start elections, so a noisy-neighbour disk (a shared disk whose latency spikes because something else on the same physical hardware is hammering it) shows up first as "leader changes", not as "disk slow". Alert on the cause, not just the symptom. - Database size vs quota. etcd has a space quota (2 GB by default). When it is exceeded, etcd raises a NOSPACE alarm and rejects writes with
mvcc: database space exceeded(mvcc, multi-version concurrency control, is etcd's storage engine; the message means the quota is full) until someone compacts (discards old key revisions the quota no longer needs), defragments (reclaims the disk space compaction just freed, since the storage file does not shrink on its own) and disarms the alarm (clears the NOSPACE flag so writes resume). That is a full outage for registrations and config changes, and it is completely predictable from a growth trend. Tracketcd_mvcc_db_total_size_in_bytesagainst the quota. - For Consul, the equivalent is
consul.raft.commitTimerising, plus disk usage on the servers' data directory.
3. Consumer-side health (what users actually feel)
The registry can be "green" while clients are broken (a bad ACL token, meaning the credential a client presents on every request has been revoked or misconfigured against the access control list, or ACL, that Consul or etcd checks it against; a firewall change; a client library that stopped re-watching). So measure from the outside:
- Synthetic lookups: a prober in each zone resolves a few canary service names every 10 seconds through the same path real clients use (DNS interface, HTTP API, or client library) and records success and latency.
- Watch staleness: the prober writes a heartbeat key (for example a timestamp) every 10 seconds, and consumers or a second prober report how old the value they see is. This catches a stuck watch or a partitioned follower serving stale reads.
- Resolution errors in client libraries: expose a counter from the shared discovery library for "lookup failed" and "served from stale cache because the registry was unreachable".
- Catalog churn: registration and deregistration rate. A sudden spike of deregistrations (for example half the instances of one service flapping) is a health-check or network problem that will soon be a traffic problem.
Page vs log: the rule and the table
The rule: page when the condition (a) already breaches the consumer SLO, or (b) removes your last margin before an outage, or (c) will cause an outage on a predictable timeline shorter than a working day. Everything else is a ticket or a log.
| Condition | Action | Why |
|---|---|---|
| No leader for more than 30 s (sustained, not a blip) | Page | All writes are failing now; deploys and failovers cannot register. |
| Failure tolerance = 0 (for example 3-member cluster with 1 member down) | Page | The next failure is total loss of writes. |
| Synthetic lookup success below SLO, burning error budget fast | Page | Consumers are impacted now. |
| DB size above 80% of quota and projected to hit 100% within 24 h | Page | Guaranteed write outage with a known ETA. |
| Watch staleness above 60 s in any zone | Page | Instances are routing on stale endpoints or config. |
| Single leader change | Log / dashboard | Normal during member restarts and upgrades. |
| More than 3 leader changes in 15 min | Ticket (page if combined with failed writes) | Flapping leadership is usually disk or network trouble building up. |
| p99 fsync above 10 ms for 15 min | Ticket, high priority | Leading indicator of elections; fix before it becomes one. |
| One member down in a 5-member cluster (tolerance still 1) | Ticket | Margin reduced but not exhausted. |
| DB size above 60% of quota | Ticket | Schedule compaction/defrag or find the writer that grew it. |
Worked example
A 3-member etcd cluster backs discovery for 40 services. Tuesday 09:00, one member's disk starts showing p99 WAL fsync of 40 ms (four times etcd's 10 ms guidance). What fires:
- Tuesday 09:15: the fsync alert raises a ticket after 15 minutes of sustained breach. Nobody is paged yet: the cluster still has a leader and quorum.
- Tuesday through Wednesday: that member keeps missing heartbeats intermittently;
etcd_server_leader_changes_seen_totalclimbs by a handful of elections a day. Each burst is still a ticket on its own, but the dashboard now correlates it with the open fsync ticket. - Wednesday 09:00 (about 24 hours after the ticket opened): the operator, following the runbook, takes the sick member out of rotation for replacement. The cluster is now 2 of 3: failure tolerance drops to 0 and the page fires, because one more loss means
floor(3/2)+1 = 2members can no longer be met.
Notice that the page came from the margin signal, not from the disk metric. The disk metric's ticket, open since Tuesday 09:15, gave about a day of warning before the operator's Wednesday action tripped the margin signal that actually woke someone.
Trade-offs and pitfalls
- Paging on leader changes is the classic noisy alert. Rolling upgrades cause them by design. Page on "no leader for N seconds" or on elections combined with failed writes.
- Monitoring only the servers misses client failures. The prober and client-library metrics are what tell you consumers are affected.
- Staleness is invisible without a heartbeat key. A follower cut off from the leader can keep serving reads it believes are current (unless clients ask for linearizable, meaning always-up-to-date, reads).
- Alert on trends for capacity. A DB-size alert at 95% gives minutes of warning; a projection alert gives hours.
- Multi-team blast radius. Because several teams depend on the registry, tie the paging alerts to a documented runbook (how to disarm a NOSPACE alarm, how to replace a member without losing quorum) so the on-call engineer is not learning Raft at 3 a.m.
An API gateway routes external traffic to internal microservices that use a registry for discovery. Describe how the gateway should discover internal services, handle caching of endpoints, and react to rapid scaling events (for example autoscale from 1 to 100 instances). Discuss cache TTLs, warm-up strategies, and circuit breaker interactions.
Sample Answer
Direct answer
The gateway should subscribe to the registry (a watch or streaming update, not a periodic re-query) and hold the endpoint list for each service in memory, so a new instance becomes routable within about a second of passing its health check. A short cache TTL (time-to-live) is then only the fallback for when the subscription breaks, and on registry failure the gateway keeps serving the last known-good list rather than dropping it. For a 1-to-100 scale-out, the danger is not discovery speed but three interactions: new instances receiving full traffic before they are warm, the one old instance being crushed by slow-start weighting, and circuit breakers sized for the old fleet tripping on healthy traffic.
Terms used below
- Registry: the service holding the current list of instances (Consul, Eureka, etcd, the Kubernetes API).
- Endpoint cache: the gateway's in-memory copy of that list.
- Warm-up: the period after start when an instance is slow (just-in-time compilation, empty caches, cold connection pools).
- Circuit breaker: a limit that fails requests fast instead of queueing them once a threshold is hit. In the Envoy proxy (used by many gateways) this means cluster-level limits on connections, pending requests and active requests, plus outlier detection, which temporarily removes (ejects) an individual instance that keeps failing. Here "cluster" is Envoy's word for the group of instances behind one backend service (an "upstream cluster" or "upstream group"), not a Kubernetes cluster or a database cluster.
1. How the gateway discovers services
flowchart LR
R[Registry] -- watch stream --> CP[Gateway control plane]
CP -- endpoint updates --> G1[Gateway node 1]
CP -- endpoint updates --> G2[Gateway node 2]
G1 --> S[Service instances]
G2 --> S
- Push over poll. One component (a control plane, meaning a management service that holds routing state and pushes it to the proxies, or each gateway node itself) holds a watch on the registry and pushes changes to gateway nodes. Polling costs load that grows with nodes × services: 20 gateway nodes polling 200 services every 30 s is 20 × 200 / 30 ≈ 133 registry requests per second, forever, and still leaves instances up to 30 s late.
- Only ready instances. The gateway must receive only instances that have passed a readiness check, not ones that have merely started.
- Registry outage = keep last good list. If the watch fails, keep routing with the cached list and let the gateway's own passive checks (errors on real traffic) drop dead instances. Emptying the cache because the registry is unreachable turns a registry outage into a full outage.
2. Caching of endpoints and TTLs
With push-based updates, a TTL is not what makes things fresh; it is the safety net. My settings:
- Refresh on push, immediately.
- Full resync every few minutes, in case a push was missed.
- Maximum staleness before alerting (not before discarding): if the list has not been confirmed for, say, 5 minutes, page someone but keep serving.
If you are stuck with polling (DNS-based discovery, for example), the TTL becomes the key number and has to be short, because it bounds two different delays:
| Direction | What a TTL of T means |
|---|---|
| Scale-out | New instances wait up to T to receive traffic (wasted capacity during a spike) |
| Scale-in | Terminated instances keep receiving requests for up to T (errors) |
The scale-in side is the one that hurts. Fix it on the service side, not with a tiny TTL: an instance being removed first deregisters, then keeps serving for a drain delay at least as long as the maximum staleness (T plus propagation time), then exits.
3. Rapid scaling from 1 to 100 instances
Warm-up
Mark an instance ready only after it has warmed: run a few synthetic requests through the hot paths (the request types that carry most of the traffic, so their caches and connection pools are what actually need priming) and prime local caches in the readiness check. That protects new instances regardless of how the gateway balances.
The slow-start trap in a 1-to-100 event
Envoy's slow start gives a new endpoint a reduced weight that grows over a configured window, starting at a minimum (10% of normal weight by default). With 1 old instance at full weight and 99 new ones at 10%:
old instance’s share=1+99×0.11=10.91≈9.2%instead of the fair 1%. At 60,000 requests per second (RPS) arriving across the fleet, that is about 5,505 RPS on the one old instance against a fair share of 600. Slow start protects the 99 newcomers by overloading the one instance that was already struggling, which is presumably why you scaled. Envoy's own documentation notes slow start is meant for a few new endpoints joining an established set. So for a large scale-out, rely on readiness-gated warm-up and turn slow start off or keep its window short; use slow start for the ordinary case of a handful of instances joining.
Circuit breakers sized for yesterday's fleet
Envoy's cluster circuit-breaker thresholds apply to the whole upstream cluster, not per instance, and default to 1024 for maximum connections, pending requests and active requests. By Little's law (concurrent requests = arrival rate × time in system), 4 gateway nodes carrying 60,000 RPS at 80 ms per request each hold
460000×0.08=1200 concurrent requestswhich exceeds 1024. The gateway starts rejecting requests with 503 errors even though 100 healthy instances sit behind it. Size these thresholds from the peak you will scale to, not the size you started at, or make them scale with the instance count.
Outlier detection during warm-up
Cold instances are often slow or return errors in their first seconds. Outlier detection may eject them (by default after 5 consecutive 5xx errors, meaning HTTP server-error responses in the 500-599 range, for 30 s the first time), but Envoy caps ejections at 10% of the cluster by default. That cap is protective in both directions: in a 100-instance cluster at most 10 are ejected, and in a small cluster it stops detection from removing the capacity you have. Do not raise it just to "clean up" a bad scale-out; fix warm-up instead.
Retries amplifying overload
When the old instance is overloaded, gateway retries add load exactly where it hurts. Use a retry budget (retries allowed only up to a fraction of active requests; Envoy's default budget is 20%) rather than a fixed retry count per request.
Worked example: timeline of a scale-out
Assume a flash sale drives traffic from 500 to 60,000 RPS.
- T+0: autoscaler requests 99 new instances.
- Instances start over the next minute or two, warm up behind their readiness check, and register.
- Each registration reaches the gateway within about a second via the watch, rather than up to 30 s later with a 30 s poll.
- With slow start off and warm-up done before readiness, traffic spreads evenly as each instance appears, and the gateway's circuit-breaker limits (sized to the 1,200-concurrent-request peak with headroom) do not trip.
- On scale-in, each instance deregisters, drains for longer than the maximum staleness, then exits. To see why the drain matters, consider the counterfactual without it: with a poll-based cache whose TTL equals the maximum-staleness window, gateway nodes refresh their caches at staggered times spread evenly across that window, so at the instant an instance actually stops, roughly half of them have already refreshed past it and half have not. With independent random picks, a request plus one retry would then hit that still-stale half both times 0.5 × 0.5 = 25% of the time. Drain-then-exit (keeping the instance serving until the whole staleness window has passed) removes this risk entirely, because no cache can still be pointing at a dead instance once it actually exits.
Trade-offs and pitfalls
- Push is more machinery than poll: a watch has to handle disconnects and resync. Worth it at this scale; for a small, stable estate, polling every few seconds is fine.
- Registry-driven routing vs a load balancer per service: pointing the gateway at a per-service load balancer (or a Kubernetes Service virtual IP) hands the scaling problem to the platform, but loses per-request balancing and outlier detection at the gateway.
- Health checks that are too eager remove warming instances and cause flapping (instances bouncing in and out of the list). Separate "ready to receive traffic" from "alive".
- Do not make the gateway's TTL the only defence against dead instances. Pair it with passive health checking and deregister-then-drain on the service side.
Clients across your fleet cache service endpoints, and after deploys or scale-out events you start seeing a spike of connection errors from clients still talking to instances that no longer exist. How would you detect that this is happening in production, and what would you change about the caching and invalidation path to reduce it? Walk through the trade-offs of the approaches you'd consider.
Sample Answer
Direct answer
First prove the pattern: correlate client connection errors by destination IP with deploy and scale events, and check whether the failing IPs were still in the clients' caches but already gone from the registry. Then fix it in the order of cost versus benefit: (1) make servers leave gracefully (stop being advertised, keep serving while callers catch up, then exit), which removes most of the errors for almost no cost; (2) move clients from TTL-based caches to watch-based (push) updates; (3) add passive ejection plus a retry on a different endpoint for idempotent requests as a safety net. Shortening TTLs alone is the tempting fix and the weakest one.
Terms
- Endpoint cache: the list of IP:port addresses a client keeps for a service so it does not look them up on every request.
- TTL (time-to-live): how long a cached entry is trusted before being refreshed.
- Watch / push: the registry notifies the client when the list changes, instead of the client waiting for its TTL to expire.
- Passive ejection (outlier detection): the client stops sending to an endpoint after it fails a few times in a row, without waiting for the registry.
- Idempotent request: one that is safe to repeat (a read, or a write carrying an idempotency key). Only these can be retried blindly.
- Connection pool: long-lived open connections a client reuses. Pools are a second, often forgotten cache of endpoints.
Step 1: detect it in production
- Error shape. Split client errors by type: connection refused and connection reset (the host exists but nothing listens, typical of a pod that just exited) and connect timeouts (the IP no longer routes anywhere). Stale endpoints produce a spike of these, not of HTTP 500s.
- Correlate with change events. Overlay the error rate with deploy start/end and autoscaler scale-in events (the autoscaler removing instances because load dropped). A stale-cache problem shows spikes that start at each scale-in or rollout step and decay over roughly one cache TTL.
- Tag errors with the destination IP and join against the registry's history: "was this IP deregistered before the failed request was sent?" If yes, the client acted on stale data. Report that as its own metric (
requests_to_deregistered_endpoint). - Measure the cache age. Have the discovery client library export "age of the endpoint list at time of use" and "time from registry change to client update". This tells you whether the problem is propagation delay (clients learn late) or shutdown ordering (the server died before anyone could have learned).
- Look for hidden caches. Language runtimes cache DNS themselves. The JVM's
networkaddress.cache.ttl, for instance, is "cache forever" when a security manager (a now-deprecated JVM feature that sandboxes what code is allowed to do) is installed and implementation-specific otherwise, so check the value your services actually run with. HTTP/2 and gRPC keep connections open for a long time and only re-resolve when a connection breaks.
Step 2: fix the server side first
Most of these errors come from ordering: the instance stops before callers know it is going. In Kubernetes, removing a pod from endpoints and sending the container SIGTERM (the standard Unix signal Kubernetes sends to ask a process to shut itself down) happen in parallel, so a process that exits immediately on SIGTERM is gone while callers are still routing to it.
The graceful sequence:
- Stop advertising: fail readiness (make the readiness probe, the periodic check Kubernetes or the registry uses to decide whether a pod should receive traffic, start reporting not-ready) / deregister from the registry.
- Keep serving for a drain window at least as long as the time callers take to notice (a
preStophook, code Kubernetes runs before sending SIGTERM, doing a sleep of, for example, 15 s, or the application ignoring SIGTERM for that window). - Stop accepting new connections, finish in-flight requests, then exit.
- For long-lived HTTP/2 or gRPC connections, send a GOAWAY (a signal telling the client to open new connections elsewhere) or set a maximum connection age on the server so clients rebalance.
On scale-out the mirror problem exists: register only after the instance is actually ready (readiness probe passes), or new instances receive traffic they cannot serve.
Step 3: change the client caching and invalidation path
| Approach | What it fixes | Cost / risk |
|---|---|---|
| Shorter TTL (30 s to 5 s) | Shrinks the stale window roughly in proportion | Multiplies lookup load on the registry or DNS by the same factor; does nothing about connection pools; still a window |
| Watch-based push (blocking queries, xDS streams, Kubernetes EndpointSlice watches) | Clients update within about a second of the registry change | Long-lived connections to the registry; the registry's fan-out load grows with watchers; needs a fallback poll if the stream breaks |
| Passive ejection | Stops sending to a dead endpoint after a few failures, regardless of what the registry says | Those first few requests still fail; per-client view |
| Retry on a different endpoint (idempotent only) | Converts a stale-endpoint failure into a small latency cost | Retry storms if unbounded; use a retry budget (for example at most 10% extra requests) |
| Evict on connection error | A refused connection immediately removes that endpoint from the local cache until the next refresh confirms it | Can briefly remove a healthy endpoint on a transient network blip |
| Stale-on-registry-outage | Keeps traffic flowing when the registry is unreachable | Opposite trade-off: deliberately uses old data, so combine with passive ejection |
Worked example: how much does each fix buy?
A service has 50 pods. A rollout replaces one pod every 20 seconds. Clients use a 30-second TTL cache, and pods exit immediately on SIGTERM.
- Each dead pod's entry stays cached for up to 30 s (the TTL), but Little's law needs the average, not the maximum. If clients' cache refreshes are staggered evenly across that window, a given dead pod's entry survives in a given cache for 15 s on average, half the TTL. A new pod dies every 20 s, which is an arrival rate of 1/20 = 0.05 dead pods per second. By Little's law (average number in the system = arrival rate × average time in the system), the average number of dead-but-still-cached pods at any moment is 0.05 × 15 = 0.75.
- Requests are spread evenly across the 50 pods, so at any moment about 0.75 / 50 = 1.5% of requests during the rollout land on a dead pod and fail.
Now apply the fixes:
- Shorter TTL (5 s): average residency drops to 2.5 s, so 0.05 × 2.5 = 0.125 dead-cached pods, 0.125 / 50 = 0.25% errors, at 6 times the lookup load.
- Graceful drain with a 30 s preStop window: the pod keeps serving for the whole time any client might still have it cached, so stale-cache errors go to about zero, with no extra registry load. This is why the server-side fix comes first.
- Retry on another endpoint: whatever remains is absorbed as one extra attempt per affected idempotent request.
Recommendation and trade-offs
Do all three layers, in this order: graceful shutdown and correct readiness (cheapest, largest effect), watch-based updates for the discovery library with a periodic full refresh as a backstop, and passive ejection with budgeted retries. Keep TTLs moderate rather than tiny: once push updates and draining exist, the TTL is only the fallback bound, and very short TTLs mainly add load on the registry and DNS. The trade-off to watch is fan-out: with thousands of clients watching a popular service, every membership change becomes thousands of notifications, so the registry or its distribution layer must be sized for watch fan-out, not just for writes.
Explain DNS-based service discovery and how DNS caching and TTLs affect failover and load distribution. If you had to choose between a 60-second and a 300-second TTL for a service record, which would you pick and why, and what else would you need to account for to make that choice actually hold up in practice?
Sample Answer
Direct answer
DNS-based service discovery means clients find a service by resolving a name such as orders.internal.example.com to the IP addresses of its healthy instances. Because every resolver and client caches the answer for the record's TTL (time-to-live), the TTL sets how long clients may keep sending traffic to an address you have already removed: it is a floor on failover time and a limit on how quickly load shifts. For a service record that must fail over, I would pick 60 seconds. The extra query load is small and a 300-second TTL means up to five minutes of errors after an instance or region dies. But a 60-second TTL only holds if clients actually honour it, so the rest of the job is making sure they do.
How DNS-based discovery works
- Instances register (directly, or via a health-checking DNS provider such as Consul's DNS interface, Kubernetes' CoreDNS, or a cloud DNS service with health checks).
- A client asks its local stub resolver (the small resolver library in the OS or language runtime), which asks a recursive resolver (the caching server, e.g. the node-local cache or the VPC (Virtual Private Cloud) resolver), which asks the authoritative server (the one that actually owns the zone, meaning the portion of the DNS namespace, such as
internal.example.com, that this server holds the real records for). - The answer, say three A records (address records, each mapping the name to one IPv4 address), is cached at every layer for the TTL. The client picks one address, usually the first, and connects.
Nothing tells a client that an address went bad: it only learns on its next lookup after the cached copy expires.
How caching and TTLs affect failover
Worst-case time for all clients to stop using a dead address is roughly:
failover time ≈ health-check detection time + TTL (+ any extra caching the client adds)
With a health check that marks an instance down after 3 failures at 10-second intervals (30 seconds):
| TTL | Worst-case failover | Lookups per second from 2,000 caching client hosts (one lookup per TTL each) |
|---|---|---|
| 60 s | 30 + 60 = 90 s | 2,000 / 60 ≈ 33 |
| 300 s | 30 + 300 = 330 s | 2,000 / 300 ≈ 6.7 |
The 60-second TTL costs about 26 more queries per second, which any resolver absorbs trivially, and cuts the worst-case failover from 5.5 minutes to 1.5 minutes. Query volume scales linearly with 1/TTL, so going lower (5 seconds) is 400 queries per second for the same fleet and starts to matter for resolvers and for paid public DNS.
How caching affects load distribution
DNS balances resolvers, not requests. If one large recursive resolver (a shared NAT, meaning a network-address-translation gateway that many separate devices sit behind and also route their DNS queries through, or a corporate DNS server) serves 500 clients, all 500 may receive the same cached answer in the same order and pile onto the first IP. Longer TTLs make this worse, because a skewed answer persists longer; round-robin record ordering, returning a random subset per response, or client-side random selection among all returned addresses spreads it out. Adding capacity has the same lag as removing it: new instances receive traffic only as caches expire.
What else must be true for the 60-second choice to hold
- Clients must honour the TTL. Language runtimes add their own caches. The JVM's positive cache (a positive cache remembers successful lookups; the negative cache below remembers failures) is governed by the
networkaddress.cache.ttlsecurity property, whose default is implementation-specific; set it explicitly (for example to 60) instead of trusting it. Some resolvers also enforce a minimum TTL that silently raises yours. - Clients must re-resolve at all. A connection pool or an HTTP/2 keep-alive connection (reusing one already-open connection for many requests instead of reconnecting for each one) opened to an IP will keep using that IP for hours, regardless of DNS TTL. Cap connection lifetime (e.g. recycle connections every few minutes) and re-resolve on reconnect, or failover never happens for long-lived clients.
- DNS must only return healthy addresses. The TTL is pointless if the authoritative server keeps handing out the dead IP; the records must be driven by health checks.
- Negative caching. If a lookup returns "no such name" (for example during a brief deregister-everything incident), resolvers cache that negative answer for a time taken from the zone's SOA (start of authority) record, per RFC 2308, and Java's
networkaddress.cache.negative.ttldefaults to 10 seconds. A long negative TTL turns a 5-second blip into minutes of "host not found". Keep it short for service zones. - Clients must retry on another address. Even with perfect TTLs, requests fail during the detection window. Returning several records and having clients try the next address on connection failure turns an outage into a latency bump.
- Lower TTLs before planned changes. A TTL only shortens after the old one expires, so for a planned migration drop the TTL a full old-TTL period ahead of time.
When I would pick 300 seconds instead
For records that rarely change and front something that fails over by other means (a load balancer's stable virtual IP, an address the load balancer itself keeps fixed while swapping which physical instance answers behind it, or an anycast address, the same IP announced from multiple physical locations so network routing sends each client to the nearest live one and simply stops routing to a failed location, no DNS change needed), 300 seconds or more is fine: the address itself never dies, so failover does not depend on DNS, and a longer TTL cuts query volume and shields you from brief DNS outages.
Trade-offs and pitfalls
- A very low TTL (under 10 seconds) makes DNS itself a hard dependency: if the resolver hiccups, every client fails its lookups at once.
- DNS offers no per-request load awareness and no health-based retries; for east-west traffic (traffic between your own services inside the system, as opposed to north-south traffic coming in from outside clients) between microservices that need fast failover, client-side load balancing (the calling service itself holds the instance list and picks one per request, no DNS or proxy hop involved) or a service mesh (a proxy next to each service that receives endpoint updates by streaming) reacts in seconds instead of a TTL period.
- Assuming "TTL 60 means failover in 60 seconds" ignores detection time, runtime caches and pooled connections, which is where most real incidents come from.
Unlock Full Question Bank
Get access to all 19 Service Discovery and Configuration Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.