Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
Describe how a service mesh such as Istio or Linkerd integrates with service discovery and dynamic configuration. Explain how sidecars discover backends, how the control plane distributes routing/policy config, and trade-offs around performance, operational complexity, and discoverability when adding a mesh to an existing platform.
Sample Answer
Direct answer
A service mesh moves discovery and traffic configuration out of application code and into proxies. A control plane watches the platform's own source of truth (in Kubernetes: Services, EndpointSlices and pod labels, plus the mesh's routing resources) and translates it into configuration that it pushes to a proxy next to every workload. The application calls a plain hostname; the local proxy already knows the healthy backend IPs, the routing rules, retry and timeout policy, and the certificates for mTLS. You gain uniform routing, security and telemetry for every language; you pay in per-pod CPU and memory, an extra network hop per side, a control plane whose config push can itself become a scaling problem, and a system that is harder to debug because the real routing decision now lives in proxy state you cannot see in your code.
Terms
- Service mesh: a layer of proxies that handle service-to-service traffic, plus a control plane that configures them. Istio and Linkerd are the two most common on Kubernetes.
- Sidecar: a proxy container injected into each application pod. Traffic in and out of the pod is redirected through it (typically using iptables, the Linux kernel's packet-filtering and routing rules, configured at pod start to send traffic through the proxy first).
- Data plane vs control plane: the proxies that carry requests are the data plane; the component that tells them what to do is the control plane (
istiodin Istio; thedestination,identityand proxy-injector components in Linkerd). - mTLS (mutual TLS): both sides of a connection present certificates, so the server knows which workload is calling, not just that the traffic is encrypted.
- xDS: the family of discovery APIs Envoy proxies use to receive configuration. Envoy is the specific proxy Istio injects as the sidecar; xDS covers LDS (listeners), RDS (routes), CDS (clusters, meaning upstream services) and EDS (endpoints, meaning the IPs behind a cluster).
- Kubernetes Service, EndpointSlice, pod labels: the platform's own primitives the mesh builds on. A Service is a stable name for a group of pods; an EndpointSlice is the live, auto-updated list of that group's actual pod IPs and their readiness; pod labels are the key/value tags (
app: reviews) Services and mesh rules select pods by. - gRPC: a binary remote-procedure-call protocol; the control plane uses it as the transport for pushing config and endpoints to every sidecar.
How sidecars discover backends
flowchart LR
K8s[Kubernetes API: Services, EndpointSlices, mesh custom resources] -->|watch| CP[Control plane]
CP -->|push config and endpoints over gRPC| PA[Sidecar in pod A]
CP -->|push| PB[Sidecar in pod B]
AppA[App A] -->|calls reviews:9080| PA
PA -->|mTLS, load balanced| PB
PB --> AppB[App B, a reviews pod]
Istio (Envoy sidecars):
istiodwatches Services and EndpointSlices, so it knows every pod IP behind every Service and which are ready.- It turns that into xDS: a cluster for each service, endpoints (EDS) listing healthy pod IPs, and routes for any VirtualService rules.
- Each Envoy sidecar holds a long-lived gRPC stream to
istiod. When a pod becomes ready or is deleted,istiodpushes an endpoint update and the sidecar's load balancer changes within seconds, without DNS TTLs and without the app doing anything. - The app still resolves
reviewsthrough normal cluster DNS, but the sidecar intercepts the connection and picks the actual backend from its own endpoint list.
Linkerd (its own Rust micro-proxy, linkerd2-proxy):
The proxy asks the control plane's destination service, over gRPC, for the endpoints and policy of each destination the application actually talks to, and keeps a watch open for that destination. That on-demand model means a proxy only holds state for the services its workload uses, which is a real difference from Istio's default behaviour.
How the control plane distributes routing and policy config
The two resources to actually know first are VirtualService and DestinationRule: together they cover almost every day-to-day routing decision. The rest of this list matters once you also need identity-based security policy or aren't yet on Istio's newer routing APIs.
- Routing: in Istio, VirtualService (match rules, weighted splits such as "5% to v2", retries, timeouts, fault injection: deliberately injecting errors or delay into some requests to test how callers handle it) and DestinationRule (load-balancing algorithm, connection pools, outlier detection which ejects endpoints that keep failing, and subsets such as
v1/v2by label). In both meshes, the Kubernetes Gateway API's HTTPRoute can express traffic splits too. - Security policy: PeerAuthentication (require mTLS) and AuthorizationPolicy (which identities may call which paths) in Istio; Server and AuthorizationPolicy resources in Linkerd. Workload certificates are issued and rotated automatically by the control plane.
- Distribution: operators apply these as Kubernetes resources; the control plane validates and translates them and pushes the result. This is dynamic configuration delivery: a traffic shift from 5% to 50% is an API write, effective fleet-wide in seconds, with no redeploy.
Worked example: why config scope matters at scale
A cluster has 500 services and 5,000 pods, and each workload really calls about 10 services with about 10 endpoints each.
- Default Istio behaviour sends each sidecar configuration for the whole mesh. Endpoint entries held across the fleet: 5,000 sidecars × 5,000 endpoints = 25,000,000. Every endpoint change also triggers a push to every sidecar.
- Scoped (Istio's
Sidecarresource restricting each workload to the hosts it uses, ordiscoverySelectorslimiting which namespaces are visible): 5,000 × (10 × 10) = 500,000 entries, 50 times less proxy memory and push work. - Linkerd's on-demand lookups give roughly the scoped behaviour by default.
The lesson: with a push-everything control plane, proxy memory and control-plane CPU grow roughly with (number of proxies) × (size of the mesh). Both of those factors grow with the same underlying number, how many pods are in the cluster, so doubling the fleet does not double the cost, it roughly quadruples it: that is what "quadratic as the platform grows" means here. Scoping is not an optimisation to add later; it belongs in the first rollout.
Trade-offs when adding a mesh to an existing platform
| Dimension | Cost | Mitigation |
|---|---|---|
| Performance | Two extra proxy hops per call (client sidecar and server sidecar), plus TLS handshakes; per-pod CPU and memory for every sidecar | Measure p99 (99th percentile) latency on your own traffic before and after on a pilot namespace; do not trust vendor benchmarks. Consider Istio's ambient mode (generally available since Istio 1.24), which replaces per-pod sidecars with a per-node ztunnel (a shared node-level proxy handling basic encrypted routing for every pod on that node) for L4 (TCP-level mTLS and policy) and optional per-namespace waypoint proxies (added only where a namespace needs HTTP-level routing decisions) only where L7 (HTTP-level routing) is needed |
| Operational complexity | A new critical control plane to upgrade and monitor; injection webhooks (the Kubernetes mechanism that automatically adds the sidecar container to a pod as it is created, so nobody edits every deployment by hand); sidecar start-up ordering (the app starts before its proxy is ready, or the proxy exits before the app finishes draining, meaning finishing in-flight requests and closing connections cleanly) | Pilot one namespace; use native sidecar support (a Kubernetes feature that starts and stops sidecar containers in the right order automatically, instead of relying on timing tricks) in recent Kubernetes versions for start/stop ordering; keep the mesh version upgrade as a rehearsed runbook |
| Discoverability (debugging) | The routing decision is no longer in code or DNS; a request can fail because of a DestinationRule nobody on the app team knows about | Standard tooling: istioctl proxy-config endpoints <pod> and istioctl proxy-status (is this sidecar in sync with istiod?), linkerd viz for per-route metrics; mesh resources owned in the same repo as the service |
| Two discovery systems | During migration, meshed pods use proxy-pushed endpoints while unmeshed callers still use DNS or a client-side library; their views of "healthy" can differ | Migrate by call path (callee first, then callers) and keep outlier detection consistent with existing health checks |
| Control-plane failure | If istiod is down, sidecars keep their last config and traffic keeps flowing, but no new endpoints or policy arrive, so scaling events go unseen | Run the control plane highly available; alert on proxy config staleness |
Recommendation
For an existing platform, adopt a mesh when you need at least two of: uniform mTLS with identity-based policy, traffic shifting for progressive delivery (rolling a change out to a growing slice of traffic while watching for regressions, rather than to everyone at once), and consistent retries, timeouts and telemetry across several languages. If you only need discovery and load balancing, Kubernetes Services plus a good client library are cheaper. If you do adopt: start with one namespace, scope config from day one, and prefer the lighter footprint (Linkerd, or Istio ambient) unless you specifically need Envoy's L7 feature set everywhere.
Compare push-based and pull-based configuration distribution models (push from central control plane vs clients polling or watching). Discuss pros and cons for scalability, consistency, security, and offline clients. Give example architectures for both and situations where each is preferable (e.g., low-latency updates vs intermittent connectivity).
Sample Answer
Direct answer
In a push model the central control plane decides when to send config and delivers it to each client; in a pull model each client asks for config, either by polling on a timer or by holding open a watch (a long-lived request the server answers when something changes). Push gives the fastest updates and central visibility of who has what, but the control plane must track and reach every client. Pull scales and handles offline or firewalled clients naturally, at the cost of update latency equal to the poll interval. My default for most fleets is the hybrid most modern systems use: the client opens the connection (pull), and the server streams changes down it (push), with a local cache of last known good config.
The two models
- Push: the control plane holds a list of targets and sends each one the new config. Examples: running an Ansible playbook (Ansible: a tool that runs a scripted list of setup steps against many hosts at once) over SSH (secure shell) to every host; a deployment system writing config to each instance's API.
- Pull by polling: each client fetches config every T seconds, usually with a conditional request (an ETag, a fingerprint of the current version, so the server can answer "not modified" cheaply). Examples: the Puppet agent (a program installed on each host that periodically fetches and applies its assigned configuration) checking in on an interval; a sidecar polling an HTTP config endpoint.
- Pull by watch (the hybrid): the client connects out and subscribes; the server pushes each change as it happens. Examples: Envoy proxies (Envoy is a widely used network proxy that runs beside a service instance) receiving configuration from a control plane over the streaming xDS APIs (Envoy's discovery service protocol); etcd watches (etcd's own long-held subscription that streams each change to the client as it happens); Consul blocking queries (a request that the server holds open until the data changes or a timeout).
Comparison
| Dimension | Push | Pull (polling) | Pull (watch / stream) |
|---|---|---|---|
| Update latency | Seconds | Up to one poll interval | Seconds |
| Scalability | Control plane does the fan-out (sending one change to many clients) and must retry every failed target | Load is steady and spread, but idle polls cost: 50,000 clients at 30 s = 50,000 / 30 ≈ 1,667 requests per second even when nothing changes | One long-lived connection per client; server memory scales with connections, change delivery is cheap |
| Consistency | Central knowledge of who got what; partial failures leave a mixed fleet the pusher must track | Every client converges within one interval; fleet is briefly mixed | Converges fast; needs resume logic for missed events on reconnect |
| Security | Control plane needs network access and credentials to every client: compromise it and you own the fleet; needs inbound ports open on clients | Clients need only outbound access and credentials to read their own config; server can scope what each client sees | Same as polling |
| Offline or intermittent clients | Push fails; must be queued and retried; the pusher must know the client came back | Natural: client catches up on its next successful poll | Client reconnects and resyncs; needs cached config to run while disconnected |
Worked example
A retailer runs 50,000 point-of-sale terminals in stores with flaky links, and 2,000 back-end service instances in a data centre.
- Store terminals: pull by polling, every 5 minutes, with on-disk cache. Terminals are behind store NAT (network address translation) routers, so the central system cannot open connections to them, and many are offline for hours. Polling load is 50,000 / 300 ≈ 167 requests per second, mostly "not modified" answers. A terminal offline overnight simply fetches the latest version when it reconnects. Update latency of up to 5 minutes is fine for price-table config.
- Back-end services: watch-based streaming. A feature kill switch (a config flag that can turn a feature off instantly, without a new deploy) or a routing change must reach all 2,000 instances within seconds. Each instance's local agent holds one streaming connection to the config service (2,000 connections, easily handled by a few servers) and applies changes as they arrive, falling back to its cached copy if the stream drops.
- Emergency push: for a security revocation the platform also keeps a push path that targets specific instances and tracks acknowledgements, because "most clients will catch up soon" is not good enough there.
When each is preferable
- Push when targets are few, always reachable, and you need a confirmed, orchestrated change (a coordinated rollout where the controller must know that host 17 got version 42 before touching host 18).
- Pull by polling for large, intermittently connected, or untrusted fleets (edge devices, stores, customer-installed agents), and where seconds of latency do not matter.
- Watch / stream for low-latency updates inside a data centre or cloud (service mesh routes, feature flags, discovery endpoints).
Trade-offs and pitfalls
- Thundering herd with polling (a large batch of clients all hitting the server in the same instant, the way a herd startles and moves together): if every client polls on the minute, the server sees spikes. Add random jitter to each client's schedule.
- Watches still need a periodic full resync: a missed event or bug can leave a client silently stale forever; a periodic re-list (every few minutes) bounds the damage.
- Pull does not mean the control plane is blind: have clients report the version they are running, so you can see fleet convergence and find stragglers.
- Validate before applying either way: a bad config pushed quickly reaches everyone quickly. Stage rollouts (canary first: release to one host or a small slice of the fleet and watch it before releasing to everyone) regardless of the delivery model.
Compare health check approaches between a registry that uses TTL heartbeats (e.g., Consul TTL checks) and Kubernetes probes (liveness, readiness, startup). For each model describe recommended frequencies, failure thresholds, trade-offs between fast failover and false positives, and how check results should influence registry state and load balancing.
Sample Answer
Direct answer
The two models put responsibility in opposite places. A Consul TTL check (Consul is a widely used service registry: a database of which instances of each service exist and are healthy, that clients query to find them) is push-based: the application itself must call Consul every TTL period to say "still passing", and silence past the TTL marks it critical, which removes it from healthy query results. Kubernetes probes are pull-based: the kubelet (the agent on each node) calls into the container, and the three probes do three different jobs: readiness controls whether the pod (Kubernetes' smallest deployable unit: one or more containers scheduled and networked together) receives traffic, liveness controls whether the container is restarted, and startup holds off the other two until a slow starter is up. The frequency and threshold choices follow from a single trade-off: detect a dead instance fast, without evicting or restarting healthy instances that merely paused.
The two models side by side
| Consul TTL check | Kubernetes probes | |
|---|---|---|
| Who acts | The app pushes a heartbeat (PUT /v1/agent/check/pass/<check_id>) | The kubelet pulls: HTTP GET, TCP connect, a gRPC health-check call, or exec (running a command inside the container and checking its exit code) |
| What failure means | No heartbeat within the TTL: status becomes critical | Consecutive failures reach failureThreshold |
| Initial state | Critical until the first heartbeat (so an unready instance never gets traffic) | Readiness starts not-ready until it succeeds |
| Effect on routing | Health-filtered queries (the ?passing filter on the health API, DNS lookups) stop returning the instance | Readiness failure removes the pod from the Service's endpoints (a Kubernetes Service is the stable name/address in front of a set of pods; its endpoints are the live list of pod IPs currently backing it) |
| Effect on process | None; the registry only reports | Liveness failure kills and restarts the container |
| What it really proves | The app's heartbeat code is running and chose to say "pass" | The kubelet can reach the endpoint and it answers correctly |
A subtle difference: a TTL check can report on things a probe cannot see (the app decides "I am passing only if my queue consumer is keeping up"), but a TTL heartbeat thread can keep running while the request-serving threads are deadlocked, and it says nothing about whether the network path from clients works. A probe exercises the real request path from the node, but only as far as the node.
Recommended frequencies and thresholds
Consul TTL checks
- Heartbeat interval about one third of the TTL, e.g. TTL 30 seconds with a heartbeat every 10 seconds. The app then tolerates two lost heartbeats (a garbage-collection (GC) pause, one dropped request) before going critical. A heartbeat at the same period as the TTL fails on the first delay.
- Note that the consecutive-failure counters Consul offers (
FailuresBeforeCritical,SuccessBeforePassing) apply to active checks such as HTTP and TCP, not to TTL checks, so the TTL-to-interval ratio is your failure threshold. - Use
warnfor degraded-but-serving (so dashboards see it) andfailfor a deliberate eviction, e.g. during graceful shutdown the app calls fail first, then drains. - Set
DeregisterCriticalServiceAfterso dead registrations get cleaned up: the minimum is 1 minute and Consul's reaper (the background process that sweeps up and removes checks that have stayed critical too long) runs every 30 seconds, so removal can take slightly longer than configured. Pick something like 10 minutes, long enough that a brief critical period does not delete the registration and force a re-register.
Kubernetes probes
Defaults are periodSeconds: 10, timeoutSeconds: 1, failureThreshold: 3, successThreshold: 1, initialDelaySeconds: 0.
- Readiness: fast and sensitive, because the action is cheap and reversible. For example period 5 s, failureThreshold 2: removed from traffic within about 10 s. It should check what this pod needs to serve (warm caches loaded, local dependencies up) but not shared downstream dependencies (see pitfalls).
- Liveness: slow and conservative, because the action (restart) is expensive and can cascade. Period 10 s, failureThreshold 3 or more, so about 30 s of consecutive failure, and it should check only "is this process wedged" (an endpoint that returns 200 if the event loop, the single loop many runtimes such as Node.js use to process one task at a time, is still responding), never dependencies.
- Startup: for slow starters, e.g.
failureThreshold: 30,periodSeconds: 10gives up to 30 × 10 = 300 seconds to start; liveness and readiness do not run until it succeeds once, so liveness can stay aggressive without killing a JVM (Java virtual machine) that is still warming up. timeoutSeconds: 1is too tight for many real endpoints under load; 2 to 3 seconds avoids counting a slow response as a failure.
Fast failover versus false positives, with numbers
Detection time is roughly period × failureThreshold for probes and TTL for TTL checks. False positives come from transient pauses shorter than that window.
Take a service whose stop-the-world GC pauses (a garbage-collection pause during which the runtime freezes every application thread, not just the collector, so nothing responds until it finishes) occasionally reach 4 seconds:
- Readiness at period 2 s, threshold 1: a 4-second pause fails at least one probe and the pod is pulled from traffic. With 50 pods pausing independently, pods are constantly flapping in and out of the endpoint list, and each change is an update pushed to every client.
- Readiness at period 5 s, threshold 2: a pod needs about 10 seconds of failure to be removed, so a 4-second pause never evicts it, and a truly dead pod is still gone in about 10 s.
- Liveness with the same aggressive settings would restart pods on every GC pause, which makes the pause problem worse (cold caches, restart storms).
The pattern: be aggressive where the action is reversible (removing from load balancing), conservative where it is destructive (restart, deregistration).
How results should drive registry state and load balancing
- Health state feeds routing, not just dashboards: a failing readiness or a critical TTL check must remove the instance from the endpoint set that clients and proxies consume. Warning states should stay routable but visible.
- Removal is not deletion: an unhealthy instance stays registered (so it can come back when it recovers); deregistration happens only after a long critical period or on shutdown.
- Protect against mass eviction: if a health signal says 90% of instances are unhealthy at once, the more likely explanation is a broken check or a shared dependency, not 90% of processes failing together. Proxies like Envoy (a widely used open-source proxy) have a panic threshold (by default, when fewer than 50% of hosts are healthy, it load-balances across all hosts anyway) precisely so that a bad health signal does not funnel all traffic onto a handful of survivors.
- Combine with passive checks: load balancers should also eject endpoints that return errors on real traffic, a technique called outlier detection: watch each endpoint's live error rate and temporarily pull out whichever ones stand out from their peers, which reacts faster than any periodic check.
- Graceful shutdown order: fail the check or readiness first, wait for the change to propagate to clients (a few seconds), then stop accepting connections.
Trade-offs and pitfalls
- Dependency checks in readiness or TTL logic: if every pod's readiness checks the shared database and the database blips, every pod goes unready at once and the whole service vanishes from discovery, turning a partial degradation into a total outage.
- Liveness equal to readiness: restarting pods because a dependency is slow is the classic restart storm.
- TTL heartbeats from a separate thread can report "passing" while the serving threads are deadlocked; tie the heartbeat to a signal from the real request path.
- Probes prove local health only: a pod can pass every probe while clients cannot reach it, for example a network policy (cluster rules restricting which pods may talk to which, even when both ends are healthy) blocking the path, or a partitioned node. Pair with client-side outlier detection.
Design a service discovery and routing strategy for microservices deployed across multiple Kubernetes clusters in different regions. Consider DNS vs service mesh, global load balancers, cross-cluster health checks, stale registry entries, and latency/consistency trade-offs for routing decisions.
Sample Answer
Direct answer
Route with a strict locality preference: a call goes to an endpoint in the same cluster if one is healthy, then another cluster in the same region, and crosses regions only on failure. For traffic between services (east-west) I would use a service mesh running one control plane per cluster that also discovers the other clusters' endpoints, because it can fail over per request in seconds and eject bad endpoints based on real traffic. For users entering the system (north-south) I would use a global load balancer (anycast: one IP address advertised from every region, so the network itself delivers each user to the nearest one; or geo-DNS with health checks) in front of per-region ingress (the entry point that receives external traffic for that region, before it reaches any cluster). Plain DNS is the fallback for teams without a mesh and for coarse, region-level failover, not for fine-grained cross-cluster routing, because TTL (time-to-live) caching makes it slow to react.
Terms used below
- Service mesh: a sidecar proxy (a small proxy process running beside each service instance, such as Envoy) that handles every outgoing call, plus a control plane that tells all proxies which endpoints exist. Istio is a common example.
- East-west / north-south: traffic between internal services versus traffic entering from users.
- Outlier detection: a proxy watches real responses and temporarily ejects an endpoint that returns consecutive errors, without waiting for a health check.
- Locality: region, then zone, then cluster; "locality-aware" routing prefers the nearest healthy option.
Setup I am designing for
Three regions, two Kubernetes clusters per region (six clusters), tens of microservices, each deployed to every cluster. Requirements: low latency, survive losing a cluster or a region, no cross-region hop in the normal path.
Why locality matters, with an assumed cross-region round trip of 70 ms (a typical order of magnitude between distant regions, not a measurement): a user request that fans through a chain of 5 services pays nothing extra when all hops stay in-region, but about 5 × 70 = 350 ms extra if each hop is routed across regions. Random global load balancing across clusters would put most hops cross-region, since two thirds of endpoints live in other regions.
Architecture
flowchart TB
U[Users] --> GLB[Global LB: anycast or geo-DNS]
GLB --> I1[Region 1 ingress]
GLB --> I2[Region 2 ingress]
subgraph R1[Region 1]
I1 --> A1[Cluster 1a: services and sidecars]
A1 <--> B1[Cluster 1b]
end
subgraph R2[Region 2]
I2 --> A2[Cluster 2a]
A2 <--> B2[Cluster 2b]
end
A1 -. failover only .-> EW[East-west gateway in Region 2]
EW --> A2
Decisions
1. DNS versus service mesh
| DNS-based (e.g. Kubernetes Multi-Cluster Services, or health-checked DNS names per region) | Service mesh (per-cluster control planes sharing endpoint knowledge) | |
|---|---|---|
| Failover speed | Health-check detection plus TTL plus client DNS caches: typically a minute or more | Seconds: proxies get endpoint updates by streaming and retry per request |
| Granularity | Per name; the client picks an IP and keeps it on pooled connections | Per request, per endpoint, with weights and locality |
| Reacts to real errors | No | Yes, outlier detection |
| Operational cost | Low | High: control-plane upgrades, certificates, proxy resource overhead |
(Kubernetes Multi-Cluster Services above is a Kubernetes feature that makes a service's endpoints visible by name across every cluster in the set, the DNS-based counterpart to the mesh's own cross-cluster discovery.)
Choice: mesh for east-west. I would flip to DNS if the team has few services, low change rate, and cannot staff mesh operations; then use one DNS name per service per region, health-checked, and accept minute-scale failover.
2. Global load balancer for entry traffic
Anycast (one IP advertised from every region, so the network delivers users to the nearest one) fails over in seconds without depending on client DNS caches; geo-DNS with 30 to 60 second TTLs is the cheaper alternative. Either one health-checks a deep endpoint per region, and each region must hold headroom for a failed neighbour's traffic.
3. Cross-cluster health checks
- Do not build a central prober that health-checks every pod (Kubernetes' smallest deployable unit, one or more containers running together) in every cluster: it is a scaling bottleneck and a single point of failure, and a partition between the prober and a healthy cluster would falsely evict it.
- Each cluster's kubelet (the agent that runs on every node and starts, stops and health-checks the pods scheduled to it) readiness probes remain the source of truth for its own pods; the local control plane publishes only ready endpoints.
- Across clusters, rely on (a) the health of each cluster's east-west gateway (the ingress point that receives traffic from other clusters) and (b) passive outlier detection on real requests. A remote endpoint that errors is ejected by the caller's proxy within a few failed requests.
4. Stale registry entries
Each cluster's control plane reads the other clusters' API servers (the Kubernetes API server: the per-cluster endpoint that holds and serves all of that cluster's state, including which pods and services exist) to learn their endpoints. When a remote cluster becomes unreachable, its endpoint list freezes, and those entries go stale.
- Keep the last known list rather than deleting it (a network blip should not remove a healthy cluster), but track its age and demote stale localities: they receive traffic only if no fresher locality is healthy.
- Outlier detection ejects stale endpoints that are actually dead on the first few failed calls.
- Guard against mass removal: if an update would drop most endpoints of a service at once, hold the old list and alert. (Envoy's panic threshold, which load-balances across all hosts when fewer than 50% are healthy by default, is a built-in version of this idea.)
- Product-specific gotcha, not a general rule: in Istio specifically, locality failover only engages when outlier detection is configured on the destination; without it, traffic keeps flowing to the preferred locality even when it is failing. Check this setting explicitly if you adopt Istio; other meshes wire the two together differently.
5. Latency versus consistency in routing decisions
Routing data is eventually consistent by design: endpoint lists propagate in seconds and can briefly disagree between clusters. That is acceptable because every call is protected by retries on connection failure and outlier detection. Making routing strongly consistent (a global registry with consensus on every endpoint change) would put cross-region commits into the discovery path and stall discovery during a partition, which is far worse than routing on a view that is a few seconds old.
Trade-offs and pitfalls
- Flat global load balancing across all clusters: simple and wrong; most hops go cross-region, adding latency and egress cost (the fee cloud providers charge for data leaving a region).
- Failover cascades: when region 1 fails over to region 2, region 2 must have spare capacity, or it fails too. Cap cross-region spillover and shed load (deliberately reject or slow down some requests rather than let every node overload) rather than letting it cascade.
- One global mesh control plane: a single failure or bad config push affects every cluster; run one control plane per cluster so a control-plane problem stays local.
- Trust across clusters: cross-cluster calls need mutual TLS (both sides present certificates) with a shared root of trust (a common certificate authority that every cluster's certificates are issued from, so any cluster can verify a certificate it has never seen before), or clusters cannot authenticate each other's workloads.
Define service discovery in the context of microservices. Explain the differences between client-side discovery and server-side (proxy) discovery, including who is responsible for load balancing, how health checks are handled, and examples of real-world tools (DNS, Consul, Kubernetes Service). When would you choose each approach and why?
Sample Answer
Direct answer
Service discovery is how one service finds the current network addresses of another when instances come and go all the time (autoscaling, deploys, crashes). A service registry keeps the list of live instances. In client-side discovery, the calling service asks the registry for the list and picks an instance itself, so the client does the load balancing. In server-side (proxy) discovery, the caller sends the request to one stable address (a load balancer or proxy), and that component looks up the instances and picks one. My default is server-side through the platform (a Kubernetes Service or a load balancer) because it keeps clients simple and language-independent; I switch to client-side when the extra hop or the lack of per-request control actually costs something.
Why it is needed
In a static world you would hard-code payments = 10.0.1.5. In a microservice fleet, instances get new IPs on every deploy and the count changes with load, so addresses must be looked up at runtime, and the lookup must only return instances that are healthy (able to serve).
The two approaches
Client-side discovery
- Each
paymentsinstance registers itself (or is registered by the platform) in the registry. - The
checkoutservice queries the registry (or a local cache of it) forpaymentsand gets a list, such as[10.0.1.5, 10.0.1.9, 10.0.2.4]. checkout's own library picks one (round robin: cycle through the list in order; or least outstanding requests: send to whichever instance currently has the fewest in-flight calls) and connects directly.
Server-side (proxy) discovery
- Instances register the same way.
checkoutcalls one stable name, such aspayments.default.svc.cluster.localor a load-balancer address.- The load balancer or proxy (which watches the registry) chooses an instance and forwards the request.
Who does what
| Client-side | Server-side (proxy) | |
|---|---|---|
| Who load-balances | The calling service's library | The load balancer or proxy |
| Who knows which instances are healthy | Registry health checks, plus the client's own view (it sees its own timeouts and errors immediately) | Registry health checks, plus the proxy's own active checks or error tracking |
| Extra network hop | None | One (unless the proxy runs on the same host as the caller) |
| Logic lives in | Every client, in every language you use | One shared component |
| Changing the algorithm | Redeploy every client | Reconfigure the proxy |
On health checks: the registry usually learns health by probing instances (an HTTP call to /health every few seconds) or by requiring heartbeats. The client or proxy adds a second, faster signal: passive health checking, which stops sending to an instance that just returned several errors in a row, without waiting for the next probe.
Real-world tools
- DNS: a service name resolves to several A records (IP addresses) and the client picks one. That is a simple form of client-side discovery, but with weak health handling and caching delays (clients keep an answer until its TTL, time-to-live, expires, and some ignore TTLs).
- Consul: supports both. Clients can query its HTTP API or its DNS interface (
payments.service.consul, returning only healthy instances) and pick an instance themselves: client-side. Or traffic goes through Consul's service mesh, where a sidecar proxy (a helper process running next to each service) does the choosing: server-side, but on the same host. - Kubernetes Service: server-side. A normal Service gets one stable virtual IP (the ClusterIP);
kube-proxy(the Kubernetes component that runs on every node and keeps its routing rules in sync with the cluster) programs rules, either iptables (the Linux kernel's built-in packet-filtering and NAT table) or IPVS (the Linux kernel's built-in load balancer), that spread connections across the pods (a pod is Kubernetes' smallest deployable unit, one or more containers scheduled together) that are passing their readiness probe (a periodic check, such as an HTTP call, that decides whether a pod is currently able to serve traffic; failing it removes the pod from the Service without killing it). A headless Service (no ClusterIP) instead returns the pod IPs through DNS, which is client-side, and is what gRPC (a high-performance RPC framework built on HTTP/2, which multiplexes many concurrent requests over a single long-lived TCP connection instead of opening a new one per call) clients often use to balance per request. - Cloud load balancers (for example, an AWS Application Load Balancer with a target group): server-side, with the cloud provider registering instances.
- Client-side libraries: gRPC's built-in load-balancing policies, Spring Cloud LoadBalancer (which replaced the Netflix Ribbon library that is no longer developed).
Worked example
checkout runs 200 instances and calls payments, which runs 50.
- Client-side: every
checkoutinstance may keep connections to everypaymentsinstance: 200 × 50 = 10,000 connections. Whenpaymentsinstance 17 starts timing out, eachcheckoutinstance notices within its own next few requests and stops using it, without waiting for the registry. But the balancing code has to exist, and be kept correct, in every languagecheckout-like services are written in. - Server-side with a Kubernetes Service:
checkoutknows one address. Connections are spread bykube-proxy, which balances per connection, not per request: it picks a pod once, when a TCP connection is first opened, and every request sent over that connection afterward goes to that same pod. Since gRPC keeps one HTTP/2 connection open and multiplexes many requests over it (see above), acheckoutinstance that opens one long-lived HTTP/2 connection sends all of its requests, forever, to whichever single pod answered first, which is a classic cause of uneven load for gRPC. Fixes: a headless Service plus client-side balancing, or a proxy that balances per request.
When to choose each
- Choose server-side when you have many languages, want one place to change routing, or run on a platform that already provides it (Kubernetes, cloud load balancers). This is the right default for most teams.
- Choose client-side when the extra hop matters (very latency-sensitive internal calls), when you need per-request balancing of long-lived connections (gRPC), or when you want each caller to react to failures instantly.
- A service mesh blends them: the logic lives in a proxy (server-side style, one implementation for all languages), but the proxy sits next to each client (client-side style, no central hop). It costs you operating the mesh.
Pitfalls
- Registering an instance as soon as the process starts, before it can serve (use readiness, not liveness, to decide when it receives traffic: a liveness probe only answers "is this process stuck and needs restarting", it says nothing about whether the process is ready to handle requests right now).
- Clients that cache the instance list with no refresh, so they keep calling instances that were removed.
- Treating the registry as the only health signal: a registry probe every 10 s means up to 10 s of requests to a dead instance unless the caller also reacts to errors.
Unlock Full Question Bank
Get access to all 16 Service Discovery and Configuration Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.