Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
You're tasked with selecting and driving adoption of a company-wide service discovery and configuration platform. Multiple teams prefer different tools (DNS, Consul, custom service mesh). How would you evaluate options technically and operationally, build a migration plan, address concerns such as operational burden and vendor lock-in, and gain cross-team buy-in while minimizing disruption?
Sample Answer
Direct answer
I would not run this as a tool bake-off. I would first write down what the company needs discovery to do (the requirements every team can agree on), score the options against those with weights the teams help set, and then commit to a layered standard: one authoritative service registry as the source of truth, DNS as the universal read interface so every language and legacy system works on day one, and a service mesh as an opt-in layer for teams that need mutual TLS and traffic shaping. Migration runs in reversible phases (dual-register, shadow, cut over by tier, decommission) led by a pilot with the most sceptical team, with success measured and published before anything is mandated.
Terms first
- Service discovery: how a caller finds the current network address of a service it wants to talk to.
- Service registry: the database of which instances of each service are alive and where (Consul is a common example; Kubernetes keeps its own).
- Service mesh: a proxy (a sidecar) runs next to every instance and handles discovery, retries, encryption and routing, programmed by a central control plane. Istio and Linkerd are examples.
- mTLS (mutual TLS): both sides of a connection prove their identity with certificates, not just the server.
- Vendor lock-in: when leaving a tool later would cost more than it is worth, usually because application code or data formats depend on it.
1. Evaluate: requirements first, then options
Run short interviews with each camp and turn their preferences into requirements. "We like DNS" usually means "it works with every client and needs no library"; "we want the mesh" usually means "we need mTLS and canary routing". Requirements, not tools, get weights.
| Criterion (weight) | Plain DNS | Consul registry (+ its DNS interface) | Custom service mesh |
|---|---|---|---|
| Works with every language and legacy host (25%) | yes | yes, via DNS; richer via API | needs a sidecar on every host |
| Health-aware, fast failover (20%) | no health, TTL-bound | health checks, seconds | health plus per-request retries |
| Multi-datacenter support (15%) | manual | built-in federation and failover queries (Consul's own mechanism for linking each datacenter's cluster to the others and automatically querying a nearby one when the local cluster has no healthy instances) | depends on the build |
| Security: mTLS, identity (15%) | none | optional (Consul Connect: Consul's built-in service-mesh feature that adds mutual TLS between services) | core feature |
| Operational burden on platform team (15%) | lowest | moderate: a consensus cluster to run | highest: proxies, control plane, certs |
| Exit cost / lock-in (10%) | none | low if apps only use DNS names | high for a custom build: you own it forever |
The operational side matters as much as the technical: who is on call for it, how upgrades happen, what happens when it is down (does traffic keep flowing on cached data, or does everything stop?), and whether the team that built the custom mesh is still staffed to maintain it. A custom mesh scores well on features and badly on bus factor (how many key people could leave or be unavailable before nobody left understands the system).
Turning that into a decision: score each cell 0 (worst) to 10 (best), consistent with the qualitative read above, multiply by its criterion's weight, and sum:
| Criterion (weight) | Plain DNS | Consul registry | Custom service mesh |
|---|---|---|---|
| Works with every language (25%) | 10 | 9 | 5 |
| Health-aware failover (20%) | 2 | 8 | 9 |
| Multi-datacenter support (15%) | 3 | 9 | 6 |
| Security: mTLS, identity (15%) | 1 | 5 | 9 |
| Operational burden, inverted so lower burden scores higher (15%) | 9 | 6 | 2 |
| Exit cost / lock-in, inverted so lower lock-in scores higher (10%) | 10 | 8 | 2 |
| Weighted total | 0.25x10 + 0.20x2 + 0.15x3 + 0.15x1 + 0.15x9 + 0.10x10 = 5.85 | 0.25x9 + 0.20x8 + 0.15x9 + 0.15x5 + 0.15x6 + 0.10x8 = 7.65 | 0.25x5 + 0.20x9 + 0.15x6 + 0.15x9 + 0.15x2 + 0.10x2 = 5.80 |
Consul wins clearly, 7.65 against 5.85 and 5.80, which is why it becomes the source of truth below rather than a coin flip between three close options. DNS and the mesh land near each other for opposite reasons: DNS wins on universal compatibility and zero lock-in but loses hard on security and failover; the mesh wins on security and failover but loses hard on operational burden and lock-in. That is exactly why the recommendation layers them instead of picking one outright: Consul underneath for its lead on health-awareness and multi-datacenter support, DNS names as the interface every client already speaks, and the mesh opt-in only where a team's own security or failover need would flip its individual score.
Recommendation for a typical mixed estate: Consul (or the platform's native registry) as source of truth, DNS names as the contract applications code against, mesh opt-in. What would flip it: if more than roughly 80% of workloads already run on Kubernetes and a regulator requires mTLS everywhere, standardise on a supported mesh as the default instead, and treat DNS as the legacy bridge.
2. Migration plan
(Two terms used below: tier-1 means the most critical, highest-traffic services, where an outage matters most; the long tail means the large number of smaller, lower-priority services that are slow to migrate voluntarily.)
- Contract first. Publish the naming scheme (
<service>.service.<dc>.<domain>), health-check requirements, and the rule that application code references names, never tool-specific APIs. This is the lock-in defence: if apps only know DNS names, the registry behind them can be swapped. - Dual registration. Services register in both old and new systems, automatically from the deploy pipeline, so nobody hand-edits two places. Nothing reads the new system yet.
- Shadow reads and diffing. A job resolves every service through both systems and reports mismatches. Cutover of a service is blocked until its diff has been clean for a week.
- Cut over by tier. Start with internal tools, then non-critical services, then tier-1. Each cutover is a config flag per caller, so rollback is a flag flip, not a redeploy.
- Decommission the old path per service once its traffic on the old path is zero for 30 days, with a published date.
Worked example: sizing the timeline
Say 300 services owned by 40 teams. The pilot (weeks 1 to 6) moves 10 services with the platform team doing the work. After that, a paved-road script means a team can move one service in about half a day, and the platform team can support about 15 cutovers per week without degrading review quality.
- Remaining after pilot: 300 - 10 = 290 services.
- At 15 per week: 290 / 15 = 19.3, so 20 weeks of rollout.
- Total: 6 + 20 = 26 weeks, roughly two quarters, plus the 30-day decommission tail.
State it as a plan with that arithmetic, so leadership can see what adding or removing platform capacity does to the date.
3. Addressing the specific concerns
- Operational burden: the platform team owns the registry cluster, its upgrades and on-call, with a published SLO. Application teams own only their health checks. Quantify it: a 5-server consensus cluster per datacenter, which tolerates the loss of 2 servers (a write only needs a majority, 3 of the 5, to agree before it commits; losing any 2 still leaves 3, a majority, so the cluster keeps accepting writes, while losing 3 would leave only 2, which can no longer form one).
- Vendor lock-in: the primary defence is contractual, not technical: apps depend on DNS names and standard health endpoints, not on a client SDK, so the registry behind those names can be swapped without touching application code. As a secondary, deeper defence inside the mesh itself, it also uses open standards (Envoy's xDS configuration APIs, a standard protocol for pushing routing config to sidecar proxies, so the proxies do not have to change even if the control plane does; and SPIFFE identities, a vendor-neutral standard for verifiable service identity, so certificates are not tied to one mesh vendor) so even the mesh's control plane is replaceable. Write down the exit plan now, with its cost.
- "Our custom mesh works fine": respect the investment. Offer to make it a candidate for the opt-in mesh layer if it passes the same criteria, and ask its owners to lead the mesh working group. Most resistance is about losing ownership, not about technology.
4. Gaining buy-in
- Publish the decision as an RFC (request for comments) with the weighted scoring, and let teams challenge the weights before the scores. Changing a weight is a cheap, visible concession.
- Pilot with the loudest sceptic. If their pain is solved, that is the most credible endorsement you can get.
- Make the paved road (the supported, recommended way of doing something, made deliberately the easiest path) cheaper than the old road: generated configs, dashboards and alerts for free on the new platform; the old one gets security fixes only.
- Measure and publish: services migrated, lookup failure rate, time to failover, incidents attributable to discovery, before and after.
- Escalate last. Executive mandate is for the long tail after the default is clearly better, not for the opening move.
Pitfalls
- Choosing by committee vote instead of weighted requirements: the loudest team wins, and the others disengage.
- Big-bang cutover: no rollback, and one bad week kills trust in the platform for years.
- Letting apps import the registry's client library "just for one feature": that is how lock-in re-enters.
- Forgetting non-Kubernetes workloads (VMs, databases, batch jobs): DNS is what keeps them in the plan.
Compare Consul, etcd, ZooKeeper, and Eureka as backends for service discovery and configuration storage. Which are CP and which are AP, and what does that mean operationally, for example ZooKeeper being unavailable during leader election versus Eureka's self-preservation mode serving stale data during a partition? Discuss watch/notify semantics, performance characteristics, operational complexity, and typical failure modes.
Sample Answer
Direct answer
etcd, ZooKeeper and Consul's server cluster are CP systems: they use a consensus protocol (an algorithm, such as Raft or ZAB, that gets every replica to agree on one ordered history of writes, normally by electing a single leader that orders and replicates each write before it counts as committed, which is why "no leader" means "no writes" in these systems), so during a network partition the side without a majority stops accepting writes rather than risk disagreeing. Eureka is AP: every server accepts registrations and keeps answering during a partition, even if the answer is stale. Operationally, CP means "correct or unavailable" and AP means "always answers, possibly wrong". My default is a CP store for configuration, locks and leader election, and discovery clients that tolerate staleness (cache the last good answer, retry on a different instance), because for routing a slightly old list beats no list.
CP and AP in plain terms
The CAP theorem says that when a network partition splits a cluster (some nodes cannot talk to others), a replicated system must choose between consistency (every read sees the latest write, or fails) and availability (every request gets an answer). CP systems choose consistency, AP systems choose availability. When there is no partition, both behave well; the choice only shows during failures, which is exactly when you are debugging at 3 a.m.
Quorum is the majority needed to make a decision: 3 of 5 nodes, 2 of 3. A 5-node CP cluster survives 2 node failures; a partition that splits it 3 and 2 leaves the 3-node side working and the 2-node side unable to write.
The four systems side by side
The table below is dense; for the CP-versus-AP question itself, the first three rows (Replication, CAP stance, Reads) carry the answer. The remaining rows (health/liveness through operational load) are what you'd use to justify picking one system over another, not to argue CP versus AP.
| etcd | ZooKeeper | Consul | Eureka | |
|---|---|---|---|---|
| Replication | Raft consensus | ZAB (ZooKeeper Atomic Broadcast) consensus | Raft among servers; gossip (Serf) among all agents for membership, where each node regularly swaps "who is alive" news with a few random peers | Peer-to-peer copying between servers, no consensus |
| CAP stance | CP | CP | CP for the catalog and KV (key-value) store; reads can opt into stale mode | AP |
| Reads | Linearizable (always latest) by default; serializable (local, possibly stale) on request | Served by whichever server the client is connected to, so possibly stale; call sync() first to catch up | default (strong except briefly during leader changes), consistent (strong, extra round trip), stale (any server, works even with no leader) | From the local registry copy, refreshed every 30 s by delta (an incremental update, not a full resend) |
| Watch / notify | Streaming watch from a given revision; resumable after disconnect; all events ordered, none skipped | Standard watches are one-time triggers: re-register after each event and you can miss intermediate changes. Persistent and recursive watches (addWatch, since 3.6) fix that | Blocking queries: long-poll (a request the server holds open, without answering, until something changes or a timeout passes, instead of the client asking repeatedly) with ?index= from the last X-Consul-Index (a response header carrying the current index, which the client echoes back to mean "tell me only about changes after this point"); a return does not guarantee a change, and an index that goes backwards must reset to 0 | None: clients poll for deltas (an incremental update containing just what changed, not the whole registry) every 30 s |
| Health / liveness | Leases (a TTL, time-to-live, the client must keep alive) | Ephemeral nodes deleted when the client's session expires | Agent runs health checks (HTTP, TCP, script) locally on each node | Client heartbeats every 30 s; evicted after 90 s without one |
| Built for | Kubernetes state, config, coordination | Coordination for Hadoop-era systems | Service discovery across datacenters, with DNS and health checks built in | Service registry for Spring Cloud / Netflix JVM stacks (a family of Java frameworks for building microservices; Eureka was originally Netflix's own registry) |
| Performance shape | Every write waits for the leader to replicate to a majority and fsync to disk (force the write out to physical disk, not just an in-memory buffer), so disk and network latency set write speed; linearizable reads also check with the leader, serializable reads scale with members | Reads served locally, so adding servers (or non-voting observers, ZooKeeper servers that receive the data stream and serve reads but don't count toward quorum) scales reads; writes all go through the leader; the whole data set must fit in memory | Writes go through the Raft leader; stale reads spread across all servers; many blocking queries on a large service list cost server CPU | Reads come from each client's local copy, so they are nearly free; registrations copy between servers asynchronously, so they are cheap but slow to converge |
| Operational load | Low-moderate: 3 or 5 members, watch disk latency and defragment (reclaim disk space after old data has been compacted away, since the storage file doesn't shrink on its own); 2 GiB default storage quota (8 GiB suggested maximum) | Moderate-high: JVM tuning (adjusting the Java Virtual Machine's memory and garbage-collection settings, since ZooKeeper runs on the JVM and a mistuned JVM causes long pauses), whole data tree held in memory, session timeouts (how long the server waits for a disconnected client to reconnect before treating its ephemeral state as gone) to tune | Moderate: servers plus an agent on every node, ACLs (access control lists), gossip encryption | Low to run, but needs clients that handle staleness |
What CP vs AP means during a partition
ZooKeeper during leader election. When the leader dies or is cut off, the ensemble cannot accept writes until a new leader is elected by a quorum. Clients connected to servers on the minority side get disconnected and, if the session times out, their ephemeral nodes (the entries that represent "this instance is alive") are deleted. A discovery design built on ephemeral nodes can therefore remove healthy instances because of a partition between the instances and ZooKeeper, not because the instances died.
Eureka self-preservation. Eureka expects heartbeats at a steady rate. If renewals fall below a threshold (by default, when more than 15% of registered instances are overdue), the server assumes the problem is the network, not mass instance death, and stops evicting anyone. The registry then serves dead instances until renewals recover. That is the AP trade: during a real mass failure, clients keep receiving addresses that no longer answer.
etcd. The minority side refuses writes and, with default linearizable reads, refuses reads too. A Kubernetes control plane (the components that make scheduling and management decisions, such as the API server and scheduler) on the minority side cannot schedule anything, but the workloads already running keep running, because the data plane (the running workloads themselves and the networking that serves their live traffic) does not depend on etcd for every request.
Consul. Writes to the catalog need the Raft leader. Clients that use stale reads (and Consul's DNS interface allows stale reads by default) keep getting answers from any server, even with no leader, so discovery continues while registration of new instances pauses.
Worked example: when does Eureka stop evicting?
100 instances, heartbeat every 30 s, so the server expects
100×3060=200 renewals per minuteand with the default 85% threshold it enters self-preservation below
0.85×200=170 renewals per minute(Eureka's exact computation of the expected count has more detail, but this is the shape.) If a switch fault cuts off 20 instances, renewals drop to 80 × 2 = 160, below 170, so nothing is evicted. If those 20 instances were genuinely dead, a client choosing uniformly at random hits a dead one 20% of the time until they return or an operator acts. The client-side mitigation is a short connect timeout plus a retry on a different instance: with independent random picks, two failures in a row happen 0.2 × 0.2 = 4% of the time. The same fault on ZooKeeper would have taken the other path: those 20 instances' sessions expire and they vanish from the registry, which is correct if they are dead and a self-inflicted outage if only their link to ZooKeeper failed.
Failure modes to name in an interview
- etcd: slow disk (fsync latency) causes leader elections; storage quota reached puts the cluster into alarm mode where it rejects writes until you compact (discard old key revisions no longer needed) and defragment; large watch fan-out on one prefix loads the leader.
- ZooKeeper: long JVM garbage-collection pauses look like dead servers or expired sessions; a "herd effect" when many clients watch one node and all wake up at once.
- Consul: gossip misconfiguration (encryption keys, ports blocked) causing agents to flap between alive and failed; a lost quorum of servers blocks all registrations; a very large catalog makes blocking queries on the whole service list expensive.
- Eureka: stale entries during self-preservation; up to roughly 90 seconds, three heartbeat intervals at the default 30-second heartbeat, for a change to reach every client (Spring Cloud Netflix's Eureka documentation: "a service is not available for discovery by clients until the instance, the server, and the client all have the same metadata in their local cache (so it could take 3 heartbeats)"), and longer still when an intermediate server's own response-cache refresh cycle adds another hop on top.
Recommendation and what would flip it
- Running on Kubernetes: use the platform's own discovery (Services backed by etcd) and do not add another registry unless you span clusters.
- Many datacenters, VMs and containers mixed, want DNS-based discovery and health checks: Consul.
- New design choosing ZooKeeper: rarely justified now; the main reason is an existing system that depends on it. Kafka, its best-known user, removed ZooKeeper mode in Kafka 4.0 in favour of its own Raft-based controller.
- Eureka: fine inside an existing Spring Cloud estate where clients already retry and tolerate staleness; I would not introduce it into a new, non-JVM system.
Pitfalls
- Putting service discovery on the critical path of every request instead of caching locally, so a registry outage becomes a full outage.
- Treating a ZooKeeper one-time watch as a change stream and silently missing updates.
- Using an even number of consensus nodes: 4 nodes need 3 for quorum and still tolerate only 1 failure, the same as 3 nodes but with more to go wrong.
- Assuming "CP" means "always correct for clients": ZooKeeper reads can still be stale without
sync(), and Consul'sdefaultmode has a brief stale window during leader changes.
Design a service discovery and routing strategy for microservices deployed across multiple Kubernetes clusters in different regions. Consider DNS vs service mesh, global load balancers, cross-cluster health checks, stale registry entries, and latency/consistency trade-offs for routing decisions.
Sample Answer
Direct answer
Route with a strict locality preference: a call goes to an endpoint in the same cluster if one is healthy, then another cluster in the same region, and crosses regions only on failure. For traffic between services (east-west) I would use a service mesh running one control plane per cluster that also discovers the other clusters' endpoints, because it can fail over per request in seconds and eject bad endpoints based on real traffic. For users entering the system (north-south) I would use a global load balancer (anycast: one IP address advertised from every region, so the network itself delivers each user to the nearest one; or geo-DNS with health checks) in front of per-region ingress (the entry point that receives external traffic for that region, before it reaches any cluster). Plain DNS is the fallback for teams without a mesh and for coarse, region-level failover, not for fine-grained cross-cluster routing, because TTL (time-to-live) caching makes it slow to react.
Terms used below
- Service mesh: a sidecar proxy (a small proxy process running beside each service instance, such as Envoy) that handles every outgoing call, plus a control plane that tells all proxies which endpoints exist. Istio is a common example.
- East-west / north-south: traffic between internal services versus traffic entering from users.
- Outlier detection: a proxy watches real responses and temporarily ejects an endpoint that returns consecutive errors, without waiting for a health check.
- Locality: region, then zone, then cluster; "locality-aware" routing prefers the nearest healthy option.
Setup I am designing for
Three regions, two Kubernetes clusters per region (six clusters), tens of microservices, each deployed to every cluster. Requirements: low latency, survive losing a cluster or a region, no cross-region hop in the normal path.
Why locality matters, with an assumed cross-region round trip of 70 ms (a typical order of magnitude between distant regions, not a measurement): a user request that fans through a chain of 5 services pays nothing extra when all hops stay in-region, but about 5 × 70 = 350 ms extra if each hop is routed across regions. Random global load balancing across clusters would put most hops cross-region, since two thirds of endpoints live in other regions.
Architecture
flowchart TB
U[Users] --> GLB[Global LB: anycast or geo-DNS]
GLB --> I1[Region 1 ingress]
GLB --> I2[Region 2 ingress]
subgraph R1[Region 1]
I1 --> A1[Cluster 1a: services and sidecars]
A1 <--> B1[Cluster 1b]
end
subgraph R2[Region 2]
I2 --> A2[Cluster 2a]
A2 <--> B2[Cluster 2b]
end
A1 -. failover only .-> EW[East-west gateway in Region 2]
EW --> A2
Decisions
1. DNS versus service mesh
| DNS-based (e.g. Kubernetes Multi-Cluster Services, or health-checked DNS names per region) | Service mesh (per-cluster control planes sharing endpoint knowledge) | |
|---|---|---|
| Failover speed | Health-check detection plus TTL plus client DNS caches: typically a minute or more | Seconds: proxies get endpoint updates by streaming and retry per request |
| Granularity | Per name; the client picks an IP and keeps it on pooled connections | Per request, per endpoint, with weights and locality |
| Reacts to real errors | No | Yes, outlier detection |
| Operational cost | Low | High: control-plane upgrades, certificates, proxy resource overhead |
(Kubernetes Multi-Cluster Services above is a Kubernetes feature that makes a service's endpoints visible by name across every cluster in the set, the DNS-based counterpart to the mesh's own cross-cluster discovery.)
Choice: mesh for east-west. I would flip to DNS if the team has few services, low change rate, and cannot staff mesh operations; then use one DNS name per service per region, health-checked, and accept minute-scale failover.
2. Global load balancer for entry traffic
Anycast (one IP advertised from every region, so the network delivers users to the nearest one) fails over in seconds without depending on client DNS caches; geo-DNS with 30 to 60 second TTLs is the cheaper alternative. Either one health-checks a deep endpoint per region, and each region must hold headroom for a failed neighbour's traffic.
3. Cross-cluster health checks
- Do not build a central prober that health-checks every pod (Kubernetes' smallest deployable unit, one or more containers running together) in every cluster: it is a scaling bottleneck and a single point of failure, and a partition between the prober and a healthy cluster would falsely evict it.
- Each cluster's kubelet (the agent that runs on every node and starts, stops and health-checks the pods scheduled to it) readiness probes remain the source of truth for its own pods; the local control plane publishes only ready endpoints.
- Across clusters, rely on (a) the health of each cluster's east-west gateway (the ingress point that receives traffic from other clusters) and (b) passive outlier detection on real requests. A remote endpoint that errors is ejected by the caller's proxy within a few failed requests.
4. Stale registry entries
Each cluster's control plane reads the other clusters' API servers (the Kubernetes API server: the per-cluster endpoint that holds and serves all of that cluster's state, including which pods and services exist) to learn their endpoints. When a remote cluster becomes unreachable, its endpoint list freezes, and those entries go stale.
- Keep the last known list rather than deleting it (a network blip should not remove a healthy cluster), but track its age and demote stale localities: they receive traffic only if no fresher locality is healthy.
- Outlier detection ejects stale endpoints that are actually dead on the first few failed calls.
- Guard against mass removal: if an update would drop most endpoints of a service at once, hold the old list and alert. (Envoy's panic threshold, which load-balances across all hosts when fewer than 50% are healthy by default, is a built-in version of this idea.)
- Product-specific gotcha, not a general rule: in Istio specifically, locality failover only engages when outlier detection is configured on the destination; without it, traffic keeps flowing to the preferred locality even when it is failing. Check this setting explicitly if you adopt Istio; other meshes wire the two together differently.
5. Latency versus consistency in routing decisions
Routing data is eventually consistent by design: endpoint lists propagate in seconds and can briefly disagree between clusters. That is acceptable because every call is protected by retries on connection failure and outlier detection. Making routing strongly consistent (a global registry with consensus on every endpoint change) would put cross-region commits into the discovery path and stall discovery during a partition, which is far worse than routing on a view that is a few seconds old.
Trade-offs and pitfalls
- Flat global load balancing across all clusters: simple and wrong; most hops go cross-region, adding latency and egress cost (the fee cloud providers charge for data leaving a region).
- Failover cascades: when region 1 fails over to region 2, region 2 must have spare capacity, or it fails too. Cap cross-region spillover and shed load (deliberately reject or slow down some requests rather than let every node overload) rather than letting it cascade.
- One global mesh control plane: a single failure or bad config push affects every cluster; run one control plane per cluster so a control-plane problem stays local.
- Trust across clusters: cross-cluster calls need mutual TLS (both sides present certificates) with a shared root of trust (a common certificate authority that every cluster's certificates are issued from, so any cluster can verify a certificate it has never seen before), or clusters cannot authenticate each other's workloads.
Design a global multi-region service discovery and configuration system for an application deployed across 5 regions serving 1M RPS. Requirements: region-local discovery for low latency, automatic regional failover, versioned global config propagation with audit history, and ability to perform region-scoped rollbacks. Discuss DNS TTLs, geo-DNS, control plane replication, and trade-offs between strong and eventual consistency.
Sample Answer
Direct answer
Split it into two systems with opposite consistency needs. Service discovery is region-local and eventually consistent: each of the 5 regions runs its own registry, clients discover only same-region endpoints, and no request ever waits on another region. Configuration has one global, strongly consistent write path that assigns every change a version and an audit record, then replicates asynchronously to a read-only config store in each region; each region points at a version, so a rollback moves one region's pointer back without touching the others. Users reach a region through geo-DNS with a short TTL (time-to-live) plus health checks, and every region is provisioned to absorb a failed neighbour's share.
Requirements and numbers I am designing to
- 1M requests per second (RPS) globally over 5 regions: about 200,000 RPS per region in steady state.
- Regional failover: if one region dies, its traffic spreads over the other 4, so each must handle 1,000,000 / 4 = 250,000 RPS. Each region is therefore provisioned at 25% above its normal load (250,000 / 200,000 = 1.25).
- Assume about 4,000 service instances per region (an illustrative sizing, not a measurement). Discovery load is driven by instance churn and watches, not by the 1M RPS: clients cache endpoint lists, so a request never calls the registry.
- Config writes are rare (tens per hour); config reads happen at process start and on every change, never per request.
Terms used below
- Registry: the database of "which healthy instances serve service X", e.g. Consul or the Kubernetes API server.
- Raft: a consensus protocol; a small cluster (3 or 5 nodes) elects a leader and commits a write only when a majority (the quorum) has stored it. It gives linearizable writes: once a write is acknowledged, every later read sees it.
- Strong vs eventual consistency: strong means every reader sees the latest committed value; eventual means readers may lag briefly but converge.
- Geo-DNS: an authoritative DNS service that answers the same name with different addresses depending on where the client is, and stops returning a region that fails health checks.
Architecture
flowchart TB
U[Users] --> G[Geo-DNS with health checks]
G --> R1[Region A edge LB]
G --> R2[Region B edge LB]
subgraph CP[Global config control plane]
W[Config API, Git review] --> S[(Raft store: versions and audit log)]
end
S -->|async replication| C1[(Region A config cache)]
S -->|async replication| C2[(Region B config cache)]
R1 --> A1[Services in A]
A1 --> D1[(Region A registry)]
A1 --> C1
R2 --> B1[Services in B]
B1 --> D2[(Region B registry)]
B1 --> C2
Two regions are drawn for readability; the other three are identical. Each region's edge LB (edge load balancer: the boundary layer that terminates incoming user connections and forwards them to services inside that region) sits behind geo-DNS and in front of the region's own services.
1. Region-local discovery
- Each region runs its own registry cluster (a 5-node Raft cluster across that region's availability zones, meaning physically separate data centers within the region, each with independent power and networking, so the cluster survives one of them failing, or the Kubernetes control plane per cluster). There is no global registry: a cross-region consensus round trip (tens to over a hundred milliseconds of network delay per commit) has no place in the discovery path, and a trans-oceanic partition must not break discovery inside a region.
- Health does not go through Raft. If 4,000 instances each wrote a heartbeat through consensus every 10 seconds, that is 4,000 / 10 = 400 consensus writes per second per region for nothing but liveness. Instead, a local agent or the kubelet (the per-node Kubernetes agent that already runs on every node and reports pod health) probes each instance and only state changes (healthy to unhealthy) become registry writes.
- Clients or sidecar proxies (a proxy process deployed next to each service instance, as in a service mesh) hold a watch on the registry and keep an in-memory endpoint list, so discovery reads cost nothing per request and survive a registry outage using the last known list.
- Endpoint entries carry version tags (
orders v42), so a rolling upgrade can route by version, keep old and new side by side, and drain old instances cleanly.
2. Automatic regional failover
- User to region: geo-DNS answers with the nearest healthy region. Health checks probe a deep endpoint in each region (one that exercises the region's real dependencies) every 10 seconds and fail it after 3 misses. DNS TTL is 30 seconds, so worst-case steering time is about 30 s detection + 30 s TTL = 60 s, plus clients that hold connections open (cap connection lifetime so they re-resolve). An anycast global load balancer (one IP address announced from every region at once, so network routing itself, not DNS or client caching, sends each user to the nearest healthy region and can redirect within seconds of a region failing) can replace geo-DNS and fail over in seconds without depending on client DNS caches; I would use one where available and keep geo-DNS as the fallback.
- Service to service: stays in-region. If a single service is down in region A but the region is otherwise up, the local proxy can fail over to that service in the nearest region, but only when the registry's failover policy allows it, because it adds cross-region latency and egress cost (what cloud providers charge for data leaving a region) and can overload the neighbour.
- Capacity: the 25% headroom above is non-negotiable; without it a regional failover becomes a multi-region outage as the surviving regions overload one after another.
3. Versioned global config with audit history
- Config lives in a single global control-plane store: a 5-node Raft cluster with one node per region. Writes need 3 of 5 regions, so a write commits even with 2 regions unreachable, and each commit costs one cross-region round trip (the leader sends the write to the other 4 regions in parallel and only has to wait for the fastest 2 replies, since itself plus those 2 is already a majority of 3; because the sends and replies happen in parallel, it is one round trip, not four sequential ones), acceptable for tens of writes per hour.
- Every write is: reviewed change in Git, then applied through the config API as
(key, value, version, author, change_id, reason)with a monotonically increasing version, appended to an append-only audit log in the same transaction. Nothing edits config directly in a region. - Each region runs a read-only config replica fed by an asynchronous stream from the control plane. Services read only from their region's replica (via a local agent that watches it and writes a file or pushes to the process). If the control plane is unreachable, regions keep serving last known good config: config staleness is acceptable, config unavailability is not.
- Replicas publish their applied version; an alert fires if any region's version lags the global head by more than a few minutes.
4. Region-scoped rollouts and rollbacks
- The control plane stores immutable versions plus, per region, a pointer:
region-a -> v118,region-b -> v117. A rollout advances pointers region by region with a bake time between (a deliberate pause after each region's change, watching its metrics before touching the next one), starting with the lowest-traffic region. - Rollback of region B is a single write:
region-b -> v117. The versions themselves never change, so the rollback is exact, audited like any other change, and does not affect regions A, C, D, E. - Every config read reports the version it served, so "which version is region B running" is answered from telemetry, not from what the control plane intended.
Strong versus eventual consistency: where each belongs
| Concern | Consistency | Why |
|---|---|---|
| Config writes, versions, audit log, region pointers | Strong (global Raft) | Two operators must not both create "v119"; audit must be complete; rollbacks must target an exact version |
| Config reads in a region | Eventual, bounded by replication lag | Reading from a local replica keeps reads fast and independent of cross-region links |
| Service discovery | Eventual, region-local | Endpoint lists are stale by seconds anyway (health checks take seconds); availability matters more |
| Geo-DNS | Eventual, bounded by TTL | Clients cache by design |
The rule behind the table: pay for strong consistency only where a stale or conflicting value is a correctness bug, and never on the request path.
Trade-offs and pitfalls
- A single global registry (one Raft cluster spanning all 5 regions for discovery) is the common wrong design: every instance change pays a cross-region commit and a partition can stall discovery everywhere.
- Async config replication means a brief window where regions run different versions. That is intended (it is how region-by-region rollout works), but configs must be backward compatible across adjacent versions.
- Low DNS TTLs are not failover by themselves: pooled connections and runtime DNS caches ignore them, and the TTL only starts after health checks detect the failure.
- What would change the design: if config must be read-your-writes consistent across regions (read-your-writes consistent: a client that just wrote a value is guaranteed to see that same value on its very next read from any region, never a stale one; for example a security revocation), send it through the control plane with a synchronous "applied in all regions" acknowledgement instead of the async path; if regions are fully independent tenants, drop the global Raft and run five independent config stores synced from Git.
Design a service discovery strategy for a microservices platform operating across multiple datacenters. Compare DNS-based discovery, client-side discovery with a registry like Consul/etcd, and service-mesh-based discovery. Discuss trade-offs in latency, consistency, failover behavior, and operational complexity.
Sample Answer
Direct answer
Run one independent registry per datacenter (never one consensus cluster stretched across datacenters), make every lookup local-first, and fail over to the nearest healthy datacenter through an explicit, per-service failover policy. On top of that, choose the lookup mechanism by client type: DNS as the universal baseline, client-side discovery from the registry for services that need fast, health-aware balancing, and a service mesh where you also need mutual TLS (both sides of a connection proving their identity with certificates, not just the server) and fine-grained traffic control. For a typical platform I would commit to: Consul (or etcd-backed equivalent) per datacenter, federated (each datacenter's registry stays independent but can answer queries about the others through a lightweight cross-datacenter link); its DNS interface for everyone; mesh adopted per service where the security or routing need justifies the extra moving parts.
Terms first
- Service registry: the live list of healthy instances per service (Consul, etcd, ZooKeeper, or Kubernetes' own API).
- Client-side discovery: the calling service queries the registry itself (usually through a library) and picks an instance.
- Server-side / DNS discovery: the caller just resolves a name; something else (DNS server, load balancer) decides.
- Service mesh: a proxy sidecar next to every instance does discovery, load balancing, retries and encryption, fed by a control plane.
- Raft quorum: registries like Consul and etcd agree on writes using the Raft protocol, a consensus algorithm in which one elected leader orders every write and a majority of servers must acknowledge it before it is durable.
- Failover: sending traffic elsewhere when the preferred instances are unhealthy.
- Watch: a long-held request a client keeps open against the registry; instead of the client re-asking on a timer, the registry pushes each change down that open connection the moment it happens.
- Outlier detection (ejection): a proxy watches the real responses a backend returns and temporarily stops sending it traffic after enough of them look like failures, without waiting for a separate health check to catch up.
- Locality / locality-aware failover: preferring an endpoint in the caller's own datacenter first, and reaching into another one only when nothing local is healthy; "locality-weighted" means splitting traffic across localities by a configured ratio rather than all-or-nothing.
Why not one global registry
A Raft cluster commits a write only after a majority acknowledges it. Spread 5 servers across 3 datacenters that are 70 ms apart, a typical 2-2-1 split, and no single datacenter holds the 3-of-5 majority a write needs on its own, so every registration, deregistration and health update pays at least one cross-datacenter round trip (about 70 ms) before it commits, and a network partition can leave the minority side (whichever datacenters are left with fewer than 3 reachable servers) unable to register anything at all. Keeping each datacenter's registry local means local writes commit in about a millisecond and a partition leaves every datacenter fully functional on its own. Consul's design follows this: servers form a Raft cluster per datacenter, and datacenters are joined by a lighter WAN (wide area network) gossip layer (nodes periodically exchange state with a few random peers instead of every node talking to every other one, so it stays cheap even at scale) used for cross-datacenter queries, not for shared consensus.
Comparison
| DNS-based | Client-side with registry | Service mesh | |
|---|---|---|---|
| Lookup latency | one DNS query, then cached for the TTL (time-to-live); near zero on cache hits | in-memory lookup in the library, kept fresh by a watch; near zero | proxy already has endpoints; adds one local proxy hop per request (sub-millisecond to low milliseconds) |
| Consistency / freshness | stale up to the TTL plus any client caching that ignores TTL | seconds: watch pushes changes as they commit | seconds: control plane pushes endpoint updates to every proxy |
| Failover behaviour | coarse: slow (TTL-bound) and per name, no per-request retry | fast and per request, but only as good as each language's library | fastest and uniform: retries, outlier ejection and locality failover in the proxy |
| Cross-DC failover | needs a geo-aware DNS answer or registry-generated records | library applies a failover list | locality-aware load balancing built in |
| Operational complexity | lowest | medium: registry cluster plus a client library per language to maintain | highest: control plane, sidecar on every pod, certificate authority (the component that issues and signs the certificates every proxy uses to prove its identity for mTLS), upgrades of both |
| Language coverage | every client | only languages with a maintained library | every client (it is out of process) |
Worked example: failover time after an instance dies
Assume health checks every 10 s, marked critical after 3 failures, DNS TTL of 30 s.
- DNS: detection 3 x 10 = 30 s, then clients may keep the old answer for up to 30 s more: worst case about 60 s of some traffic still going to the dead instance (longer if a client caches past the TTL).
- Client-side with watch: detection 30 s, then the watch delivers the change in well under a second: about 30 s, and a library that retries on connection failure hides most of that window.
- Mesh: same 30 s for the registry path, but the proxy also does outlier detection on live traffic (Envoy, a widely used sidecar proxy, ejects a host by default after 5 consecutive 5xx errors, the HTTP status codes that signal a server-side failure), so a hard-down instance stops receiving requests after a handful of failed requests, often before the health check notices.
The recommended design
flowchart TB
subgraph DC1[Datacenter 1]
R1[Registry cluster, 5 servers] --- A1[Services: DNS or library or sidecar]
end
subgraph DC2[Datacenter 2]
R2[Registry cluster, 5 servers] --- A2[Services: DNS or library or sidecar]
end
R1 <-->|WAN federation: cross-DC queries only| R2
A1 -->|failover only when local instances unhealthy| A2
- Local-first resolution.
payments.service.consul(or the mesh equivalent) returns only instances in the caller's datacenter by default. - Explicit failover per service. In Consul this is a prepared query with a failover list ("nearest 2 datacenters if no local healthy instances"); in a mesh it is locality-weighted load balancing. Opt-in per service, because failing over a service whose database is local can make things worse (cross-DC latency on every query, or writes against a read-only replica).
- Mechanism by client type. Legacy and third-party software uses DNS. Latency-sensitive internal services use the client library or the mesh. Services needing zero-trust networking (verifying every call cryptographically instead of trusting it just because it came from inside the network perimeter; in practice, mTLS between every pair) use the mesh.
- Configuration stays per datacenter too. Consul's key/value store is not replicated across datacenters by default; replicate deliberately (a tool that copies a prefix from a primary datacenter) and treat the primary as the only writer.
Trade-offs and pitfalls
- Cross-DC failover can cause a cascade. If DC1 fails over its 10,000 requests per second onto DC2, DC2 needs that headroom. Cap failover traffic or pre-provision for it.
- DNS cache ignorance is the classic outage: some runtimes pin an answer far longer than the TTL. Test every major language's behaviour, do not assume.
- A mesh doubles your upgrade surface (every sidecar plus the control plane), and a control-plane outage must leave proxies serving their last-known endpoints, not empty ones.
- What would change the recommendation: a single datacenter makes most of the federation machinery unnecessary; a hard requirement for mTLS everywhere makes the mesh the default rather than the exception; an all-Kubernetes estate can use Kubernetes' native discovery per cluster plus a multi-cluster layer instead of Consul.
That is every published Service Discovery and Configuration Management question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.