Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
As a staff engineer you must decide whether to centralize configuration via a strongly-consistent coordination service or to push configuration updates via an eventually-consistent CDN to thousands of services. Draft a migration plan, risk analysis, rollback strategy, and operational changes required (monitoring, runbooks, team training).
Sample Answer
Direct answer
I would not pick one model for everything. My decision: keep a strongly consistent store as the single source of truth for writes (small, audited, one writer path), but distribute configuration to the thousands of services through signed, versioned snapshots served by a CDN (content delivery network), with a lightweight change notification so updates arrive in seconds rather than at the next poll. The coordination service stays in the picture only for the minority of settings that truly need linearizable reads: leader election (a group of instances agreeing on exactly one of them being in charge), locks (only one process may hold a given resource at a time), and fencing (stopping a process that has lost its lock, for example after a long pause, from acting as if it still holds it). It is not needed for ordinary tuning config. Then migrate in reversible phases with dual-read verification, a one-step rollback at every phase, and new monitoring built around config age and version skew.
Terms first
- Strongly consistent coordination service: a store such as etcd, ZooKeeper or Consul where every write is agreed by a majority of servers (a quorum) before it is acknowledged, so all readers see one agreed order of changes.
- Linearizable: once a write is acknowledged, every later read anywhere sees it (or something newer). Needed for locks; rarely needed for a timeout value.
- Eventually consistent: every reader converges to the latest value, but for a while different readers can see different versions.
- CDN: a global network of caches that serves the same file to huge numbers of clients cheaply and close to them.
- Version skew: at a given moment, part of the fleet runs config version N and part runs N-1.
- Watch: a long-lived open connection a client holds with the coordination service so it is pushed a notification the instant a value changes, instead of asking repeatedly.
- Egress: data a service sends out to clients; the thing most networks and clouds meter and charge for, so "egress per day" below is both a bandwidth number and a cost number.
The decision and its arithmetic
Assume 3,000 services with 10 instances each, a 200 KiB full config snapshot, 300 config changes per day, polling every 30 s through a CDN edge TTL (time-to-live) of 10 s.
services, instances_per_service = 3_000, 10
clients = services * instances_per_service
snapshot_kib = 200
poll_s, cdn_ttl_s = 30, 10
changes_per_day = 300
print(f"config clients: {clients:,}")
print(f"consistent store: {clients:,} long-lived watches; on 5 servers = {clients // 5:,} per server")
print(f"CDN polling load: {clients / poll_s:,.0f} requests/s (mostly 304 Not Modified, meaning the file has not changed so nothing is resent)")
print(f"egress per change if every client fetches a full snapshot: {clients * snapshot_kib / 2**20:.1f} GiB")
print(f"egress per day at {changes_per_day} changes: {changes_per_day * clients * snapshot_kib / 2**30:.1f} TiB")
print(f"worst-case staleness, poll only: {poll_s + cdn_ttl_s} s")
Output (Python 3, run as shown):
config clients: 30,000
consistent store: 30,000 long-lived watches; on 5 servers = 6,000 per server
CDN polling load: 1,000 requests/s (mostly 304 Not Modified, meaning the file has not changed so nothing is resent)
egress per change if every client fetches a full snapshot: 5.7 GiB
egress per day at 300 changes: 1.7 TiB
worst-case staleness, poll only: 40 s
What the numbers say:
- Centralised reads concentrate risk. 30,000 clients all holding watches on a 5-server consensus cluster (6,000 each) turns the config store into a hard dependency of every service; a quorum loss or a reconnect storm (every dropped client trying to reconnect within the same few seconds) after a network blip hits all of them at once. The CDN absorbs 1,000 mostly-empty requests per second without effort, and the CDN being slow never blocks a service that already has its config.
- Full snapshots are the wrong unit. 1.7 TiB per day of egress comes from every client re-downloading everything on every change. Partition snapshots per service (each service fetches only its own namespace plus a small shared file), so a change to one service's config is fetched by its 10 instances, not 30,000.
- Staleness is the price. Poll-only gives up to 40 s of lag. A notification channel (a lightweight "version N+1 of namespace
paymentsexists" message over a pub/sub system or server-sent events) cuts typical lag to seconds, while polling remains the backstop if notifications are lost.
What would flip the decision: if most config changes must be applied atomically across many services at the same instant (rare in practice), or the fleet is a few hundred instances, the consistent store with watches is simpler and fine.
Risk analysis
| Risk | Impact | Mitigation |
|---|---|---|
| Stale config served from a CDN edge | services act on old values for longer than expected | versioned immutable file names (payments/v143.json) plus a tiny short-TTL pointer file (a small object that just says "the current version is v143"; everything heavy is the immutable file it points to, so only the tiny pointer needs a short cache time); monitor config age per instance |
| Version skew across the fleet during rollout | two behaviours at once, e.g. old and new rate limits | design config changes to be safe to run side by side; for the few that are not, use the coordination service |
| Tampered or corrupted snapshot | attacker or bug pushes config to thousands of hosts | sign every snapshot; clients verify signature and schema before applying and keep last good on failure |
| Signing key compromise | forged config accepted fleet-wide | key in a hardware-backed key management service (a managed system that generates and stores the signing key inside tamper-resistant hardware, so the raw key material is never exposed even to whoever operates it), short-lived signing via the publish pipeline only, rotation runbook |
| Bad config published | fleet-wide incident within seconds | staged rollout by cohort (canary services or percentage of instances) with automatic halt on error-rate regression |
| Cold start (a brand-new instance booting with nothing cached locally yet) when CDN and store are unreachable | new instances cannot boot | bake last-known-good config into the image or a local disk cache as a bootstrap fallback |
Migration plan
- Inventory and classify every key: "tuning" (timeouts, limits, feature toggles), "needs linearizable" (leader election, locks, anything used for mutual exclusion, meaning only one process may act at a time), "secret" (moves to the secret manager, never to a CDN). Expect the second group to be small.
- Build the pipeline behind the store: on every committed write, the publisher renders per-service snapshots, signs them, uploads them as immutable versions, updates the pointer, and sends the notification.
- Shadow mode: the client library reads from both paths and reports mismatches while still using the old one. Run until mismatch rate is zero for two weeks.
- Cut over by tier: internal tools, then non-critical services, then tier-1. Each service switches with a per-service flag.
- Decommission the watches on the coordination service once no client has used them for 30 days, and shrink it to the linearizable use cases.
Rollback strategy
- During migration: flip the per-service flag back to the coordination-service path. The old path stays fully live until step 5, so rollback is instant and needs no redeploy.
- For a bad config value after migration: re-point the pointer file to the previous version. Because versions are immutable, "previous" is always intact, and the rollback propagates through the same notification path in seconds.
- For a broken publisher: clients simply keep last good config; the pipeline can be fixed without an outage, which is the main resilience gain over a centralised read path.
Operational changes
- Monitoring: per-instance applied config version and config age (seconds since the version it runs was published); fleet version-skew dashboard (how many instances on each version, how long the newest took to reach 100%); fetch errors and signature failures; publisher lag from write to pointer update. Alert when config age exceeds, say, 5 minutes on more than 1% of a service's instances.
- Runbooks: rolling back a config version; "CDN unreachable" (services keep running; stop non-urgent changes); signing key rotation; forcing a fleet-wide refresh; emergency change when the publisher is down (a break-glass path: a rehearsed, fully audited manual override for exactly this situation).
- Team training: a short session per team on the new mental model (config is eventually consistent: never rely on two services flipping at the same instant), the client library's last-good behaviour, and how to read the version-skew dashboard. Run one game day (a planned exercise where the team deliberately breaks something on purpose to rehearse the response) where the CDN path is blocked and teams practise the runbook.
Pitfalls
- Moving locks or leader election onto the eventually consistent path. That turns rare staleness into two leaders at once.
- Letting a CDN cache the pointer file for long; everything else can be cached forever because it is immutable.
- Treating the migration as done at cutover. The payoff (a smaller blast radius: less of the system affected when any single thing fails) only arrives when the old watches are gone.
You're tasked with selecting and driving adoption of a company-wide service discovery and configuration platform. Multiple teams prefer different tools (DNS, Consul, custom service mesh). How would you evaluate options technically and operationally, build a migration plan, address concerns such as operational burden and vendor lock-in, and gain cross-team buy-in while minimizing disruption?
Sample Answer
Direct answer
I would not run this as a tool bake-off. I would first write down what the company needs discovery to do (the requirements every team can agree on), score the options against those with weights the teams help set, and then commit to a layered standard: one authoritative service registry as the source of truth, DNS as the universal read interface so every language and legacy system works on day one, and a service mesh as an opt-in layer for teams that need mutual TLS and traffic shaping. Migration runs in reversible phases (dual-register, shadow, cut over by tier, decommission) led by a pilot with the most sceptical team, with success measured and published before anything is mandated.
Terms first
- Service discovery: how a caller finds the current network address of a service it wants to talk to.
- Service registry: the database of which instances of each service are alive and where (Consul is a common example; Kubernetes keeps its own).
- Service mesh: a proxy (a sidecar) runs next to every instance and handles discovery, retries, encryption and routing, programmed by a central control plane. Istio and Linkerd are examples.
- mTLS (mutual TLS): both sides of a connection prove their identity with certificates, not just the server.
- Vendor lock-in: when leaving a tool later would cost more than it is worth, usually because application code or data formats depend on it.
1. Evaluate: requirements first, then options
Run short interviews with each camp and turn their preferences into requirements. "We like DNS" usually means "it works with every client and needs no library"; "we want the mesh" usually means "we need mTLS and canary routing". Requirements, not tools, get weights.
| Criterion (weight) | Plain DNS | Consul registry (+ its DNS interface) | Custom service mesh |
|---|---|---|---|
| Works with every language and legacy host (25%) | yes | yes, via DNS; richer via API | needs a sidecar on every host |
| Health-aware, fast failover (20%) | no health, TTL-bound | health checks, seconds | health plus per-request retries |
| Multi-datacenter support (15%) | manual | built-in federation and failover queries (Consul's own mechanism for linking each datacenter's cluster to the others and automatically querying a nearby one when the local cluster has no healthy instances) | depends on the build |
| Security: mTLS, identity (15%) | none | optional (Consul Connect: Consul's built-in service-mesh feature that adds mutual TLS between services) | core feature |
| Operational burden on platform team (15%) | lowest | moderate: a consensus cluster to run | highest: proxies, control plane, certs |
| Exit cost / lock-in (10%) | none | low if apps only use DNS names | high for a custom build: you own it forever |
The operational side matters as much as the technical: who is on call for it, how upgrades happen, what happens when it is down (does traffic keep flowing on cached data, or does everything stop?), and whether the team that built the custom mesh is still staffed to maintain it. A custom mesh scores well on features and badly on bus factor (how many key people could leave or be unavailable before nobody left understands the system).
Turning that into a decision: score each cell 0 (worst) to 10 (best), consistent with the qualitative read above, multiply by its criterion's weight, and sum:
| Criterion (weight) | Plain DNS | Consul registry | Custom service mesh |
|---|---|---|---|
| Works with every language (25%) | 10 | 9 | 5 |
| Health-aware failover (20%) | 2 | 8 | 9 |
| Multi-datacenter support (15%) | 3 | 9 | 6 |
| Security: mTLS, identity (15%) | 1 | 5 | 9 |
| Operational burden, inverted so lower burden scores higher (15%) | 9 | 6 | 2 |
| Exit cost / lock-in, inverted so lower lock-in scores higher (10%) | 10 | 8 | 2 |
| Weighted total | 0.25x10 + 0.20x2 + 0.15x3 + 0.15x1 + 0.15x9 + 0.10x10 = 5.85 | 0.25x9 + 0.20x8 + 0.15x9 + 0.15x5 + 0.15x6 + 0.10x8 = 7.65 | 0.25x5 + 0.20x9 + 0.15x6 + 0.15x9 + 0.15x2 + 0.10x2 = 5.80 |
Consul wins clearly, 7.65 against 5.85 and 5.80, which is why it becomes the source of truth below rather than a coin flip between three close options. DNS and the mesh land near each other for opposite reasons: DNS wins on universal compatibility and zero lock-in but loses hard on security and failover; the mesh wins on security and failover but loses hard on operational burden and lock-in. That is exactly why the recommendation layers them instead of picking one outright: Consul underneath for its lead on health-awareness and multi-datacenter support, DNS names as the interface every client already speaks, and the mesh opt-in only where a team's own security or failover need would flip its individual score.
Recommendation for a typical mixed estate: Consul (or the platform's native registry) as source of truth, DNS names as the contract applications code against, mesh opt-in. What would flip it: if more than roughly 80% of workloads already run on Kubernetes and a regulator requires mTLS everywhere, standardise on a supported mesh as the default instead, and treat DNS as the legacy bridge.
2. Migration plan
(Two terms used below: tier-1 means the most critical, highest-traffic services, where an outage matters most; the long tail means the large number of smaller, lower-priority services that are slow to migrate voluntarily.)
- Contract first. Publish the naming scheme (
<service>.service.<dc>.<domain>), health-check requirements, and the rule that application code references names, never tool-specific APIs. This is the lock-in defence: if apps only know DNS names, the registry behind them can be swapped. - Dual registration. Services register in both old and new systems, automatically from the deploy pipeline, so nobody hand-edits two places. Nothing reads the new system yet.
- Shadow reads and diffing. A job resolves every service through both systems and reports mismatches. Cutover of a service is blocked until its diff has been clean for a week.
- Cut over by tier. Start with internal tools, then non-critical services, then tier-1. Each cutover is a config flag per caller, so rollback is a flag flip, not a redeploy.
- Decommission the old path per service once its traffic on the old path is zero for 30 days, with a published date.
Worked example: sizing the timeline
Say 300 services owned by 40 teams. The pilot (weeks 1 to 6) moves 10 services with the platform team doing the work. After that, a paved-road script means a team can move one service in about half a day, and the platform team can support about 15 cutovers per week without degrading review quality.
- Remaining after pilot: 300 - 10 = 290 services.
- At 15 per week: 290 / 15 = 19.3, so 20 weeks of rollout.
- Total: 6 + 20 = 26 weeks, roughly two quarters, plus the 30-day decommission tail.
State it as a plan with that arithmetic, so leadership can see what adding or removing platform capacity does to the date.
3. Addressing the specific concerns
- Operational burden: the platform team owns the registry cluster, its upgrades and on-call, with a published SLO. Application teams own only their health checks. Quantify it: a 5-server consensus cluster per datacenter, which tolerates the loss of 2 servers (a write only needs a majority, 3 of the 5, to agree before it commits; losing any 2 still leaves 3, a majority, so the cluster keeps accepting writes, while losing 3 would leave only 2, which can no longer form one).
- Vendor lock-in: the primary defence is contractual, not technical: apps depend on DNS names and standard health endpoints, not on a client SDK, so the registry behind those names can be swapped without touching application code. As a secondary, deeper defence inside the mesh itself, it also uses open standards (Envoy's xDS configuration APIs, a standard protocol for pushing routing config to sidecar proxies, so the proxies do not have to change even if the control plane does; and SPIFFE identities, a vendor-neutral standard for verifiable service identity, so certificates are not tied to one mesh vendor) so even the mesh's control plane is replaceable. Write down the exit plan now, with its cost.
- "Our custom mesh works fine": respect the investment. Offer to make it a candidate for the opt-in mesh layer if it passes the same criteria, and ask its owners to lead the mesh working group. Most resistance is about losing ownership, not about technology.
4. Gaining buy-in
- Publish the decision as an RFC (request for comments) with the weighted scoring, and let teams challenge the weights before the scores. Changing a weight is a cheap, visible concession.
- Pilot with the loudest sceptic. If their pain is solved, that is the most credible endorsement you can get.
- Make the paved road (the supported, recommended way of doing something, made deliberately the easiest path) cheaper than the old road: generated configs, dashboards and alerts for free on the new platform; the old one gets security fixes only.
- Measure and publish: services migrated, lookup failure rate, time to failover, incidents attributable to discovery, before and after.
- Escalate last. Executive mandate is for the long tail after the default is clearly better, not for the opening move.
Pitfalls
- Choosing by committee vote instead of weighted requirements: the loudest team wins, and the others disengage.
- Big-bang cutover: no rollback, and one bad week kills trust in the platform for years.
- Letting apps import the registry's client library "just for one feature": that is how lock-in re-enters.
- Forgetting non-Kubernetes workloads (VMs, databases, batch jobs): DNS is what keeps them in the plan.
Describe a time when you led an improvement to configuration management or service discovery at your organization. Explain the original problem, your technical and organizational approach, measures of success (for example reduced incident count or faster deployments), and lessons learned including what you automated afterward.
Sample Answer
Direct answer
A strong answer tells one specific story in the STAR shape (Situation, Task, Action, Result) and makes four things explicit because the question asks for all four: the original problem with evidence that it was real, the technical and organizational approach (both, because config and discovery are shared infrastructure and changing them means changing how several teams work), measures of success that you defined before you started and actually tracked, and lessons learned, including what you automated afterward. Interviewers are listening for ownership, for how you brought other teams along, and for whether you measured the outcome honestly.
How to structure it
1. The original problem (about 20% of your time)
Anchor it in evidence, not opinion. Good openings sound like "we had N incidents in two quarters where the root cause was a config change" or "new services took a week to get discoverable because registration was a ticket to another team". Say who was hurt: on-call engineers, customers, developer velocity.
Name the mechanism, briefly, for a technical audience. Typical ones in this area:
- Configuration was edited by hand on hosts or in a shared file with no review, no versioning and no rollback.
- Service endpoints were hard-coded or kept in a static list that went stale after scaling events.
- Secrets or config were baked into container images, so any change needed a rebuild and redeploy.
- Each team had invented its own discovery client, with different caching and retry behaviour.
2. Your approach: technical AND organizational
Technical (what you built or changed):
- Moving config into version control with review (config as code), validated by a schema check in CI (continuous integration, the automated test pipeline that runs on every change).
- Introducing a service registry (a central directory of running instances, such as Consul, etcd or the Kubernetes API) so instances register themselves and clients look them up instead of using static lists.
- Staged rollout of config changes (a small slice of instances first, then everyone) with automatic rollback on error-rate increase.
Organizational (how you got others to adopt it), which is what separates a senior story from a junior one:
- How you got buy-in: data from incident reviews, a one-page proposal, a pilot with one willing team.
- How you handled resistance: a team that said "our current setup works" and what you offered them (migration tooling, doing the first migration with them).
- How you sequenced the migration so nothing broke: dual-read periods (running the old and new config sources side by side and comparing their output before cutting over to the new one), a deprecation date announced in advance.
3. Measures of success
Pick metrics that map directly to the problem you described, and say how they were measured:
- Count of incidents whose root cause was a config or discovery change, from the incident tracker, before vs after.
- Median time from "config change merged" to "live everywhere".
- Time to roll back a bad config change.
- Lead time to make a new service discoverable.
- Adoption: how many of the services were migrated by a given date.
Be honest about attribution. If incident counts dropped while traffic also changed, say so. "I believe the change contributed, and here is why" is stronger than an implausibly exact number.
4. Lessons learned and what you automated afterward
The question explicitly asks what you automated afterward, so plan a concrete answer: typically the manual steps that remained after the first phase. Examples: a CI check that blocks a config change which fails schema validation; a bot that opens a pull request to delete a feature flag (a config switch that turns a piece of code on or off without a new deploy) 30 days after it reached 100%; automatic deregistration of instances that miss health checks; a dashboard or alert on config propagation lag.
Worked example (an illustrative story skeleton)
Situation. At a mid-sized e-commerce company, about 60 services read configuration from files copied onto hosts by a deploy script. In the incident reviews for one half-year, I counted 7 incidents whose root cause was a config change, including a checkout outage where one host kept an old database hostname after a migration.
Task. I was the SRE on the platform team. I proposed owning a fix and got my manager to give me most of a quarter for it.
Action (technical). We moved all config into a Git repository with a JSON Schema per service (a machine-checkable description of what valid config looks like: required fields, types, allowed values), validated in CI. A small agent on each host watched Consul's key-value store (Consul: a central store that holds the current config and notifies any watcher the moment it changes) and rewrote the local file on change, so the application did not need to change how it read config. Changes rolled out to one canary host per service first (a canary: a small trial release to a single host, watched closely, before the change reaches everyone else), with an automatic revert if that host's error rate rose.
Action (organizational). I presented the incident data at the engineering leads' meeting and asked for two pilot teams rather than a mandate. The payments team pushed back because they were mid-audit; I agreed they would go last and paired with their lead on the migration plan. We published a deprecation date for the old deploy script three months out.
Result. Over the following half-year, the same incident-review tagging showed 2 config-caused incidents, both caught at the canary stage and reverted automatically. Rollback went from "find the old file and redeploy", which in the checkout incident took about 40 minutes, to reverting a commit and letting the agent propagate it. By the deprecation date, 55 of the 60 services had moved; the remaining 5 were legacy services we scheduled for retirement instead.
Lessons and automation. Two lessons: I underestimated how much teams valued keeping their existing file format, and adopting the agent-writes-a-file approach is what made migration cheap. Second, schema validation caught far more problems than the canary did, so afterwards I automated schema generation from each service's config class, and added a CI job that flags config keys nobody has read in 90 days so dead config gets deleted.
Every number in the story should be one you can explain if the interviewer asks "how did you count that?". The counts above come from incident-review tags and a migration checklist, which is the kind of source you should be ready to name.
Trade-offs and pitfalls
- All technology, no people. Describing a Consul migration in detail but not how you got 10 teams to adopt it undersells a leadership question.
- Fabricated precision. "Reduced incidents by 83.4%" invites the follow-up "over what baseline?". Round, sourced numbers are more credible.
- No failure. A story where everything went perfectly sounds rehearsed. Include one thing that went wrong and how you adjusted.
- "We" all the way through. Credit the team, but make your own decisions and actions distinguishable.
- Skipping the automation part. The question asks for it explicitly; leaving it out is a missed requirement, not a style choice.
That is every published Service Discovery and Configuration Management question for Engineering Manager so far. Browse the other topics in this category, or practice this one interactively.