Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
Explain how Kubernetes service discovery works for ClusterIP and headless services. Given YAML expressed inline as apiVersion: v1; kind: Service; metadata.name: backend; spec.clusterIP: None; selector.app: backend, describe how DNS records are created, how clients discover endpoints (A vs SRV), and when to use headless services vs ClusterIP.
Sample Answer
Direct answer
A Kubernetes Service is a stable name for a changing set of pods (pod: Kubernetes' smallest deployable unit, one or more containers that run and are scheduled together) picked by a label selector (a rule that matches objects by their key/value labels, such as app=backend). The cluster DNS server (CoreDNS in most clusters) publishes a record for it. For a normal ClusterIP service, the name resolves to one virtual IP, and the node's networking layer spreads connections across the pods behind it. For a headless service (clusterIP: None, like the backend example), there is no virtual IP: the name resolves directly to the IP of every ready pod, so the client sees the individual pods and chooses among them itself. Use ClusterIP by default; use headless when the client must address specific pods (databases, StatefulSets) or balance per request itself (gRPC).
The example
Written out as YAML, the service in the question is:
apiVersion: v1
kind: Service
metadata:
name: backend
spec:
clusterIP: None # this line makes it headless
selector:
app: backend # pods carrying label app=backend are its endpoints
How the records get created
- The EndpointSlice controller watches pods matching
app=backendand writes their IPs (and whether each is ready, i.e. passing its readiness probe, a periodic check the kubelet, the per-node agent that runs and monitors pods, performs against the pod to decide if it should receive traffic) into EndpointSlice objects. - CoreDNS watches Services and EndpointSlices through the Kubernetes API and answers queries from that in-memory view. Nothing is written to a zone file; records change as soon as CoreDNS sees the update.
- Names follow the pattern
<service>.<namespace>.svc.<cluster-domain>, where the cluster domain is usuallycluster.local. Inside the same namespace, a pod can just usebackend, because its resolver search path (a list of domain suffixes the pod's DNS client tries in order, configured automatically by Kubernetes) fills in the rest.
Say backend lives in namespace shop and has three ready pods at 10.244.1.5, 10.244.2.7 and 10.244.3.9. (The SRV row below assumes the service declares a named port; the note right after the table covers what changes for the question's actual YAML, which declares none.)
| Query | ClusterIP service (clusterIP auto-assigned, say 10.96.40.12) | Headless service (clusterIP: None) |
|---|---|---|
A record for backend.shop.svc.cluster.local | one answer: 10.96.40.12 | three answers: 10.244.1.5, 10.244.2.7, 10.244.3.9 (ready pods only) |
| Per-pod A records | none | <hostname>.backend.shop.svc.cluster.local for each pod; pods without an explicit hostname get a system-assigned name (CoreDNS uses the dashed IP, such as 10-244-1-5) |
SRV record for _http._tcp.backend.shop.svc.cluster.local | one record: port 8080, target backend.shop.svc.cluster.local | one record per ready pod, each targeting that pod's own name |
Important detail about the question's YAML: it declares no ports. SRV records are only published for named ports, so this exact service has A records but no SRV records. Adding
ports:
- name: http
port: 8080
protocol: TCP
creates _http._tcp.backend.shop.svc.cluster.local. With N ready pods and M named ports, a headless service publishes N x M SRV records.
How clients discover endpoints: A vs SRV
- A lookup (what almost every client does): resolve the name, get IP addresses, connect to a port the client already knows (from config). With ClusterIP you always get the same single IP, which is simple and cache-safe. With headless you get the current pod list and must pick one yourself.
- SRV lookup: resolve
_http._tcp.backend...and get the port and target name as well, so the port does not have to be configured separately. Useful for clients that support SRV and for services whose port differs by environment. Most HTTP libraries do not look up SRV, so treat it as an option, not the default.
What actually happens on a ClusterIP connection
The ClusterIP is not a real machine. kube-proxy (the per-node component that programs each node's packet-forwarding rules for Services) or an eBPF replacement such as Cilium (eBPF: a way to run small, fast programs inside the Linux kernel to handle packets, which projects like Cilium use instead of kube-proxy's traditional rules) programs rules on every node so that a packet sent to 10.96.40.12:8080 is rewritten to one of the pod IPs. That choice happens per connection. So a client holding one long-lived connection (gRPC, a framework built on HTTP/2 that keeps one TCP connection open and sends many requests over it instead of opening a new connection per request; or a database connection pool) sticks to one pod for every request on that connection, which is the main reason people reach for headless services: the client, not the cluster's per-connection routing, gets to decide where each request goes.
When to use which
| Use ClusterIP when | Use headless when |
|---|---|
| clients are ordinary HTTP services that open many short connections | clients need a stable identity per pod (StatefulSet: a Kubernetes workload type that gives each of its pods a stable, numbered name and stable storage instead of treating them as interchangeable, e.g. db-0, db-1 for replication or Kafka brokers, the server processes that make up a Kafka message-queue cluster, each one individually addressable) |
| you want the cluster to hide pod churn completely | the client does its own per-request load balancing across long-lived connections (gRPC client-side balancing) |
| DNS caching in the client is uncontrolled (a single stable IP never goes stale) | a peer-discovery protocol needs to list all members (clustered caches, databases forming a quorum: a majority of members that must agree before the group accepts a decision, such as electing a leader) |
Pitfalls
- Client DNS caching with headless services. Some runtimes cache DNS answers for a long time (older JVM, the Java Virtual Machine that runs Java programs, settings were the classic case: it cached a resolved name indefinitely by default unless configured otherwise). The pod list then goes stale and traffic keeps hitting pods that are gone. Set a short cache time and re-resolve periodically.
- Not-ready pods are hidden. Headless DNS returns ready pods only. For clusters that must find peers before they are ready (bootstrapping a quorum), set
publishNotReadyAddresses: trueon the service. - No ports means no SRV, as above. It is a common surprise when a client expects SRV.
- Headless does not balance anything. If the client just takes the first IP, all traffic lands on one pod.
Design a dynamic configuration service that supports versioned configurations, atomic rollouts, and quick rollback. Requirements: services can query a specific version or subscribe to change notifications, audit history is preserved, and rollbacks across many instances complete within minutes. Sketch the APIs, storage approach, and how you'd coordinate rollout and rollback safely.
Sample Answer
Direct answer
Make every configuration version an immutable, content-addressed document, and make "what is live" a separate, tiny release pointer per config and per rollout stage. Publishing never edits a live document: it creates a new version. Rolling out means moving the pointer for a growing slice of instances; rolling back means moving the pointer back to a version every instance has already run, which is one small write plus a notification, not a new edit. Instances learn about changes through a watch stream and also poll on a fixed interval, so the rollback time has a hard upper bound even when the push path is broken.
Requirements I am designing to
- Scale assumption: about 20,000 instances across 300 services; configs are typically under 50 KiB; tens of publishes per hour, not thousands.
- Clients can fetch a specific version and subscribe to changes.
- Atomic rollout: an instance sees all of version N or all of version N+1, never a mix of keys from both.
- Audit history: who changed what, when, why, and what was live at any past moment.
- Rollback across all instances within minutes; I will target under 2 minutes worst case.
Terms
- Content-addressed: a version's ID is a hash of its bytes (for example SHA-256). Same bytes, same ID; a client can verify it got exactly what was published.
- Watch / subscription: a long-lived connection (a gRPC stream, using gRPC, a binary remote-procedure-call protocol; SSE which is server-sent events over HTTP; or long polling where the server holds a request open until something changes) over which the server tells clients "the version you should run is now X".
- Bake (bake time): the period you hold a change at a given stage, watching it, before trusting it enough to widen to the next stage.
- CAS (compare-and-swap): "update the pointer to X only if it is currently W". It stops two operators from overwriting each other's rollout.
- Canary: a small first slice of instances that gets a change before everyone else.
Architecture
flowchart LR
Op[Operator or deploy pipeline] -->|publish, release, rollback| API[Config API]
API -->|validate schema| API
API --> DB[(Postgres: versions, releases, audit log)]
API -->|pointer change event| Bus[Change notifier]
Bus --> Edge[Distribution nodes]
Edge -->|watch stream: version id| Agent[Client library in each instance]
Agent -->|GET blob by hash| Edge
Agent -->|health and applied version| Mon[Rollout controller]
Mon -->|advance or roll back| API
Storage
config_versions(config_name, version_id, content_hash, body, schema_version, author, message, created_at): append-only. Nothing is ever updated or deleted in this table (retention aside), which is what makes audit and rollback trivial.releases(config_name, stage, version_id, updated_by, updated_at, revision): one small row per config and rollout stage (for examplecanary,stage-10pct,all). Therevisioncolumn makes CAS a singleUPDATE ... WHERE revision = ?.audit_log: one row per publish, release move and rollback, including the ticket or reason.- Postgres is enough at this write rate. The read load does not go to Postgres: distribution nodes (a small fleet of cache servers placed close to the instances, the same role a CDN edge plays for web assets) cache version blobs by hash (they are immutable, so the cache never needs invalidating) and hold the current pointers in memory.
APIs
| Call | Purpose |
|---|---|
POST /configs/{name}/versions (body, message) | Validate against the schema, store immutable version, return version_id. Does not change what is live. |
GET /configs/{name}/versions/{version_id} | Fetch an exact version (the "query a specific version" requirement). |
POST /configs/{name}/releases (version_id, plan, expected revision) | Start a staged rollout; CAS on the release row. |
POST /configs/{name}/rollback (to_version_id, reason) | Point every stage at a previous version in one transaction. |
WATCH /configs?names=a,b&known=a:v12,b:v4 | Stream "you should now run version X" events. Carries IDs, not bodies. |
GET /configs/{name}/history | Audit trail: versions, releases, rollbacks with who and why. |
How an instance applies a version atomically
- The watch stream says "for your stage,
payments-limitsshould bev13". - The client library downloads the blob by hash (or reads it from its local cache), verifies the hash, and parses and validates it completely before touching live state.
- It swaps one in-memory reference from the old parsed object to the new one. Application code always reads through that reference, so a request sees either v12 or v13 in full.
- It reports
applied=v13(orrejected=v13, reason) to the rollout controller. - It writes the version to local disk as last known good, so a restart while the service is unreachable still boots with valid config.
Coordinating rollout safely
sequenceDiagram
participant Op as Operator
participant API as Config API
participant C as Canary instances
participant R as Rollout controller
participant A as All instances
Op->>API: release v13 (plan 1% then 10% then 100%)
API->>C: watch event: run v13
C->>R: applied v13, error rate, latency
R->>R: bake 10 min, compare with instances still on v12
R->>API: advance to 10%, then 100%
API->>A: watch event: run v13
Note over R,API: on regression, controller calls rollback to v12
- Stage membership is deterministic:
hash(instance_id + config_name) mod 100 < stage_percent. For example, hashing"i-04821" + "payments-limits"lands at bucket 48 out of 100: that instance stays out at the 1% and 10% stages (48 is not below either threshold), and first receives the new version once the rollout reaches 50% (48 < 50). The same instances are always in the 1% for a given config, so a rollback cleanly shrinks the set (dropping every instance whose bucket is no longer below the smaller percentage) instead of reshuffling it. - Health gates compare canary instances against the instances still on the old version over the same window (error rate, latency, crash count, and "rejected config" reports). The comparison against a live control group avoids blaming a config for a traffic spike that hit everyone.
- Automatic rollback fires on a gate failure; humans can also trigger it with one call.
- Only one rollout per config at a time, enforced by the CAS on the release row, so two operators cannot interleave.
- Blast-radius limits (blast radius: how much of the system a single bad change can affect): a config that many services depend on can require a longer bake and a smaller first stage.
Worked example: can rollback finish within minutes?
Assumptions: 20,000 instances, a 50 KiB config, watch streams held by distribution nodes, and a fallback poll every 60 seconds for any instance whose stream is broken.
- Push path. Rollback is one row update plus 20,000 small "run v12" messages. Each instance already has v12 cached locally because it ran v12 before this rollout, so it needs no download at all.
- Worst case if a client did need to download: 20,000 instances × 50 KiB = 1,000,000 KiB = 976.6 MiB, about 0.95 GiB of total transfer (binary units used consistently throughout: 1,024 KiB per MiB, 1,024 MiB per GiB). Spread over roughly 30 seconds by distribution nodes serving immutable, cacheable blobs, that is about 976.6 / 30 ≈ 33 MiB per second fleet-wide, trivial for a small cache fleet to absorb.
- Bound when push fails. The 60-second poll means every healthy instance converges within about 60 s plus apply time, even if every watch stream were down. Those polls cost 20,000 / 60 ≈ 333 requests per second fleet-wide, each a tiny "is my version still current?" check answered from memory.
- Result: typical rollback is seconds (push); the guaranteed bound is roughly one poll interval, comfortably inside "a few minutes". If a tighter bound is needed, shorten the poll to 30 s at double the request rate (about 667 per second).
Trade-offs and pitfalls
| Decision | Chosen | Why | What would change it |
|---|---|---|---|
| Push vs pull | Push notification, pull the body, plus periodic poll | Push gives speed, poll gives a guaranteed bound | Very small fleets can poll only |
| Store | Postgres + edge cache | Low write rate, strong audit queries | Needing multi-region writes might push toward etcd or Spanner-style storage (a globally distributed database that can accept correct writes from multiple regions at once, at the cost of extra latency and operational complexity) |
| Rollback | Pointer move to an old version | Old version is already validated and cached | None: "roll back by editing forward" is how outages get extended |
| Atomicity unit | Whole document per config | Easy to reason about | Cross-config atomicity (two configs that must change together) needs a "bundle" version that pins both |
Pitfalls:
- Rolling back by publishing a "fixed" edit. It is a new, untested version under pressure. Roll back to the known-good ID first; fix forward afterwards.
- Config that can crash the parser. Validate at publish time and in the client before swapping; a client that rejects a version must keep running the old one.
- Everyone applies at the same instant. A change that, for example, resets connection pools can cause a synchronized spike. Add a few seconds of random jitter to the apply step.
- Config and code coupling. If v13 needs a code change that only half the fleet has, the rollout controller must know the minimum client version for a config version and skip instances that are too old.
Describe a schema evolution strategy for configuration data and feature flags that ensures backward and forward compatibility across microservices written in different languages. Include approaches for typed versus untyped configs, deprecation policies, defaulting strategies, validation tooling, and how to handle clients that cannot be upgraded simultaneously.
Sample Answer
Direct answer
Treat configuration and feature flags as a public API with a schema, and evolve that schema with the same rules you would use for a wire protocol (the format two programs agree on to exchange data over a network): only additive changes are allowed in place; every removal, rename or type change goes through expand, migrate, contract, and the contract step is gated on evidence that no running client still reads the old shape. Every field needs a safe default defined in the schema (not scattered through code), a validator runs in CI (continuous integration, the automated checks on every change) and again at publish time, and old clients are protected either by staying tolerant readers or by the config service serving them a shape they understand.
Key terms
- Backward compatible: new code can read old config. (A client upgraded today still works with last month's config documents.)
- Forward compatible: old code can read new config. (A client that has not been upgraded still works after someone publishes a config containing new fields.) In a polyglot fleet (services written in several different languages) you need both, because you never upgrade every service at once.
- Typed config: the structure is declared in a schema language (Protocol Buffers/Protobuf, Google's compact binary encoding format with a schema; Avro, a similar binary format popular in the Kafka/Hadoop ecosystem; JSON Schema, a way to declare rules for plain JSON; CUE, a newer schema and validation language) and code is generated or validated against it.
- Untyped config: free-form key/value strings or JSON blobs that each service interprets however it likes.
- Tolerant reader: a client that ignores fields it does not recognise and applies a default for fields that are missing, instead of failing.
- Expand/contract: a change done in two phases. Expand adds the new shape alongside the old; clients migrate; contract removes the old shape once nobody reads it.
1. Typed vs untyped: pick typed, with an escape hatch
| Typed (Protobuf / JSON Schema / CUE) | Untyped (string key/value, raw JSON) | |
|---|---|---|
| Catching mistakes | At publish time, before any service sees it | At runtime, in whichever service parses it first |
| Cross-language | Generated code gives every language the same field names, types and defaults | Each language re-implements parsing; subtle divergence (is "1" true?) |
| Evolution rules | Enforceable by tooling (breaking-change detectors) | Enforced only by review and memory |
| Flexibility | Needs a schema change to add a field | Anything goes |
Recommendation: typed by default. Use Protobuf (Protocol Buffers: each field is declared with a small integer field number as well as a name, and that number, not the name, is what actually goes out on the wire) when services already speak it. Its binary encoding has well-known compatibility rules: a reader identifies fields by number, skips any field number it does not recognise instead of failing, and proto3 (the current version of the Protobuf language) runtimes preserve those unknown fields when they re-serialise a message, so a service that only forwards a config through does not silently drop fields it does not understand. Otherwise use JSON Schema when configs are hand-edited JSON/YAML. Keep a small, explicitly untyped extensions map for team-local experiments, with the rule that anything read by more than one service must graduate into the schema.
2. The compatibility rules (write them down, enforce them in CI)
Safe in place:
- Add an optional field with a default.
- Add an enum value, provided every reader has an explicit "unknown value" branch.
- Widen a numeric range that readers already validate against (for example, if
max_retrieswas declared 0 to 5 and every reader already validates against that range, raising the allowed maximum to 10 is safe: an old reader that never sees a value above 5 behaves exactly as before, and a new reader that receives 8 accepts it correctly; narrowing the range is the unsafe direction, since an old reader might already rely on a value the new, smaller range would reject).
Never in place (always expand/contract):
- Removing or renaming a field. In Protobuf, also never reuse a removed field's number or name: mark them
reserved(a declaration that tells the compiler "never assign this number or name again"). Otherwise, because Protobuf identifies fields by number rather than by name, an old client that still expects field 3 to meanretrieswill read whatever new field 3 means today as if it were stillretries, silently misinterpreting it. - Changing a type (
int32milliseconds to a duration string, a bool to an enum). - Changing the meaning or unit of a field while keeping its name. This is the most dangerous because every tool says it is compatible.
- Making an optional field required.
Tooling: buf breaking (a command from the buf CLI, a common Protobuf tool) checks Protobuf schemas against the previous version and fails the build on incompatible changes. For JSON Schema, a CI job validates every stored config document against both the new schema and the previous one. The config service validates again at publish time, so nobody can bypass CI by editing the store directly.
3. Defaulting strategy
- Defaults live in the schema, and generated code in every language reads them from there, so Go and Python agree on what "missing" means.
- Distinguish "absent" from "zero". In proto3, a plain
int32 retries = 3;cannot tell "not set" from "set to 0": both look identical on the wire. Declare such fieldsoptional(which gives presence tracking: the generated code adds a separate "was this field actually set" flag alongside the value) or wrap them (use a message type like Protobuf's well-knownInt32Valueinstead of a bareint32, so "unset" is a distinct, checkable state), because "0 retries" and "use the default" are different decisions. - A default must be safe, not just valid. The default for a timeout should be the value that keeps the system working if the field disappears, not
0. - Unknown enum values fall back explicitly, for example to the most conservative behaviour, and increment a metric so you know it happened.
4. Feature flags: same discipline, shorter life
Feature flags are configuration that changes behaviour per request or per user.
- Type every flag (boolean, string, number, structured JSON variant) and require a code-side default at every evaluation site. The OpenFeature standard (a vendor-neutral flag evaluation API from the CNCF, the Cloud Native Computing Foundation) builds this in: calls look like
getBooleanValue("new-checkout", false), so an unreachable flag service or a deleted flag returns the default instead of throwing. - Never change a flag's type or meaning under the same key. Create a new key and retire the old one.
- Deprecation has an owner and an expiry date set at creation. Release flags (temporary switches for a rollout) should be removed within weeks of reaching 100%; permanent operational flags (kill switches) are documented as permanent.
5. Deprecation policy
- Mark the field deprecated in the schema (Protobuf has
[deprecated = true]; generated code then warns at compile time in most languages). - Announce a removal date, with a support window long enough for the slowest-releasing consumer. A common rule is "two release cycles of the slowest consumer or 90 days, whichever is longer".
- Measure readers before removing. The client library records which config fields each service actually reads and reports it (field name, service name, library version). Removal is allowed only when that report shows zero readers for the window.
- Remove, and reserve the name/number.
6. Clients that cannot be upgraded at the same time
This is the real constraint, and there are three tools for it, in order of preference:
- Tolerant readers everywhere. Every client library ignores unknown fields and applies schema defaults. This alone makes all additive changes forward compatible.
- Expand/contract with dual-write. During the expand window, publishers write both the old and new fields; old clients read the old one, new clients prefer the new one.
- Version negotiation at the config service. Each client sends the schema version it was built with (
schema_version=2) when it fetches config. The service stores config in the newest shape and down-converts on read for older clients, so an embedded device or a vendor-supplied service that can never be rebuilt still gets a shape it understands. This costs converter code in the service, so reserve it for consumers that genuinely cannot move.
Worked example: renaming timeout_ms (an integer) to timeout (a duration string)
Three published config documents and two client generations. The old client was built when only timeout_ms existed and defaults it to 1000 if missing; the new client prefers timeout and falls back to timeout_ms. Version 2 also adds an unrelated new field, retry_budget (the fraction of a service's own request volume it is allowed to spend on retries, so retries cannot multiply into a bigger overload than the original traffic), to show that an old client ignoring a field it does not recognise matters just as much as how it handles a renamed one.
import json
# Config documents as the config service would store them.
v1 = {"schema_version": 1, "timeout_ms": 2500}
v2 = {"schema_version": 2, "timeout_ms": 2500, "timeout": "2.5s", "retry_budget": 0.1} # expand
v3 = {"schema_version": 3, "timeout": "2.5s", "retry_budget": 0.1} # contract
def parse_duration_ms(text):
if text.endswith("ms"):
return int(float(text[:-2]))
if text.endswith("s"):
return int(float(text[:-1]) * 1000)
raise ValueError(f"bad duration {text!r}")
def old_reader(doc):
"""A client built against version 1: knows only timeout_ms."""
return {"timeout_ms": doc.get("timeout_ms", 1000)} # default if missing
def new_reader(doc):
"""A client built against version 2: prefers the new field, falls back to the old."""
if "timeout" in doc:
timeout_ms = parse_duration_ms(doc["timeout"])
else:
timeout_ms = doc.get("timeout_ms", 1000)
return {"timeout_ms": timeout_ms, "retry_budget": doc.get("retry_budget", 0.2)}
for name, doc in [("v1", v1), ("v2", v2), ("v3", v3)]:
print(name, "old client ->", old_reader(doc), "| new client ->", new_reader(doc))
Output:
v1 old client -> {'timeout_ms': 2500} | new client -> {'timeout_ms': 2500, 'retry_budget': 0.2}
v2 old client -> {'timeout_ms': 2500} | new client -> {'timeout_ms': 2500, 'retry_budget': 0.1}
v3 old client -> {'timeout_ms': 1000} | new client -> {'timeout_ms': 2500, 'retry_budget': 0.1}
Read the three lines as the three phases:
- v1 (before): both clients agree on 2500 ms.
- v2 (expand): both fields are written; both clients still agree on 2500 ms, and the new client also picks up the new
retry_budgetfield while the old client simply ignores it (forward compatibility). - v3 (contract, done too early): the old field is gone. The old client does not crash. It silently falls back to its default of 1000 ms, so requests that need 2.5 s start timing out. That silent wrong default is the typical way schema evolution fails in production, and it is why the contract step is gated on measured reader counts (section 5), not on a calendar date alone.
Trade-offs and pitfalls
- "Compatible" by the tool is not compatible by meaning. Changing
timeoutfrom seconds to milliseconds passes every schema check. Put units in field names or use typed durations. - Required fields are forever. Adding one breaks every old publisher; removing one breaks every old reader. Prefer optional with a safe default.
- Down-conversion is a long-term cost. Every schema version you promise to serve is converter code to test. Cap it (for example, the current version plus two older ones) and publish that policy.
- Validation at publish time is the last line of defence, because a bad config reaches every instance within seconds, while a bad code deploy usually rolls out gradually.
- Flag debt is schema debt. Hundreds of stale flags are hundreds of untested code paths. The expiry date at creation is what keeps them from accumulating.
Describe strategies to test and debug behavior driven by externalized configuration: missing keys, incorrect types, live reload, secrets rotation, and mixed-version deployments. Include automated tests, canary rollout strategies, safe defaults, and how to handle secrets securely in CI while enabling meaningful tests.
Sample Answer
Direct answer
Treat configuration as code with a schema: every key has a declared type, range and safe default; every change is validated in CI (continuous integration) against the schema of every binary version currently running; the application validates again on load and keeps the last good config if a live reload is bad; and every change rolls out as a canary with automatic comparison and rollback. For secrets, CI tests use short-lived, scoped test credentials obtained through identity federation (the CI job proves who it is with a token it already holds, and trades that for a short-lived credential, instead of a long-lived password stored in CI settings), and rotation is tested by actually rotating in a staging environment, so real production secrets never enter CI.
Terms first
- Externalized configuration: settings that live outside the binary (environment variables, files, a config service such as Consul or AWS AppConfig) so behaviour changes without a rebuild.
- Live reload: the running process picks up a config change without restarting.
- Secrets rotation: replacing a credential (database password, API key) with a new one on a schedule or after exposure.
- Mixed-version deployment: during a rolling deploy, old and new binaries run side by side and read the same config.
- Canary: apply a change to a small slice first, compare it with the rest, then widen or roll back.
Strategy per failure mode
| Failure mode | Prevent (automated test) | Contain (runtime) | Debug (when it happens) |
|---|---|---|---|
| Missing keys | schema test: every key the code reads is declared with a default; a unit test loads an empty config and asserts the service starts with safe behaviour | defaults are conservative: new features off, timeouts bounded, retries low | log the effective config (with defaults applied, secrets redacted) at startup and on each reload |
| Incorrect types | strict parsing tests: "1500" for an int and "false" for a bool are rejected, not coerced ("false" coerced to bool is True in many languages) | validate before swap; reject and keep last good | expose "config rejected" as a counter with the reason |
| Live reload | test that a reload is atomic (readers see old or new, never half), and that a bad reload leaves the old config active | build a new immutable config object, validate, then swap one reference | expose the applied config version per instance |
| Secrets rotation | integration test in staging that rotates and asserts zero failed requests | overlap window: new and old credential both valid, clients re-read on auth failure | alert on auth failures per credential version, never logging the secret |
| Mixed versions | CI validates the new config against the schema of every deployed version | expand, then contract (add the new capability everywhere before anything uses it, remove the old one only once nothing still uses it): ship code that understands a key (default off) before anyone sets the key | break down errors by binary version and config version together |
Worked example: validation plus last-known-good reload
A typed config with safe defaults, strict parsing, range checks, and a holder that only swaps after validation. Three Python-specific mechanics matter here even if you do not read Python day to day: frozen=True makes the object immutable once built, so nothing can quietly mutate a config that is already live; __post_init__ is a hook the dataclass runs right after construction, used here to enforce the range checks; and fields(Config) asks the class itself which keys it declares, so the parser's allow-list always matches the schema instead of drifting out of sync with it.
from dataclasses import dataclass, fields
@dataclass(frozen=True)
class Config:
timeout_ms: int = 2000 # safe default: bounded, not infinite
max_retries: int = 2
new_pricing: bool = False # new behaviour defaults OFF
def __post_init__(self):
if not 50 <= self.timeout_ms <= 30_000:
raise ValueError(f"timeout_ms out of range: {self.timeout_ms}")
if not 0 <= self.max_retries <= 5:
raise ValueError(f"max_retries out of range: {self.max_retries}")
def parse(raw: dict) -> Config:
known = {f.name: f.type for f in fields(Config)}
unknown = set(raw) - set(known)
if unknown: # typo'd key: fail loudly, don't ignore
raise ValueError(f"unknown keys: {sorted(unknown)}")
out = {}
for key, typ in known.items():
if key not in raw:
continue # missing: dataclass default applies
val = raw[key]
if typ is bool and not isinstance(val, bool):
raise TypeError(f"{key}: expected bool, got {val!r}")
if typ is int and (isinstance(val, bool) or not isinstance(val, int)):
raise TypeError(f"{key}: expected int, got {val!r}") # bool is a subclass of int in Python, so it is excluded explicitly or True/False would silently pass as 1/0
out[key] = val
return Config(**out)
class Holder:
"""Live reload: validate first, then swap the whole object atomically."""
def __init__(self, cfg): self.current = cfg
def reload(self, raw):
try:
self.current = parse(raw)
return "applied"
except (TypeError, ValueError) as e:
return f"rejected, kept last good ({e})"
cases = {
"some keys missing": {"timeout_ms": 1500}, # max_retries and new_pricing are simply absent
"wrong type": {"timeout_ms": "1500"},
"string bool": {"new_pricing": "false"},
"typo'd key": {"timout_ms": 1500},
"out of range": {"timeout_ms": 0},
}
h = Holder(parse({}))
for name, raw in cases.items():
print(f"{name:<18} -> {h.reload(raw)}; now {h.current}")
Output (Python 3, run as shown):
some keys missing -> applied; now Config(timeout_ms=1500, max_retries=2, new_pricing=False)
wrong type -> rejected, kept last good (timeout_ms: expected int, got '1500'); now Config(timeout_ms=1500, max_retries=2, new_pricing=False)
string bool -> rejected, kept last good (new_pricing: expected bool, got 'false'); now Config(timeout_ms=1500, max_retries=2, new_pricing=False)
typo'd key -> rejected, kept last good (unknown keys: ['timout_ms']); now Config(timeout_ms=1500, max_retries=2, new_pricing=False)
out of range -> rejected, kept last good (timeout_ms out of range: 0); now Config(timeout_ms=1500, max_retries=2, new_pricing=False)
The first case supplies only timeout_ms; max_retries and new_pricing are simply absent from that dict, so the dataclass fills in the defaults it declared for them, and the change applies. Every later case is rejected and the service keeps running on the last good values. In Python, assigning self.current is a single reference swap, so a reader sees either the old or the new object, never a mix.
Why strict unknown-key rejection is safe with mixed versions: it would break if someone set a new key while old binaries (which do not know it) were still running. The expand-then-contract rule prevents that ordering: the binary that knows new_pricing (defaulting to off) is fully deployed first, and CI refuses a config that uses a key unknown to any currently deployed version. Removing a key runs the same way in reverse: stop reading it, deploy everywhere, then delete it.
Canary rollout for config changes
Config changes cause as many incidents as code changes, so they get the same pipeline: apply to 1% of instances (or one availability zone, an isolated location within a cloud region), compare error rate and latency with the rest for a fixed bake time (the soak period you watch before trusting the change and moving on, for example 10 minutes), then 10%, 50%, 100%. The rollout controller halts and reverts automatically when the canary is clearly worse. The revert is "re-publish the previous version", which must always be one command.
Secrets in CI, securely and still meaningfully
- No production secrets in CI, ever. Tests hit ephemeral dependencies (a database in a container) whose credentials are generated for that run and die with it.
- When CI must call a real service (a cloud API in an integration test), use workload identity federation: the CI system presents a signed OIDC (OpenID Connect) token proving "this is job X in repo Y on branch main", and the cloud's identity service exchanges it for credentials valid for minutes, scoped to a test account. No long-lived key is stored in CI settings.
- Masking is a backstop, not a control. Log masking misses a secret that is base64-encoded (a reversible text encoding, not encryption: anyone can decode it back) or printed character by character, so assume anything a test can read can leak, and scope accordingly.
- Test rotation for real: a staging job rotates the database password on a schedule while a load generator runs, and asserts zero failed requests. The mechanics it proves: the secret store holds two valid versions during an overlap window, clients reload on a timer and retry once with a fresh fetch on an authentication failure, and the old version is revoked only after metrics show nothing is still using it.
Debugging checklist when behaviour looks config-driven
- Which config version is applied on the misbehaving instances (not which one was published)?
- Do errors split by config version or binary version? That separates "bad value" from "incompatible code".
- What does the effective config dump show after defaults were applied?
- Were any reloads rejected recently (the counter), meaning part of the fleet is stuck on an older version?
Trade-offs and pitfalls
- Fail-closed vs fail-open on startup. Fail open means keep serving on the values you already trust; fail closed means refuse to proceed rather than run on data you cannot trust. On reload, fail open: keep last good. On a cold start with no valid config, fail closed: refuse to start rather than run on half-validated values; otherwise a bad config silently becomes production behaviour on every new instance.
- Coercion is the silent killer: permissive parsers that turn
"false"intoTrueor"30s"into30milliseconds pass every happy-path test. - Defaults drift. A default in code and a value in the config store can disagree without anyone noticing; the startup dump of effective config is how you see it.
- Canary bake time vs speed: a 10-minute bake per stage makes a four-stage rollout take 40 minutes; keep an audited fast path for emergency reverts only.
Implement in Python a thread-safe ServiceResolver class that resolves a service name to IP addresses using socket.getaddrinfo and caches results per-service with a configurable TTL. Provide methods: resolve(service_name) -> list of IPs, and invalidate(service_name). The cache should refresh after TTL expiry and avoid duplicate simultaneous DNS queries when multiple threads request the same service concurrently.
Sample Answer
Direct answer
Keep one lock that protects two small dictionaries: a cache of service -> (IPs, expiry time) and an in-flight table of service -> pending lookup. A caller that finds a fresh cache entry returns it immediately. A caller that finds nothing fresh either becomes the leader for that service (registers a pending lookup, releases the lock, calls getaddrinfo) or, if a lookup is already pending, becomes a follower and waits on it. This pattern is called single-flight: however many threads ask for the same name at once, only one DNS query goes out. The DNS call must happen outside the lock, otherwise one slow lookup would block every other service's cache hits.
Approach
- Why a TTL cache at all?
socket.getaddrinfoasks the operating system resolver, which may go over the network to a DNS server. Doing that on every request adds latency and load. Caching for a TTL (time-to-live), a fixed number of seconds after which the answer is treated as expired, trades a little freshness for a lot of speed.getaddrinfodoes not return the DNS record's own TTL, so the class takes a configured one. - Monotonic clock. Expiry uses
time.monotonic(), which never jumps backwards when the wall clock is corrected by NTP (the network time protocol). Usingtime.time()could make entries live far too long or expire instantly after a clock step. - Leader/follower with
threading.Event. The leader stores the result (or the exception) on a small_InFlightobject and sets its event. Followers wake up and read the same result, so all concurrent callers see one consistent answer. - Errors are propagated, never cached. A failed lookup is raised to every waiting caller, and the next call retries. (Caching negative results briefly is a deliberate extension, discussed below.)
invalidate()also detaches an in-flight lookup. If someone invalidates while a lookup is running, that lookup's answer might predate the change that triggered the invalidation, so the leader checks that its flight is still the registered one before caching.- Injectable
lookupandclock. This makes the concurrency and TTL behaviour testable without a network and without sleeping through real TTLs.
Code
import socket
import threading
import time
class _Entry:
# __slots__ is optional: it just tells Python this class only ever has these two
# attributes, so instances skip the usual per-object __dict__ and use less memory.
# Correctness does not depend on it; it is worth doing here because many of these
# can exist at once, one per cached service.
__slots__ = ("ips", "expires_at")
def __init__(self, ips, expires_at):
self.ips = ips
self.expires_at = expires_at
class _InFlight:
"""One pending lookup that concurrent callers wait on (single-flight)."""
__slots__ = ("done", "ips", "error")
def __init__(self):
self.done = threading.Event()
self.ips = None
self.error = None
class ServiceResolver:
def __init__(self, ttl_seconds=30.0, port=None,
lookup=socket.getaddrinfo, clock=time.monotonic):
if ttl_seconds <= 0:
raise ValueError("ttl_seconds must be positive")
self._ttl = ttl_seconds
self._port = port
self._lookup = lookup # injectable for tests
self._clock = clock # monotonic: immune to wall-clock jumps
self._lock = threading.Lock() # guards _cache and _inflight only
self._cache = {} # service_name -> _Entry
self._inflight = {} # service_name -> _InFlight
def resolve(self, service_name):
with self._lock:
entry = self._cache.get(service_name)
if entry is not None and self._clock() < entry.expires_at:
return list(entry.ips) # fresh hit, copy out
flight = self._inflight.get(service_name)
leader = flight is None
if leader:
flight = _InFlight()
self._inflight[service_name] = flight
if not leader: # follower: wait for the leader
flight.done.wait()
if flight.error is not None:
raise flight.error
return list(flight.ips)
try: # leader: DNS call OUTSIDE the lock
# type=socket.SOCK_STREAM asks getaddrinfo for TCP-capable results only (as
# opposed to UDP), which is also what keeps it from returning the same IP
# twice for two socket types we do not care about.
infos = self._lookup(service_name, self._port,
type=socket.SOCK_STREAM)
# getaddrinfo returns a list of 5-tuples: (family, socktype, proto, canonname,
# sockaddr). sockaddr is itself a tuple, (ip, port) for IPv4 or (ip, port,
# flowinfo, scopeid) for IPv6, so info[4] is the sockaddr and info[4][0] is
# always the IP string. The set comprehension collects and dedupes them.
ips = sorted({info[4][0] for info in infos})
if not ips:
# gaierror ("getaddrinfo error") is the same exception type getaddrinfo
# itself raises on a real failure, so every caller handles an empty
# result the same way it would handle a genuine DNS failure.
raise socket.gaierror(f"no addresses for {service_name}")
flight.ips = ips
with self._lock:
# Only cache if nobody called invalidate() while we were resolving.
if self._inflight.get(service_name) is flight:
self._cache[service_name] = _Entry(ips, self._clock() + self._ttl)
except Exception as exc:
flight.error = exc # errors are never cached
finally:
with self._lock:
if self._inflight.get(service_name) is flight:
del self._inflight[service_name]
flight.done.set()
if flight.error is not None:
raise flight.error
return list(flight.ips)
def invalidate(self, service_name):
with self._lock:
self._cache.pop(service_name, None)
# Detach any lookup already running: its answer may predate the
# change that triggered this invalidation, so it must not be cached.
self._inflight.pop(service_name, None)
# ---------------- demo driver: deterministic, no network needed ----------------
if __name__ == "__main__":
calls = {"n": 0}
calls_lock = threading.Lock()
gate = threading.Event()
def fake_lookup(host, port, type=0):
with calls_lock:
calls["n"] += 1
gate.wait() # hold the lookup open
return [(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("10.0.0.7", 0)),
(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("10.0.0.5", 0)),
(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("10.0.0.5", 0))]
now = {"t": 1000.0}
r = ServiceResolver(ttl_seconds=30, lookup=fake_lookup, clock=lambda: now["t"])
results = []
def worker():
results.append(r.resolve("payments.internal"))
threads = [threading.Thread(target=worker) for _ in range(20)]
for t in threads:
t.start()
while calls["n"] == 0: # wait until the leader is inside the lookup
time.sleep(0.001)
time.sleep(0.05) # let the other 19 queue up as followers
gate.set()
for t in threads:
t.join()
print("threads:", len(results), "| lookups:", calls["n"], "| answer:", results[0])
print("all threads got the same list:", all(x == results[0] for x in results))
now["t"] += 10
r.resolve("payments.internal")
print("cached hit at t+10s, lookups still:", calls["n"])
now["t"] += 21
r.resolve("payments.internal")
print("after TTL expiry (t+31s), lookups:", calls["n"])
r.invalidate("payments.internal")
r.resolve("payments.internal")
print("after invalidate(), lookups:", calls["n"])
# invalidate() while a lookup is in flight: that answer must not be cached
slow_gate = threading.Event()
def slow_lookup(host, port, type=0):
slow_gate.wait()
return [(socket.AF_INET, socket.SOCK_STREAM, 6, "", ("10.0.0.9", 0))]
r2 = ServiceResolver(ttl_seconds=30, lookup=slow_lookup, clock=lambda: now["t"])
t = threading.Thread(target=r2.resolve, args=("orders.internal",))
t.start()
while "orders.internal" not in r2._inflight:
time.sleep(0.001)
r2.invalidate("orders.internal")
slow_gate.set()
t.join()
print("invalidate during in-flight lookup, cached afterwards:", "orders.internal" in r2._cache)
def failing_lookup(host, port, type=0):
raise socket.gaierror("Name or service not known")
bad = ServiceResolver(ttl_seconds=5, lookup=failing_lookup)
try:
bad.resolve("nope.internal")
except socket.gaierror as e:
print("failure propagates, not cached:", e, "| cache size:", len(bad._cache))
Output from running the file as shown:
threads: 20 | lookups: 1 | answer: ['10.0.0.5', '10.0.0.7']
all threads got the same list: True
cached hit at t+10s, lookups still: 1
after TTL expiry (t+31s), lookups: 2
after invalidate(), lookups: 3
invalidate during in-flight lookup, cached afterwards: False
failure propagates, not cached: Name or service not known | cache size: 0
The demo pins everything that matters: the fake lookup returns fixed addresses, the clock is a variable the driver advances, and the 20 threads are held at a gate so that they genuinely overlap. The output shows one DNS call for 20 concurrent callers, deduplicated and sorted IPs, a cache hit inside the TTL, a refresh after it, a forced refresh after invalidate(), an invalidation racing an in-flight lookup, and a failure that is raised but not cached.
Key points
- Lock scope is tiny. The lock is held only to read or change the two dictionaries, never across I/O. Different services never wait on each other's DNS.
- Copies out.
resolve()returnslist(...)so a caller that mutates its result cannot corrupt the cache. - Deduplication.
getaddrinfocan return one entry per (family, socket type, protocol) combination, so the same IP often appears more than once; the set removes duplicates. Asking forSOCK_STREAMalso cuts duplicates at the source. - IPv4 and IPv6. Both families are returned by default. Pass
family=socket.AF_INETin the lookup call if callers cannot use IPv6.
Complexity
- Cache hit: O(1) dictionary lookup under the lock, plus O(k) to copy k addresses.
- Miss: one
getaddrinfocall per service per TTL window, regardless of how many threads ask, plus O(m log m) to sort m results. - Memory: O(S × k) for S cached services with k addresses each, plus O(S) for pending lookups.
Edge cases
| Case | Behaviour |
|---|---|
| Many threads ask for the same service at once | One lookup; the rest wait on the event and share the result. |
| Many threads ask for different services | Lookups run in parallel because the lock is released during I/O. |
Lookup raises (gaierror, timeout) | Every waiter gets the same exception; nothing is cached; the next call retries. |
| Lookup returns an empty list | Treated as an error rather than caching "no instances". |
invalidate() during an in-flight lookup | The running lookup's answer is not cached; the next caller starts a fresh lookup. |
| TTL of zero or negative | Rejected in the constructor. |
| Wall-clock jumps | Irrelevant, because expiry uses the monotonic clock. |
Trade-offs and extensions
- Stale-while-revalidate. At expiry, every caller for that service waits for the refresh. For hot services you can instead return the expired value immediately and refresh in the background, so latency never spikes at the TTL boundary. The cost is serving data slightly older than the TTL.
- Serve-stale-on-error. If the DNS server is down, returning the last known addresses (and counting that you did) keeps traffic flowing. The risk is routing to instances that have gone away, so cap how stale you will go.
- Negative caching. Caching a failure for a few seconds protects the DNS server from a retry storm when a name genuinely does not exist.
- Timeouts.
getaddrinfohas no timeout parameter. A hung resolver would hold followers forever, so production code often runs the lookup in a worker pool and has followers usedone.wait(timeout). - Jittered expiry. If thousands of processes cache the same name with the same TTL, they refresh in lockstep. Adding a small random jitter to the expiry spreads that load.
Unlock Full Question Bank
Get access to all 15 Service Discovery and Configuration Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.