Service Discovery and Configuration Management Questions
Letting services find and configure each other at runtime: service registries, client-side versus server-side discovery, DNS-based discovery, dynamic configuration, feature flags, and secrets distribution. Covers how services stay wired together as instances come and go, how config changes propagate safely, and how to monitor and diagnose the outages that stale endpoints or bad config pushes cause. The connective plumbing of a microservices deployment.
Explain DNS-based service discovery and how DNS caching and TTLs affect failover and load distribution. If you had to choose between a 60-second and a 300-second TTL for a service record, which would you pick and why, and what else would you need to account for to make that choice actually hold up in practice?
Sample Answer
Direct answer
DNS-based service discovery means clients find a service by resolving a name such as orders.internal.example.com to the IP addresses of its healthy instances. Because every resolver and client caches the answer for the record's TTL (time-to-live), the TTL sets how long clients may keep sending traffic to an address you have already removed: it is a floor on failover time and a limit on how quickly load shifts. For a service record that must fail over, I would pick 60 seconds. The extra query load is small and a 300-second TTL means up to five minutes of errors after an instance or region dies. But a 60-second TTL only holds if clients actually honour it, so the rest of the job is making sure they do.
How DNS-based discovery works
- Instances register (directly, or via a health-checking DNS provider such as Consul's DNS interface, Kubernetes' CoreDNS, or a cloud DNS service with health checks).
- A client asks its local stub resolver (the small resolver library in the OS or language runtime), which asks a recursive resolver (the caching server, e.g. the node-local cache or the VPC (Virtual Private Cloud) resolver), which asks the authoritative server (the one that actually owns the zone, meaning the portion of the DNS namespace, such as
internal.example.com, that this server holds the real records for). - The answer, say three A records (address records, each mapping the name to one IPv4 address), is cached at every layer for the TTL. The client picks one address, usually the first, and connects.
Nothing tells a client that an address went bad: it only learns on its next lookup after the cached copy expires.
How caching and TTLs affect failover
Worst-case time for all clients to stop using a dead address is roughly:
failover time ≈ health-check detection time + TTL (+ any extra caching the client adds)
With a health check that marks an instance down after 3 failures at 10-second intervals (30 seconds):
| TTL | Worst-case failover | Lookups per second from 2,000 caching client hosts (one lookup per TTL each) |
|---|---|---|
| 60 s | 30 + 60 = 90 s | 2,000 / 60 ≈ 33 |
| 300 s | 30 + 300 = 330 s | 2,000 / 300 ≈ 6.7 |
The 60-second TTL costs about 26 more queries per second, which any resolver absorbs trivially, and cuts the worst-case failover from 5.5 minutes to 1.5 minutes. Query volume scales linearly with 1/TTL, so going lower (5 seconds) is 400 queries per second for the same fleet and starts to matter for resolvers and for paid public DNS.
How caching affects load distribution
DNS balances resolvers, not requests. If one large recursive resolver (a shared NAT, meaning a network-address-translation gateway that many separate devices sit behind and also route their DNS queries through, or a corporate DNS server) serves 500 clients, all 500 may receive the same cached answer in the same order and pile onto the first IP. Longer TTLs make this worse, because a skewed answer persists longer; round-robin record ordering, returning a random subset per response, or client-side random selection among all returned addresses spreads it out. Adding capacity has the same lag as removing it: new instances receive traffic only as caches expire.
What else must be true for the 60-second choice to hold
- Clients must honour the TTL. Language runtimes add their own caches. The JVM's positive cache (a positive cache remembers successful lookups; the negative cache below remembers failures) is governed by the
networkaddress.cache.ttlsecurity property, whose default is implementation-specific; set it explicitly (for example to 60) instead of trusting it. Some resolvers also enforce a minimum TTL that silently raises yours. - Clients must re-resolve at all. A connection pool or an HTTP/2 keep-alive connection (reusing one already-open connection for many requests instead of reconnecting for each one) opened to an IP will keep using that IP for hours, regardless of DNS TTL. Cap connection lifetime (e.g. recycle connections every few minutes) and re-resolve on reconnect, or failover never happens for long-lived clients.
- DNS must only return healthy addresses. The TTL is pointless if the authoritative server keeps handing out the dead IP; the records must be driven by health checks.
- Negative caching. If a lookup returns "no such name" (for example during a brief deregister-everything incident), resolvers cache that negative answer for a time taken from the zone's SOA (start of authority) record, per RFC 2308, and Java's
networkaddress.cache.negative.ttldefaults to 10 seconds. A long negative TTL turns a 5-second blip into minutes of "host not found". Keep it short for service zones. - Clients must retry on another address. Even with perfect TTLs, requests fail during the detection window. Returning several records and having clients try the next address on connection failure turns an outage into a latency bump.
- Lower TTLs before planned changes. A TTL only shortens after the old one expires, so for a planned migration drop the TTL a full old-TTL period ahead of time.
When I would pick 300 seconds instead
For records that rarely change and front something that fails over by other means (a load balancer's stable virtual IP, an address the load balancer itself keeps fixed while swapping which physical instance answers behind it, or an anycast address, the same IP announced from multiple physical locations so network routing sends each client to the nearest live one and simply stops routing to a failed location, no DNS change needed), 300 seconds or more is fine: the address itself never dies, so failover does not depend on DNS, and a longer TTL cuts query volume and shields you from brief DNS outages.
Trade-offs and pitfalls
- A very low TTL (under 10 seconds) makes DNS itself a hard dependency: if the resolver hiccups, every client fails its lookups at once.
- DNS offers no per-request load awareness and no health-based retries; for east-west traffic (traffic between your own services inside the system, as opposed to north-south traffic coming in from outside clients) between microservices that need fast failover, client-side load balancing (the calling service itself holds the instance list and picks one per request, no DNS or proxy hop involved) or a service mesh (a proxy next to each service that receives endpoint updates by streaming) reacts in seconds instead of a TTL period.
- Assuming "TTL 60 means failover in 60 seconds" ignores detection time, runtime caches and pooled connections, which is where most real incidents come from.
Show how DNS SRV and TXT records can be used together for discovery and lightweight configuration. Provide an example SRV record and associated TXT metadata for a service named db.example, and explain how a client would parse priority, weight, port, and configuration metadata from these records.
Sample Answer
Direct answer
An SRV record answers "where is this service?": for a name like _postgres._tcp.db.example it returns one or more (priority, weight, port, target host) tuples. A TXT record on the same name carries small key=value strings the client can read as lightweight configuration (TLS mode, pool size, a timeout). The client looks up both, reads the TXT settings, then orders the SRV targets: lowest priority number first, and within the same priority it picks randomly in proportion to weight, falling back to the next priority tier only when every host in the current tier fails.
The pieces, in plain language
- DNS (Domain Name System) is the phone book every machine already has. An
Arecord maps a name to an IP address, but it cannot say which port to use or which server to prefer. - SRV record (RFC 2782, "service location") adds exactly that. Its owner name has the shape
_<service>._<protocol>.<domain>, and its data is four fields:
| Field | Meaning | Rule the client applies |
|---|---|---|
| priority | preference tier | lower number is tried first; higher tiers are backups |
| weight | share of traffic inside one tier | pick randomly, proportional to weight; weight 0 still gets a small, non-zero chance among records that share its priority, it does not mean "skip unless nothing else works": that all-else-fails behaviour is what a separate, higher priority value gives (see db-dr below) |
| port | TCP/UDP port to connect to | used as-is, so the port no longer has to be hard-coded |
| target | host name of the server | resolved separately via its A/AAAA record |
- TXT record holds one or more quoted strings, each up to 255 bytes. DNS itself gives them no structure; the convention (used by DNS-based service discovery, RFC 6763) is one
key=valuepair per string, with a barekeymeaning a boolean "true". - TTL (time-to-live): how many seconds resolvers (the DNS servers, often run by an ISP or cloud provider, that look up records on a client's behalf and cache the answer) and clients may cache the answer. It is the propagation delay for any change you make.
Example records for db.example
_postgres._tcp.db.example. 300 IN SRV 10 60 5432 db1.example.
_postgres._tcp.db.example. 300 IN SRV 10 40 5432 db2.example.
_postgres._tcp.db.example. 300 IN SRV 20 0 5433 db-dr.example.
_postgres._tcp.db.example. 300 IN TXT "v=1" "sslmode=require" "pool_max=20" "connect_timeout_ms=3000" "prefer_same_zone"
Reading it: db1 and db2 share priority 10 and split traffic 60/40. db-dr (a disaster-recovery replica on port 5433) is priority 20, so it only receives connections when both priority-10 hosts are unreachable. The TXT record says: metadata schema version 1, TLS is mandatory, cap the connection pool at 20, give up a connect attempt after 3 seconds, and prefer a replica in the caller's zone (zone here means an availability zone, a distinct data-center-like unit inside a cloud region used for placement and failure isolation, not a DNS zone; this TXT string is just an application-level hint and has nothing to do with how DNS itself organizes names into zones).
How a client parses and uses them
-
Query
SRVandTXTfor_postgres._tcp.db.example(in parallel; both are cached for the TTL). -
Parse TXT: split each string at the first
=; lower-case the key; treat a string with no=as a boolean flag. Check thev=key first and refuse, or fall back to safe defaults, on a version you do not understand. Apply a default for every missing key so a deleted TXT string never produces a crash. -
Order SRV targets per RFC 2782: group by priority ascending; inside a group, sum the weights, draw a random integer from 0 to that sum inclusive, and take the first record whose running weight total reaches the draw. Remove it and repeat for the rest of the group. Zero-weight records are placed first in the list, which is what gives them a small, non-zero chance rather than none, not "only when nothing else works."
Worked draw, using the priority-10 pair above (
db1weight 60,db2weight 40): sorted ascending by weight, the group is[db2, db1]. The running totals are 40 (afterdb2) then 100 (afterdb1), and the draw is a random integer from 0 to 100 inclusive, 101 possible values. A draw of 25 stops atdb2(its running total, 40, is the first to reach 25). A draw of 85 stops atdb1(only its running total, 100, reaches 85). Becausedb2's running total is checked first,db2wins on any draw from 0 to 40 (41 of the 101 values, about 40.6%) anddb1wins on 41 to 100 (60 of 101, about 59.4%), which is why 10,000 simulated lookups land near 5951/4049 rather than an exact 6000/4000: the+1from the inclusive draw shifts the true odds slightly toward whichever record is checked first. The same mechanism is what gives a weight-0 record a chance at all when it shares a tier with a non-zero-weight record: it wins only on the single draw value of 0. -
Resolve the chosen target's
A(IPv4 address) orAAAA(the same idea for an IPv6 address) record, connect on the SRV port, and on failure move to the next target in the ordered list.
Runnable demonstration
This parses the four records above (as text, the way dig, the standard command-line DNS lookup tool, prints them), applies the TXT config with defaults, and checks the selection maths over 10,000 simulated lookups with a fixed seed.
import random, shlex
from collections import Counter
ANSWER = """
_postgres._tcp.db.example. 300 IN SRV 10 60 5432 db1.example.
_postgres._tcp.db.example. 300 IN SRV 10 40 5432 db2.example.
_postgres._tcp.db.example. 300 IN SRV 20 0 5433 db-dr.example.
_postgres._tcp.db.example. 300 IN TXT "v=1" "sslmode=require" "pool_max=20" "connect_timeout_ms=3000" "prefer_same_zone"
"""
def parse(text):
srv, cfg = [], {}
for line in text.strip().splitlines():
name, ttl, _cls, rtype, *rest = shlex.split(line)
if rtype == "SRV":
prio, weight, port, target = rest
srv.append({"priority": int(prio), "weight": int(weight),
"port": int(port), "target": target.rstrip(".")})
elif rtype == "TXT":
for s in rest: # each quoted string is one key=value pair
key, sep, val = s.partition("=")
cfg[key.lower()] = val if sep else True # bare key = boolean flag
return srv, cfg
def order(records, rng):
"""RFC 2782 ordering: lowest priority first, weighted random within a priority."""
out = []
for prio in sorted({r["priority"] for r in records}):
group = sorted((r for r in records if r["priority"] == prio),
key=lambda r: r["weight"]) # weight-0 entries first
while group:
total = sum(r["weight"] for r in group)
pick, running = rng.randint(0, total), 0
for i, r in enumerate(group):
running += r["weight"]
if running >= pick:
out.append(group.pop(i))
break
return out
srv, cfg = parse(ANSWER)
if cfg.get("v") != "1":
raise SystemExit("unknown metadata version")
pool_max = int(cfg.get("pool_max", 10))
timeout_ms = int(cfg.get("connect_timeout_ms", 5000))
print("config:", {"sslmode": cfg["sslmode"], "pool_max": pool_max,
"connect_timeout_ms": timeout_ms, "prefer_same_zone": cfg.get("prefer_same_zone", False)})
rng = random.Random(7)
first = Counter(order(srv, rng)[0]["target"] for _ in range(10_000))
print("first choice over 10,000 lookups:", dict(first))
down = {"db1.example", "db2.example"}
chosen = next(r for r in order(srv, rng) if r["target"] not in down)
print("with both priority-10 hosts down ->", f'{chosen["target"]}:{chosen["port"]}')
Output (Python 3, run as shown):
config: {'sslmode': 'require', 'pool_max': 20, 'connect_timeout_ms': 3000, 'prefer_same_zone': True}
first choice over 10,000 lookups: {'db1.example': 5951, 'db2.example': 4049}
with both priority-10 hosts down -> db-dr.example:5433
The 5951/4049 split is the roughly-60/40 weights showing through random draws (the exact odds are 59.4/40.6, worked out above), and db-dr is never a first choice while a priority-10 host exists. In production you would use a real resolver library (software that sends the actual DNS queries over the network and parses the standard binary wire format, rather than parsing text; for example dnspython in Python or net.LookupSRV in Go, which already returns records in RFC 2782 order) instead of parsing text.
Trade-offs and pitfalls
- Caching delays every change. With a 300-second TTL, draining
db1(removing it from rotation, typically before taking it down for maintenance, so it stops receiving new connections) by deleting its SRV record can take 5 minutes (longer if some resolver ignores TTLs). Keep TTLs short (30 to 60 s) for records you expect to change, and keep the old host alive through the drain. - Many clients ignore SRV. Browsers and most HTTP clients never look up SRV; libraries must opt in. Check your drivers before designing around it (the MongoDB
mongodb+srv://connection string is a well-known example that does use SRV and TXT this way). - DNS has no health checking. The records keep listing a dead host until someone (or an automation) removes it, so clients must still time out and fail over themselves.
- TXT is public and unauthenticated. Anyone who can query the zone can read it, and without DNSSEC (DNS Security Extensions, which sign records) a spoofed answer is possible. Never put passwords or tokens in TXT; restrict it to non-sensitive tuning knobs, and validate every value (range-check
pool_max, reject non-integers) because a typo in a zone file (the text file where a domain's DNS records, including these SRV and TXT records, are defined) becomes a production config change. - Size and type limits. 255 bytes per string, everything is a string, no nesting. Once config needs types, per-environment overrides, or audit history, move it to a real configuration store and keep DNS for location only.
That is every published Service Discovery and Configuration Management question for Network Engineer so far. Browse the other topics in this category, or practice this one interactively.