DNS, DHCP, and Name Resolution Questions
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
Design the authoritative DNS for a SaaS product serving 100M queries a day, with sub-50ms global latency and 99.99% availability. What choices do you make and what trade-offs do you accept?
Sample Answer
Direct answer
At 100 million queries a day the load is small: 100,000,000 / 86,400 s is about 1,157 queries per second on average, and a planning assumption of a peak ten times the average is about 11,600 per second, which any one modern authoritative server handles. The hard parts are the two other requirements. I would run anycast authoritative DNS from two independent managed providers, fed from one hidden primary (the server where the zone is edited, not listed in the NS records, so the public never queries it) through a tested pipeline, with short TTLs (time-to-live, how long caches may keep an answer) on the product's own records and a clear plan for the failure that multi-provider does not remove: a bad zone push. The 50 ms target is met by anycast proximity plus resolver caching, and 99.99% (52.56 minutes of allowed downtime a year, 4.32 minutes in a 30-day month) is met by not depending on a single provider.
The numbers behind it
| Quantity | Value | Working |
|---|---|---|
| Average load | 1,157 queries per second | 100,000,000 / 86,400 |
| Peak (assumed 10x, an assumption not a measurement) | about 11,600 queries per second | 1,157 x 10 |
| Allowed downtime at 99.99% | 52.56 minutes per year | 365 x 24 x 60 x 0.0001 |
| Allowed downtime at 99.99%, 30-day month | 4.32 minutes | 30 x 24 x 60 x 0.0001 |
| Two providers failing together, if fully independent | 0.0001 x 0.0001 = 0.00000001 (99.999999% available) | an upper bound, not the real figure, see "What multi-provider does not fix" |
Most queries never reach you. A recursive resolver caches an answer for its TTL, so what reaches the authoritative servers is the cache-miss traffic of the world's resolvers. The 50 ms target therefore applies to the round trip from a resolver to the nearest authoritative node, which is what anycast optimises. A worked picture, with illustrative round-trip times: a resolver in Frankfurt querying a single unicast server in Virginia pays about 90 ms per cache miss. If the provider announces the same address from Frankfurt as well, routing sends that resolver to the Frankfurt site and the miss costs about 8 ms. The end user only waits for that miss when their resolver's cache is cold; a warm cache answers in the resolver's own time. Measure it from outside rather than assuming it (choice 8).
Choices
For a first version, choices 1, 2, 3, 5 and 8 go live together (the providers, the delegation, the pipeline, TTLs and monitoring). Choices 4, 6 and 7 are layered on once those work.
1. Anycast, provided by managed providers. Anycast (RFC 4786) advertises the same server address from many sites, and routing delivers each query to a topologically near one. Building a global anycast network yourself needs your own AS (autonomous system) number, address space, BGP (Border Gateway Protocol, the Internet's routing protocol) peering and sites on every continent, which is a large cost for a service that a provider already runs. Buy it. RFC 4786 also warns that rapid BGP withdrawals and announcements (flaps) can destabilise routing, and that TCP exchanges may land on different nodes during instability, so choose providers that publish how they withdraw failed nodes.
2. Two providers, four name server names. Delegate the zone to name servers from both providers (for example two from each). The delegation is the set of NS records at the parent, and for four illustrative names it reads:
example.com. IN NS ns1.dns-a.example.net.
example.com. IN NS ns2.dns-a.example.net.
example.com. IN NS ns1.dns-b.example.org.
example.com. IN NS ns2.dns-b.example.org.
The same four names are also listed as NS records inside the zone, and named-checkzone accepted that zone. Resolvers that cannot reach one server try another and tend to prefer faster ones, so a provider outage costs retries and some latency rather than a full outage. This is the main defence against the largest single failure cause, the provider itself.
3. A hidden primary and a pipeline. Keep the zone's source of truth in version control, validate it with a checker (for BIND, named-checkzone), and push to both providers. Two ways to distribute:
- Zone transfer from a hidden primary. The provider pulls the zone with AXFR/IXFR (full and incremental zone transfer), authenticated by TSIG (a shared-secret signature on each transfer request) and a firewall allow-list of provider addresses. A configuration excerpt that passes
named-checkconf(the addresses are documentation placeholders):
zone "example.com" {
type primary;
file "example.com.zone";
allow-transfer { key provider-key; };
also-notify { 198.51.100.10; 203.0.113.10; };
};
- API push through a declarative zone tool, where each provider is a target. Required when a provider does not support transfers, and when you use provider-specific record types.
Whichever you choose, stagger the push: update provider A, verify with queries, then provider B. A pipeline that writes the same bad zone to both providers at the same second defeats the redundancy.
4. Geo answers and latency. Geographic or latency-based steering for the application's own records is a proprietary feature of each provider and is not carried by zone transfer. Two ways to handle it: keep the product's hostname (a name below the zone apex, such as app.example.com; the apex is the bare domain example.com, which must carry SOA and NS records, and a CNAME cannot share its name with other records, so the apex cannot be a CNAME) a plain CNAME or A record set that points at a global load balancer which does its own steering, or accept steering features common to both providers and configure each by API. The first is more portable, the second is more flexible but ties you to both providers' feature sets. I would choose the first for a first version, because identical zone content on both providers is simpler to verify.
5. TTLs and negative caching. Set 60 seconds on records that a failover must change, 300 seconds on stable ones, and set the SOA MINIMUM field (the last number in the SOA record) to a short value such as 60 seconds. Negative caching means a resolver remembers that a name did not exist, for the lower of the SOA record's TTL and MINIMUM (RFC 2308), so a record created after a failed lookup is not hidden for hours. In the lab, with a zone TTL of 300 and MINIMUM 60, a lookup of a name that did not yet exist returned NXDOMAIN with the SOA at TTL 60. The record was then added, and five seconds later the resolver still answered from its negative cache:
$ dig @127.0.0.1 new.example.com A +noall +comments +authority (5 s after the record was added)
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 32732
example.com. 54 IN SOA ns1.example.com. hostmaster.example.com. 2 3600 900 1209600 60
The 54 is the remaining negative-cache time, so the name stays invisible for about 55 more seconds, not 300. Short TTLs raise your query volume because caches expire sooner: at 60 seconds a busy resolver asks at most once a minute per name, so volume is bounded by the number of resolvers rather than the number of users. Illustration: if 20,000 busy resolvers each cache one name, the most they can send for it per day is 20,000 x 1,440 minutes = 28.8 million at a 60 s TTL, 5.76 million at 300 s and 0.48 million at 3,600 s. The NS and DS records at the parent (the registry, the organisation that runs the .com zone and holds your delegation) are controlled by the registry's TTLs, not yours, so a delegation change takes as long as the parent's TTL says. Plan delegation changes in advance.
6. DNSSEC with two providers. This is a decision to make at design time even if the first release ships unsigned. DNSSEC adds cryptographic signatures to DNS data so a validating resolver can detect forged answers. The zone-signing key (ZSK) signs your records, the key-signing key (KSK) signs the set of keys, and a DS record at the parent (a fingerprint of the KSK) ties the zone to the parent's chain of trust. Signing with two independent providers needs a multi-signer design (RFC 8901 describes a common KSK with per-provider ZSKs, and per-provider KSK and ZSK with several DS records). If you do not need DNSSEC at launch, make that a conscious decision, because retrofitting it across two providers is real work.
7. Health-checked failover. Let health checks withdraw an unhealthy application endpoint from DNS answers, but set a floor so that when every endpoint looks unhealthy the answers still return the full set instead of nothing. A monitoring fault should not become an outage.
8. Monitoring. Probe from many networks and regions against each provider's name servers directly (dig @each-ns +norecurse) and through public resolvers. Alert on answer differences between providers (and on a serial mismatch where transfers are used), and on probe latency per region. Define the SLO (service-level objective) on successful resolution at the resolver edge.
What multi-provider does not fix
Two providers each at 99.99% only combine to 99.999999% if their failures are fully independent, which they are not in practice. Shared dependencies set the real ceiling: the zone-build pipeline, the registrar account and its two-factor recovery, the parent's DS and NS records, an expired domain registration, and expired DNSSEC signatures. A bad zone push to both providers is more likely than two simultaneous provider outages. Mitigate with staged pushes, automated validation, a one-command rollback, registrar locks and expiry alerts, and keep DNSSEC signature expiry in monitoring.
Trade-offs I accept
- Cost and operational overhead of two providers, and the limit to features both support.
- Slightly higher resolver load from short TTLs, in exchange for fast failover. A resolver may also keep serving an expired cache entry while it cannot reach the authoritative servers (serve-stale, RFC 8767, which suggests 30-second TTLs on stale answers), but that depends on the resolver and must not be counted on.
- Slower delegation changes, because the parent's TTL is outside my control.
Design the recursive resolver service for an enterprise with millions of internal clients. What would you decide about placement, redundancy, cache sizing, security and privacy, and how would you know it is healthy?
Sample Answer
Direct answer
Run full recursive resolvers in-house as a fleet of identical caching nodes, advertised through anycast (one IP address announced from several nodes and sites, so routing sends each client to a nearby healthy one) from at least two sites, with no cache hierarchy at first. Give the nodes large caches, DNSSEC validation, QNAME minimisation and an access list, and measure health from the client side with synthetic lookups plus the resolver's own counters. The numbers below are sizing estimates built on stated assumptions, to be replaced by a load test and by measured traffic.
Two terms the rest depends on. Recursive resolution means the resolver itself walks the DNS tree for a client: it asks a root server, then a top-level-domain server, then the zone's own authoritative servers, and caches what it learns. Forwarding means the resolver does not walk the tree; it passes the client's question to another resolver and relays the answer. A fleet of full recursive resolvers does the first; a forwarder-only fleet does the second.
Placement and redundancy
- Sites. At least two failure domains (data centres or cloud regions, so that one power, network or software failure takes out only one of them), each with several nodes. Each node runs the resolver and a small health agent that announces the anycast address by BGP (the routing protocol) only while local lookups succeed, and withdraws it when they fail. Routers spread clients across the healthy nodes with ECMP (equal-cost multipath: hash each flow to one of several equal routes).
- Two addresses. Hand clients two resolver addresses (through DHCP option 6, the DNS server list, or their provisioning system), each announced from every site, so a client that loses one still has the other.
- Anycast versus unicast. Anycast gives failover at routing speed (the node withdraws its route) instead of waiting for a client's per-server timeout of several seconds. The cost is that a routing change can move a TCP session to a different node mid-flow, which is rare for DNS since TCP exchanges are short.
- Cloud. Each VPC or virtual network forwards only the company's own zones to the in-house fleet (and the cloud-internal zones to the provider's resolver), and the on-premises network forwards cloud private zones the other way. Keep forwarding rules from creating a loop between the two.
- Hierarchy. One tier, not two. A second shared cache tier raises hit rate but adds a failure domain and a hop. Add it only if measured upstream query volume exceeds what the fleet's internet egress and the authoritative servers' rate limits tolerate.
Capacity (ESTIMATE, from assumptions)
Assume 2,000,000 clients and a peak of 0.05 queries per second per client:
2,000,000 clients x 0.05 queries per second = 100,000 queries per second
With two sites of four nodes each, a healthy fleet gives 100,000 / 8 = 12,500 queries per second per node. If a whole site fails, the surviving four nodes carry 100,000 / 4 = 25,000 each. Buy and configure so a node holds 25,000 queries per second with headroom, then measure the real per-node ceiling with a query-replay tool before purchase; the 0.05 figure is an assumption, replace it with counts from your current resolvers.
Cache sizing
A resolver keeps a message cache (complete answers, such as the whole reply to "www.example.com A") and an RRset cache (the individual record sets inside those replies, which several different answers can share, for example one CNAME target set). Size from the working set, not from client count: the number of distinct names asked within a few hours dominates. Start with the rrset cache at twice the message cache because each cached answer references several record sets (a sizing choice to check against hit rate), set thread and slab counts to the CPU count (slab counts must be powers of two), and watch the hit ratio. The Unbound defaults are 4 MB for each cache and one thread, which is far too small for this fleet:
server:
# Who may ask, and where
interface: 0.0.0.0
access-control: 10.0.0.0/8 allow
access-control: 0.0.0.0/0 refuse
# Parallelism: one thread per CPU, one slab per thread (powers of two)
num-threads: 4
so-reuseport: yes
msg-cache-slabs: 4
rrset-cache-slabs: 4
infra-cache-slabs: 4
key-cache-slabs: 4
# Cache sizes
msg-cache-size: 512m
rrset-cache-size: 1024m
# Latency and resilience
prefetch: yes
serve-expired: yes
serve-expired-ttl: 86400
serve-expired-client-timeout: 1800
# Validation, privacy and abuse limits
qname-minimisation: yes
harden-dnssec-stripped: yes
aggressive-nsec: yes
edns-buffer-size: 1232
ratelimit: 1000
auto-trust-anchor-file: "/var/lib/unbound/root.key"
# Monitoring
extended-statistics: yes
What each group does (option meanings from the Unbound unbound.conf documentation):
- Who may ask.
interfaceis the address to listen on; the twoaccess-controllines allow the internal range and refuse everything else. - Parallelism.
num-threadsis how many worker threads serve clients. A slab is one lock-protected partition of a cache; more slabs mean threads wait for each other less.msg-cache-slabs,rrset-cache-slabs,infra-cache-slabsandkey-cache-slabsset that for the message cache, the record-set cache, the infrastructure cache (Unbound's memory of how fast and reachable each upstream server is) and the key cache (cached DNSSEC public keys).so-reuseportgives each thread its own listening socket so the kernel spreads incoming queries. - Cache sizes. The two sizes are in megabytes here, replacing the 4 MB defaults.
- Latency and resilience.
prefetchrefreshes a popular entry in its last 10 percent of TTL (the time a cache may keep a record), so users rarely wait on an expiry.serve-expiredlets the resolver answer from an expired entry when the authoritative servers cannot be reached, for at mostserve-expired-ttlseconds (86400 is one day).serve-expired-client-timeoutis in milliseconds: 1800 means wait 1.8 seconds for a fresh answer before sending the stale one. - Validation, privacy and abuse limits.
qname-minimisationsends upstream servers only the labels they need.harden-dnssec-strippedtreats a signed zone whose signatures have been removed as bogus instead of accepting it.aggressive-nsecreuses cached proofs of non-existence to answer for nonexistent names in signed zones without asking upstream again.edns-buffer-size: 1232is the largest UDP reply the resolver advertises.ratelimit: 1000caps queries per second that Unbound sends to the servers of one zone.auto-trust-anchor-fileholds the root zone's public key, the starting point for checking signatures, and keeps it updated. - Monitoring.
extended-statisticsadds more counters tounbound-control stats_noreset.
The access-control, cache-size and trust-anchor lines carry the essentials of a safe, correctly sized resolver (the qname-minimisation, edns-buffer-size and aggressive-nsec settings are also on by default in current Unbound, so listing them records the intent); the slab, so-reuseport, prefetch, serve-expired, aggressive-nsec and ratelimit lines are tuning for load and resilience, and the resolver works without them.
unbound-checkconf reported no errors on this configuration with the comments added (run with the interface set to 127.0.0.1 and the trust anchor at a lab path). prefetch refreshes popular records before they expire, and serve-expired answers from an expired cache entry when the authoritative servers cannot be reached (RFC 8767 describes the behaviour, and suggests a maximum stale age between 1 and 3 days; this configuration allows one day). serve-expired-client-timeout: 1800 waits 1.8 seconds before falling back to the stale answer, matching the RFC's suggested 1.8 seconds.
Security and privacy
- Not an open resolver. The
access-controllines refuse everyone outside the internal ranges; an open resolver becomes a reflector. - Validation. DNSSEC validation with the root trust anchor. In the lab, answers were counted under
num.answer.secure=2after two queries; a failed validation returns SERVFAIL to the client, so watch for it. - Query privacy toward the internet. QNAME minimisation (RFC 9156) sends each upstream server only as many labels of the name as it needs (the root servers are asked for
com., not forwww.example.com.), so the root and TLD servers do not see full names. Internal zones go to internal servers by forwarder rules and never leak outward. - Privacy toward staff. Query logs record browsing. Keep per-client logging off by default, use the aggregate counters, and put any sampled log behind access control and a retention limit.
- Abuse protection.
ratelimitlimits the rate of queries Unbound itself sends upstream while recursing, which protects authoritative servers from your own fleet when a client misbehaves;access-controlkeeps outsiders away from the fleet.
How you know it is healthy
Two views, because a resolver can look fine to itself and be broken for users.
- Synthetic checks from the clients' side, every 10 seconds per site and per address:
dig +time=2 +tries=1 @<anycast-address> <known-name>. Alert on failures and on latency. - The resolver's own counters (
unbound-control stats_noreset). From the lab, after two lookups of the same name:
total.num.queries=2
total.num.cachehits=1
total.num.cachemiss=1
total.requestlist.exceeded=0
total.recursion.time.avg=0.225819
num.answer.rcode.SERVFAIL=0
num.answer.secure=2
total.recursion.time.avg is in seconds (0.225819 is about 226 milliseconds per lookup that needed recursion); the other lines are counts. The hit ratio is hits / (hits + misses) = 1 / 2 here, only to show the arithmetic. In production, alert on a hit ratio that falls from its baseline, a rising SERVFAIL share (a starting threshold of 2 percent over 5 minutes, to be tuned against the baseline), any growth in total.requestlist.exceeded (the resolver is overloaded), and total.recursion.time.avg climbing.
Trade-offs
Anycast makes failover fast but hides which node answered, so keep the answering node's name in logs and monitoring labels. A large cache lowers upstream traffic but a cold node after a restart is slow for minutes, so stagger restarts. The alternative of forwarding every query to a cloud provider's resolver removes the fleet to run, at the cost of sharing that provider's outages and policy.
Your team is deciding between forwarding all queries to a cloud provider's resolver and running full recursive resolvers in-house. How do you weigh them?
Sample Answer
Direct answer
Run your own recursive resolvers and use forwarding only for the zones that must go somewhere specific (internal Active Directory zones to the domain controllers, cloud private zones to the cloud resolver). Forwarding everything to the cloud provider is the better choice only when nobody on the team can operate a resolver tier, or the resolvers cannot be given direct internet access on port 53. The reason to prefer recursion is that you keep control of validation, visibility, and failure behaviour; the reason to forward is that you hand the operating cost and the capacity problem to someone else.
Three terms decide this question. A recursive resolver answers a client by doing the lookup itself: it asks a root server, then a top-level-domain server, then the zone's authoritative servers, and caches the result. Forwarding means the resolver does no walking; it hands the question to a named upstream resolver (here the cloud provider's) and relays that resolver's answer. Conditional forwarding is forwarding only for the zones you name and recursing for everything else.
Weighing them on the six axes
| Axis | Forward everything to the cloud resolver | Full recursion in-house |
|---|---|---|
| Latency | A cold name costs one hop to a provider cache that is warm for popular names because it serves many customers | A cold name needs the full walk from root to TLD to authoritative. In the lab one uncached validated lookup took an average of 0.226 seconds (a single sample, not a benchmark); a warm name is answered from local cache |
| Cache hit rate | Your cache plus the provider's. For a small site the provider's much larger population gives far better hit rates | Only your population warms the cache. At millions of clients that is plenty; at a few hundred it is not |
| Privacy | One operator sees every name your users ask | Each authoritative server sees only its own part, and with QNAME minimisation (RFC 9156) only as many labels as it needs |
| DNSSEC validation | Whatever the provider does. You must confirm it validates, and a bogus answer reaches you only as an error | You validate yourself with the root trust anchor and see the failure reason in your own logs |
| Operations | Nothing to patch or scale. You are limited by the provider's query limits (the AWS documentation for Route 53 VPC Resolver lists 10,000 UDP queries per second per endpoint IP address; the worked comparison below shows what that means for a large fleet) | You patch, size, monitor and defend the fleet |
| Failure mode | A provider outage or throttling is your outage unless you have a second path | Your own bugs and capacity are your outage, but you can serve stale data and choose the fallback |
Recommendation and what flips it
Commit to in-house recursion with conditional forwarding. The Unbound configuration below sends only the company zone to the domain controllers and resolves every other name by full recursion; there is no rule for . (the root, meaning "everything else"), so those queries start at the root servers. unbound-checkconf reported no errors on it:
server:
interface: 127.0.0.1
access-control: 127.0.0.0/8 allow
trust-anchor-file: "/usr/share/dns/root.key"
domain-insecure: "corp.example.com"
forward-zone:
name: "corp.example.com."
forward-addr: 10.20.0.10
forward-addr: 10.20.0.11
trust-anchor-file is the root zone's public key, the starting point from which Unbound checks DNSSEC signatures (DNSSEC validation) on what it resolves itself. domain-insecure tells Unbound not to require a DNSSEC chain for that internal name. It is there because the internal zone sits under a signed public parent (example.com): Unbound sends the DS lookup for corp.example.com to the same forwarders (its log showed the DS queries answered by the forwarder), and an unsigned internal server cannot supply the signed proof of "no DS" that the signed parent would, so validation fails. In the lab (an unsigned corp.example.com zone served by BIND, queried through Unbound with this forward-zone, and the real signed example.com parent) the name returned SERVFAIL without the line and NOERROR with it. The price is that names under corp.example.com are not DNSSEC-validated, which matches reality because the zone is unsigned.
If a team does choose to send everything to the provider first, the variant below adds one more block (also checked with unbound-checkconf, no errors):
forward-zone:
name: "."
forward-addr: 10.0.0.2
forward-first: yes
With forward-first: yes, when a forwarded query gets a SERVFAIL (the resolver's "I could not get an answer" code), Unbound falls back to less specific resolution (the Unbound documentation's wording): it resolves the name itself and validates the result with its own trust anchor. So the fallback keeps the resolver working through a provider fault. One consequence to accept knowingly: if the provider applies filtering by returning SERVFAIL, the fallback would bypass it, so check how your provider signals a blocked name.
A worked comparison with the provider's limit, on assumed numbers: 2,000,000 clients x 0.05 queries per second = 100,000 queries per second. This treats every client query as reaching the endpoint with no cache in front of it, so it is an upper bound: a caching forwarder tier in front of the endpoint would cut the figure by its hit ratio. At 10,000 per endpoint IP address that is 100,000 / 10,000 = 10 addresses running at 100 percent. The same AWS page advises adding addresses once any one passes 50 percent of capacity, so plan for 20. One endpoint holds up to 6 addresses and the default quota is 4 endpoints per account per Region, so 24 addresses fit within the defaults, with little room to spare. The per-address limit applies to the endpoints that carry forwarded traffic, so it matters most when everything is forwarded. The same AWS page adds that where connection tracking is enforced by restrictive security group rules, or queries are routed through a Network Load Balancer, the per-address figure for an inbound endpoint can be as low as 1,500 queries per second, so 10,000 is a ceiling and not a planning number. Both the 0.05 figure and the 100 percent / 50 percent planning are assumptions to replace with your own measurements.
What flips the decision to forward-everything:
- No resolver operations capacity. A small team with no on-call for DNS is safer on the provider's fleet.
- Egress restriction. If policy forbids resolvers from reaching arbitrary internet servers on UDP and TCP port 53 (RFC 9210 requires TCP support), full recursion cannot work.
- Small user population. The provider's warm cache outperforms yours.
- The provider offers the telemetry you need. Query logging and DNS filtering built in, for example, remove the main visibility argument for self-hosting.
What flips it the other way: needing your own query visibility, wanting independence from one provider's incidents, running multi-cloud or hybrid (one resolver tier serves every environment), or needing response policy you control (rules on the resolver that rewrite or block answers for names you choose, for example to blackhole known malware domains).
Decide with a measurement, not an argument
Put a canary network segment (a small slice of users or a test network moved onto the option under test before everyone else) on each option for a week and compare, for the same query mix: median and 95th percentile resolution time, SERVFAIL share, and cache hit ratio from unbound-control stats_noreset (hits divided by hits plus misses). Include a failure drill: block the upstream in each case and see how long users take to notice.
Pitfalls
- Forwarding to the provider and also running a local cache hides who is responsible for a wrong answer: record which tier answered.
- A forwarder loop (cloud resolver forwards corp names to on-premises, which forwards cloud names back) fails slowly and confusingly; list each zone's single owner.
- Firewalls that allow only UDP to the resolver break large answers, because the client's fallback to TCP is blocked.
Design name resolution for an environment with several cloud accounts, Kubernetes clusters and an on-premises network, where internal services need private names and the same names must resolve differently outside. How do queries flow in each direction, and what would you do to avoid inconsistent answers?
Sample Answer
Direct answer
Give every zone exactly one owner, let resolvers find the owner by forwarding on the longest matching domain suffix, and generate any name that must differ inside versus outside from one source of truth. Concretely: a private root such as internal.example.com split into sub-zones (aws., k8s., onprem.), each served by the team and system that owns those hosts; on-premises resolvers forward the cloud and cluster sub-zones to the cloud resolvers' inbound endpoints; the cloud resolvers forward the on-premises sub-zone out through outbound endpoints; each Kubernetes cluster's CoreDNS (the in-cluster DNS server) answers its own names and forwards the rest to the cloud resolver. Split-horizon (same name, different answer inside and outside) is done by two zones of the same name with different audiences, and every record in both comes from one repository and is checked by an automated consistency test.
Two rules do most of the work: one owner per zone, and forwarding on the longest matching suffix. Everything else below (generated views, the checker script, TTL choices) exists to keep those two rules true over time.
Terms used below. A zone is one slice of the DNS namespace that a single set of servers is responsible for. Kubernetes terms: a cluster is a set of machines running containers; a pod is the smallest unit that runs your containers and gets its own IP address; a namespace is a named grouping of cluster objects, and kube-system is the one where Kubernetes keeps its own components; a ConfigMap is a cluster object that holds configuration text other components read (CoreDNS reads its settings from one); cluster.local is the default DNS suffix for names inside a cluster, such as my-service.default.svc.cluster.local. AWS terms: a private hosted zone is a DNS zone that only the VPCs you attach it to can see; an inbound endpoint is a set of IP addresses inside your VPC that resolvers outside it (such as on-premises ones) can send queries to; an outbound endpoint is the set of addresses from which the VPC resolver sends queries out to resolvers elsewhere; a forwarding rule says "queries for this domain go to these resolver addresses". An on-premises resolver does the same with a conditional forwarder. A stub domain in CoreDNS is a block that sends one suffix to a chosen upstream resolver. Longest-suffix match means that when several rules fit a name, the one with the most matching trailing labels wins: api.k8s.corp.lab matches both corp.lab and k8s.corp.lab, and k8s.corp.lab is chosen.
Who asks whom
| Query starts at | Asks first | What sends it onward | Answered by |
|---|---|---|---|
Pod, name ends cluster.local | CoreDNS in the cluster | nothing, CoreDNS owns it | CoreDNS |
| Pod, any other name | CoreDNS | default forward (or a stub-domain block) to the cloud resolver | cloud resolver, or the stub domain's upstream |
| Cloud workload, cloud name | VPC resolver (base plus 2) | the private hosted zone attached to that VPC | the hosted zone |
| Cloud workload, on-premises name | VPC resolver | forwarding rule, out through the outbound endpoint | on-premises resolvers |
| On-premises client, cloud or cluster name | on-premises resolver | conditional forwarder to the inbound endpoint addresses | the VPC resolver behind the endpoint |
| Anyone, public name | their resolver | normal recursion | the public zone's provider |
Query flow in each direction
The cloud side is described with Amazon Route 53's resolver (VPC Resolver, where a VPC is a Virtual Private Cloud, a private network inside a cloud account); other clouds have equivalent inbound and outbound forwarding features.
- Pod to an in-cluster name (
cluster.local): CoreDNS answers it. These names never leave the cluster. - Pod to any other name: CoreDNS forwards to the cloud resolver (in a VPC, the address at the base of the VPC range plus two, for example 10.0.0.2 for 10.0.0.0/16). Kubernetes lets you add stub domains in the
corednsConfigMap inkube-system, for example the block shown after this list, which sends one suffix to a specific upstream. CoreDNS does not support a hostname as the upstream of a stub domain, only IP addresses. - Cloud workload to a cloud name: the VPC resolver answers from the private hosted zone (a private DNS zone attached to chosen VPCs) associated with that VPC. Route 53 picks the most specific matching zone and returns NXDOMAIN (name does not exist) if that zone has no such record. It does not fall back to the public zone.
- Cloud workload to an on-premises name: a forwarding rule for
onprem.internal.example.comsends the query out of an outbound endpoint to the on-premises resolvers over the VPN or private link. If a query matches several rules the most specific one wins, and a resolver rule overrides a private hosted zone of the same name. Rules can be shared across accounts, so one set of rules serves all accounts. - On-premises client to a cloud name: the on-premises resolver has a conditional forwarder for
aws.internal.example.comandk8s.internal.example.compointing at the cloud resolver's inbound endpoint addresses. Point the forwarders at the inbound endpoint addresses, not at the VPC base-plus-two address: the inbound endpoint is the address set AWS provides for queries coming from your own network. - Anyone to a public name: resolved recursively as usual. The public zone for
example.comstays at its public provider.
The stub-domain block from flow 2 (the format Kubernetes documents for the coredns ConfigMap, whose own example uses consul.local:53 and 10.150.0.1; the suffix here is this design's):
onprem.internal.example.com:53 {
errors
cache 30
forward . 10.150.0.1
}
Reading it line by line: onprem.internal.example.com:53 { opens a block that applies only to queries for that suffix, arriving on port 53. errors makes CoreDNS log the errors it hits to standard output. cache 30 keeps answers for at most 30 seconds. forward . 10.150.0.1 sends every query in this block to the upstream resolver at 10.150.0.1 (the . means all names in the block's zone; the address is illustrative). Queries for any other suffix do not match this block and take CoreDNS's default path.
Making the longest-suffix rule concrete (executed)
Unbound configured with two forward zones, corp.lab to one authoritative server and k8s.corp.lab to another, chose by suffix. Because the lab upstreams sit on loopback, the server: block needs do-not-query-localhost: no; without it Unbound refuses to use 127.0.0.1 upstreams and answers SERVFAIL. In production the forward addresses are real resolver addresses and that line is not needed.
server:
do-not-query-localhost: no
forward-zone:
name: "corp.lab"
forward-addr: 127.0.0.1@5301
forward-zone:
name: "k8s.corp.lab"
forward-addr: 127.0.0.1@5302
The server on port 5301 held corp.lab with ldap IN A 10.0.0.20 and www IN A 10.0.0.80; the server on port 5302 held k8s.corp.lab with api IN A 10.96.0.1. dig @127.0.0.1 +short ldap.corp.lab A returned 10.0.0.20 (from the parent zone's server) and dig @127.0.0.1 +short api.k8s.corp.lab A returned 10.96.0.1 (from the cluster's server). The lab uses corp.lab; in production, build the private root under a domain you own so it cannot collide with a name someone else registers.
Split-horizon without inconsistency
The same name can legitimately differ by audience. In BIND this is a view per audience; two views of corp.lab, with the internal one matched by source address, answered www.corp.lab differently from the same server:
view "internal" { match-clients { 127.0.0.1; }; zone "corp.lab" { type primary; file "/z/db.corp.int"; }; };
view "external" { match-clients { any; }; zone "corp.lab" { type primary; file "/z/db.corp.ext"; }; };
The two zone files were identical except for one line, www IN A 10.0.0.80 in db.corp.int and www IN A 203.0.113.80 in db.corp.ext. Queried from 127.0.0.1 it returned 10.0.0.80; queried from 127.0.0.2 (dig -b 127.0.0.2 @127.0.0.2, after ip addr add 127.0.0.2/32 dev lo so that BIND listens on that address) it returned 203.0.113.80. In Route 53 the same pattern is a public and a private hosted zone with the same name. The private one wins for the VPCs it is attached to, and a name missing from it yields NXDOMAIN inside, so every public name the inside also needs must be copied into the private zone.
What I do to avoid inconsistent answers
- One owner per name. Each sub-zone has one authoritative system. A parent zone delegates with NS records (private hosted zones support NS delegation of a subdomain); nobody duplicates a sub-zone into a second place.
- Same rules everywhere. The forward rules on every resolver (on-premises, each VPC, each cluster) come from one definition, because an environment missing the
k8srule falls into the parent zone's rule and gets NXDOMAIN from a server that never held the name. - Generate both split-horizon views from one record source. Internal and external records live in the same repository and are rendered into each zone, so a name cannot exist in one view only by accident.
- Test it continuously. A checker script queries each resolver for a list of expected answers and flags any mismatch:
#!/usr/bin/env bash
# usage: dnscheck.sh expected.txt resolver [resolver...] (resolver = host or host:port)
exp=$1; shift; rc=0
while read -r name type want; do
[[ -z $name || $name == \#* ]] && continue
for r in "$@"; do
host=${r%%:*}; port=53; [[ $r == *:* ]] && port=${r##*:}
got=$(dig @"$host" -p "$port" +short +time=2 +tries=1 "$name" "$type" 2>&1 | grep -v "^;;" | sort | paste -sd, -)
if [[ $got == "$want" ]]; then s=ok; else s=MISMATCH; rc=1; fi
printf '%-9s %-22s via %-15s want=%-12s got=%s\n' "$s" "$name" "$r" "$want" "${got:-<empty>}"
done
done < "$exp"
exit $rc
The expectations file expected.txt has one name type want line per check:
# name type want
ldap.corp.lab A 10.0.0.20
api.k8s.corp.lab A 10.96.0.1
www.corp.lab A 10.0.0.80
Run as bash dnscheck.sh expected.txt 127.0.0.1:53 127.0.0.1:5354, with an on-premises resolver (the Unbound configuration above, port 53) and a cloud-side resolver (a second Unbound on port 5354 with only the corp.lab forward zone, built without the k8s rule), it printed:
ok ldap.corp.lab via 127.0.0.1:53 want=10.0.0.20 got=10.0.0.20
ok ldap.corp.lab via 127.0.0.1:5354 want=10.0.0.20 got=10.0.0.20
ok api.k8s.corp.lab via 127.0.0.1:53 want=10.96.0.1 got=10.96.0.1
MISMATCH api.k8s.corp.lab via 127.0.0.1:5354 want=10.96.0.1 got=<empty>
ok www.corp.lab via 127.0.0.1:53 want=10.0.0.80 got=10.0.0.80
ok www.corp.lab via 127.0.0.1:5354 want=10.0.0.80 got=10.0.0.80
and exited 1, so a pipeline stops. Run it from each environment against its own resolver. The MISMATCH row is the one to act on: the resolver on port 5354 returned nothing (<empty>) for the k8s name because it lacks the k8s forwarding rule.
Reading the script: exp=$1; shift; rc=0 stores the expectations file name, drops it from the argument list so "$@" holds only resolvers, and starts the exit code at 0. while read -r name type want reads each line of the file as three fields. [[ -z $name || $name == \#* ]] && continue skips blank lines and lines starting with # (comments). host=${r%%:*} cuts everything from the first : onward, leaving the host; port=53; [[ $r == *:* ]] && port=${r##*:} keeps 53 unless the resolver was written host:port, in which case ${r##*:} keeps everything after the last :. For 127.0.0.1:5354 these give 127.0.0.1 and 5354. In the dig line, +short prints only the answer data, +time=2 +tries=1 limits the wait to two seconds and one attempt, grep -v "^;;" drops dig's error lines (they start with ;;), sort orders multi-address answers, and paste -sd, - joins the lines into one comma-separated string so it can be compared with the expected value. The if compares, sets s to ok or MISMATCH and sets rc=1 on any mismatch, printf prints the aligned row, ${got:-<empty>} prints <empty> when nothing came back, and done < "$exp" feeds the file to the loop. exit $rc returns 1 if any row mismatched.
- Short TTLs on dynamic names, longer on stable ones. Cluster service names change often (the Kubernetes example above caches for 30 seconds); stable infrastructure names can be longer. Mismatched TTLs can also make two environments disagree for a few minutes even when the data is identical.
Trade-offs
- A central resolver layer is one more path to keep highly available; give each endpoint at least two addresses in different zones or sites.
- Split-horizon doubles the places a record can be wrong. If the inside and outside answers do not need to differ, publish one name and expose it only on private addresses behind access control.
- If everything is in a single cloud with no on-premises network, drop the on-premises forwarders and keep the rest.
That is every published DNS, DHCP, and Name Resolution question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.