DNS, DHCP, and Name Resolution Questions
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
How would you know that your DNS service is healthy from the user's point of view, before anyone complains, and how would you avoid paging people for noise?
Sample Answer
Direct answer
Measure what a user experiences, not what the server says about itself: a synthetic probe (a scripted lookup run on a schedule) resolves real names through the same resolvers users use, from several regions, and checks both the answer and the time. Page a human only when probes from more than one region fail persistently. Everything else (error ratios, cache hit ratio, per-region latency drift) goes to dashboards and tickets, because those signals are noisy on their own and many have causes nobody should be woken for.
What to measure, and who gets told
| Signal | What it tells you | Source | Action |
|---|---|---|---|
| Probe success: NOERROR and the expected answer | The user-visible outcome, including wrong answers | Probes in 3 or more regions | Page when at least 2 regions fail 3 probes in a row |
| Probe latency, 95th percentile (p95: 95% of probes are faster) per region | Slowness a user feels, per region | Same probes | Ticket when above baseline for 15 minutes; alert only if it nears client timeouts |
| Two probe kinds: a warm name and a unique never-cached name | Warm = resolver speed. Unique = the full path to the authoritative servers | Probe names under a zone you control, e.g. a wildcard record (explained after the table) | Compare: only the unique name slow points past the resolver |
| SERVFAIL ratio (the resolver's "could not answer" code) | Upstream or resolver trouble | Resolver counters | Ticket; page only if probes also fail |
| NXDOMAIN ratio (the "name does not exist" code), against its own baseline | A jump means a client or deploy is asking for wrong names | Resolver counters | Ticket |
| Cache hit ratio = hits / (hits + misses) | Falling hit ratio raises latency and upstream load | Resolver counters | Dashboard |
| Zone consistency: the SOA serial (the zone version number) equal on every name server (NS, the servers named as the zone's authorities) | Replication or zone-transfer problems (a zone transfer is how a secondary server copies the zone from the primary) | Probe that asks each NS | Ticket |
A unique name needs a wildcard record, one that answers for every name under a label. Put *.probe IN A 203.0.113.50 (an illustrative address) in a zone you control and every name ending in .probe.example.com resolves, so each probe can invent a name no cache has ever seen, such as probe-$RANDOM-$(date +%s).probe.example.com. A real generated name from the lab, passed through the script below, was probe-11127-1791275626.probe.example.com:
$ dns-probe.sh 127.0.0.1 probe-11127-1791275626.probe.example.com 203.0.113.50
probe resolver=127.0.0.1 name=probe-11127-1791275626.probe.example.com result=ok answer=203.0.113.50 latency_ms=0
The per-region and unique-name points are what make it "from the user's point of view": a resolver can report zero errors while one region's users wait seconds, and a cached name can hide an authoritative outage.
A probe you can run
This script checks one resolver and one name, labels the outcome and prints one machine-readable line. It passes shellcheck (a shell linter) with no findings.
#!/usr/bin/env bash
# Synthetic DNS probe: one line per (resolver, name) with outcome and latency.
# usage: dns-probe.sh RESOLVER NAME EXPECTED_ANSWER
set -u
resolver=$1 name=$2 expect=$3
out=$(dig @"$resolver" "$name" A +tries=1 +time=2 +noall +comments +answer +stats 2>&1)
status=$(grep -o 'status: [A-Z]*' <<<"$out" | awk '{print $2}')
ms=$(grep -o 'Query time: [0-9]*' <<<"$out" | awk '{print $3}')
answer=$(grep -E "^${name//./\\.}\.[[:space:]]" <<<"$out" | awk '{print $NF}' | sort | paste -sd, -)
if [ -z "$status" ]; then result=no_response
elif [ "$status" != NOERROR ]; then result=$status
elif [ "$answer" != "$expect" ]; then result=wrong_answer
else result=ok; fi
printf 'probe resolver=%s name=%s result=%s answer=%s latency_ms=%s\n' \
"$resolver" "$name" "$result" "${answer:--}" "${ms:--}"
Against a lab resolver (Unbound on 127.0.0.1 forwarding example.com to an nsd authoritative server on 127.0.0.2, with www = 203.0.113.10 and api = 203.0.113.11):
$ dns-probe.sh 127.0.0.1 www.example.com 203.0.113.10
probe resolver=127.0.0.1 name=www.example.com result=ok answer=203.0.113.10 latency_ms=1
$ dns-probe.sh 127.0.0.1 www.example.com 198.51.100.7 # expectation drifted
probe resolver=127.0.0.1 name=www.example.com result=wrong_answer answer=203.0.113.10 latency_ms=0
$ dns-probe.sh 127.0.0.1 nope.example.com -
probe resolver=127.0.0.1 name=nope.example.com result=NXDOMAIN answer=- latency_ms=0
--- authoritative server stopped
probe resolver=127.0.0.1 name=api.example.com result=no_response answer=- latency_ms=-
--- resolver address with nothing listening
probe resolver=127.0.0.9 name=www.example.com result=no_response answer=- latency_ms=-
What each line of the script does:
set -umakes an unset variable an error, so a forgotten argument stops the script instead of probing an empty name. The next line takes the three arguments.out=$(dig ...)runs one query and stores everything it prints.@"$resolver"picks the server to ask;+tries=1 +time=2means one attempt with a 2 second wait;+noall +comments +answer +statsprint only the header comments, the answer section and the statistics (including the query time);2>&1also captures error messages.status=takes the textstatus: NOERRORout of the header (grep -oprints only the matching part,<<<"$out"feeds the variable togrepas its input) andawk '{print $2}'keeps the second word,NOERROR.ms=does the same forQuery time: 12(third word,12).answer=keeps the answer-section lines that start with the probed name.${name//./\.}replaces every.in the name with\.because in a regular expression a bare dot means any character (www\.example\.com);awk '{print $NF}'keeps the last field of each line, the address;sort | paste -sd, -sorts the addresses and joins them with commas, so a multi-address answer compares the same way each time.- The
ifchain labels the outcome in this order: no status at all isno_response; any status other thanNOERRORis reported as that code; aNOERRORwhose address list differs from the expected one iswrong_answer; otherwiseok. printfprints onekey=valueline;${answer:--}means "use-when empty".
Four outcomes are visible: a correct answer, a wrong answer (the probe catches a changed record that a status-code check would miss), an error code, and no answer at all. When the authoritative server stops, a user of an uncached name sees a timeout, not a clean error, because the resolver spends longer retrying than the probe's 2 second limit. Ship the line to your metrics system and compute the ratio per region. The REFUSED code (the server declines to answer, for example because the client is not allowed to ask) and SERVFAIL both land in the result= field as their own labels.
Resolver-side counters come from the resolver itself. For Unbound, unbound-control stats_noreset prints total.num.cachehits, total.num.cachemiss and, when extended-statistics: yes is set in the server: block of unbound.conf, num.answer.rcode.SERVFAIL (all seen in the lab with that option on; without it the per-rcode lines are absent); other resolvers publish comparable counters, so check your own resolver's documentation for the exact names.
Paging without noise
Assume a probe fails by chance with probability 0.5% (an assumed figure for transient packet loss, independent between probes), runs every 30 seconds from each of 3 regions, so 2,880 probes per region per day.
- Page on any single failed probe: 2,880 x 3 x 0.005 = 43.2 expected false pages a day (8,640 probes x 0.5%).
- Page on 3 failures in a row at one region: 2,880 x 3 x 0.005^3 = 0.001 a day, at the price of a 90 second detection delay (3 x 30 s). The pieces: a day has 86,400 seconds, and one probe every 30 seconds gives 86,400 / 30 = 2,880 probes per region; 3 regions give 2,880 x 3 = 8,640 probes a day. Each probe ends one possible run of 3, so there are about 8,640 runs, and a run is all-failed with probability 0.005^3 = 1.25 x 10^-7. Then 8,640 x 1.25 x 10^-7 = 0.00108, about 0.001 a day.
- Require at least 2 of the 3 regions to agree before paging, and send a single-region failure as a ticket: a local network fault near one probe host stops paging anyone.
Real failures are correlated, so the independence assumption is optimistic. That is why the rule needs two regions and a duration, not just a threshold. Also silence the alert during announced maintenance, and keep a deliberate-NXDOMAIN probe separate so an expected "no such name" result is not counted as an outage.
If you set a service-level objective (SLO: a target for the share of good probes) of 99.9% over 30 days, the error budget (how much failure the target still allows) is 0.1% of the 30 x 24 x 60 = 43,200 minutes in the window, which is 43.2 minutes of probe-failed time.
What query logging costs
Assume an illustrative 2,000 queries per second and about 150 bytes per log line: 2,000 x 150 = 300,000 B/s, or 25.92 GB a day. Logging 1% of NOERROR queries cuts that to about 259 MB a day; keep all SERVFAIL and REFUSED lines, because they are rare: at 0.1% of traffic that is 2 per second, about 25.9 MB a day. Counters stay unsampled and are cheap, so rates and ratios never depend on the sample.
Pitfalls
Alerting on the server's own "up" check misses wrong answers and regional slowness. Paging on SERVFAIL ratio alone wakes people for one customer's broken delegation (the NS records that hand a name to the servers responsible for it). Probing only popular names measures the cache, not the service.
You are setting up DNS for a domain that must receive email and also advertise a SIP or LDAP service on non-default ports. Which record types do you use for each, how do they differ, and where do the rules about what a name can point to constrain your design?
Sample Answer
Direct answer
Use MX records for mail and SRV records for the SIP and LDAP services, because only SRV carries a port number. MX says which hosts accept mail for the domain, ranked by preference. SRV says which hosts, on which ports, with what priority and weight, provide a named service. Both point at hostnames, and those target names must have address records (A or AAAA) and must not be aliases (CNAMEs). That one rule, plus the rule that a CNAME cannot share its name with any other record, drives the design.
Record types and how they differ
| MX | SRV | |
|---|---|---|
| Owner name (the name the record is attached to) | The mail domain itself, for example example.com. | _service._proto.name, for example _sip._tcp.example.com. (RFC 2782) |
| Fields | Preference, target host | Priority, weight, port, target host |
| Port | Always the mail protocol's own port; MX cannot carry one | Carried in the record, so non-default ports work |
| Selection | Lowest preference first; equal preferences are randomised by the sender (RFC 5321 section 5.1) | Lowest priority first; weight splits traffic among equal priorities |
| Target rule | Must not be an alias (RFC 2181 section 10.3) | Target must have address records and must not be an alias (RFC 2782) |
Mail. MX records (mail exchanger records) name the receiving hosts. Lower preference numbers are preferred, and if several have the same preference the sender randomises them to spread load. If a domain has no MX records at all, a sender treats the domain's own address record as an implicit MX with preference 0 (RFC 5321 section 5.1), so a domain that must not receive mail should publish a null MX (RFC 7505: a single MX with preference 0 and target .).
SIP and LDAP. SIP (Session Initiation Protocol) sets up voice and video calls, and LDAP (Lightweight Directory Access Protocol) queries a directory such as a user database. A client looks up _sip._tcp.example.com or _ldap._tcp.example.com and gets priority, weight, port and target. The port is the point: SIP on 5070 or LDAP on 3890 is reachable without every client being configured with the number. Priority is a failover order, and weight divides traffic between targets of equal priority (RFC 2782). A target of . (a single dot, the root) is the way to say the service is decidedly not available at this domain: RFC 2782 defines it as exactly that, so a client stops looking instead of guessing. The service name begins with an underscore, as in _sip._tcp, to avoid collisions with ordinary host names that occur in nature (RFC 2782).
A zone that does all three
$TTL 300
@ IN SOA ns1.example.com. hostmaster.example.com. ( 1 3600 900 1209600 300 )
IN NS ns1.example.com.
ns1 IN A 127.0.0.2
; mail for the domain
@ IN MX 10 mx1.example.com.
@ IN MX 20 mx2.example.com.
mx1 IN A 192.0.2.25
mx2 IN A 192.0.2.26
; SIP over TCP on 5070, 60/40 split between two gateways
_sip._tcp IN SRV 10 60 5070 sipgw1.example.com.
_sip._tcp IN SRV 10 40 5070 sipgw2.example.com.
sipgw1 IN A 192.0.2.40
sipgw2 IN A 192.0.2.41
; LDAP on 3890, dir1 preferred, dir2 only if dir1 is unreachable
_ldap._tcp IN SRV 0 0 3890 dir1.example.com.
_ldap._tcp IN SRV 10 0 3890 dir2.example.com.
dir1 IN A 192.0.2.50
dir2 IN A 192.0.2.51
Reading the zone:
$TTL 300sets the default TTL for records that do not state one.@means the zone's own name,example.com.here, and a blank owner field repeats the previous owner.- The zone apex is that bare name,
example.com. An RRset is all records with the same name and type, for example the two MX records at@. - The SOA record is
ns1.example.com.(the primary name server),hostmaster.example.com.(the administrator's mailbox written as a name, which stands for hostmaster@example.com), then five numbers: serial1(the version, which secondaries compare), refresh3600(how often a secondary checks for a new version, in seconds), retry900(how soon it tries again after a failure), expire1209600(14 days: when a secondary stops answering if it cannot reach the primary) and300(the minimum, which caps how long a name that does not exist is remembered: a resolver uses the smaller of this value and the SOA record's own TTL, RFC 2308, and both are 300 here). - The SRV fields are priority, weight, port, target. SIP has priority 10 for both gateways, weights 60 and 40, so about 60% and 40% of clients pick them. LDAP uses priority 0 for dir1 and 10 for dir2, which is the failover order, and weight 0 on both because each priority level has a single target and so there is no selection to weight (RFC 2782: administrators should use weight 0 when there is no server selection to do).
Checked with BIND's zone checker, set to fail on the two alias errors, and queried:
$ named-checkzone -M fail -S fail example.com example.com.zone
zone example.com/IN: loaded serial 1
OK
$ dig @127.0.0.2 example.com MX +noall +answer +additional
example.com. 300 IN MX 20 mx2.example.com.
example.com. 300 IN MX 10 mx1.example.com.
mx1.example.com. 300 IN A 192.0.2.25
mx2.example.com. 300 IN A 192.0.2.26
$ dig @127.0.0.2 _ldap._tcp.example.com SRV +noall +answer
_ldap._tcp.example.com. 300 IN SRV 10 0 3890 dir2.example.com.
_ldap._tcp.example.com. 300 IN SRV 0 0 3890 dir1.example.com.
The answers come back in rotating order, not sorted. The client sorts by preference or priority, so never rely on the order the server returns. The server also places the target hosts' A records in the additional section (the part of a DNS reply for helpful extra records, here the addresses of the targets), which saves the client a second lookup. The +additional flag in the first dig is what printed those two A lines.
Where the "what can a name point to" rules constrain the design
- MX and SRV targets must not be aliases (CNAME records). Replace
mx2 IN A 192.0.2.26withmx2 IN CNAME mx1.example.com.and the checker rejects the zone:
zone example.com/IN: example.com/MX 'mx2.example.com' is a CNAME (illegal)
zone example.com/IN: not loaded due to errors.
The same thing happens to an SRV target:
zone example.com/IN: _ldap._tcp.example.com/SRV 'dir2.example.com' is a CNAME (illegal)
RFC 2181 section 10.3 says the name used as an MX or NS target must not be an alias, that using one neither works as well as might be hoped, and that the name must have one or more address records; RFC 2782 says the same for an SRV target. Behaviour with an alias target is therefore not guaranteed: it may work for one sender and fail for another, which is the worst kind of bug. Point targets at names that own A or AAAA records. Provider-managed hostnames that are only CNAMEs cannot be used directly as targets: ask the provider for address records, or create your own A records for the same addresses with a documented update process.
2. A CNAME cannot coexist with any other record at the same name (RFC 2181 section 10.1). Because the zone apex must carry SOA, NS and MX records, the apex cannot be a CNAME. The checker confirms it: with @ IN CNAME mx1.example.com. alongside the SOA and NS records, loading fails with example.com: CNAME and other data. A provider hostname that must be reached through a CNAME can be used for www, but not for the apex: the apex needs A or AAAA records, or a provider-specific alias feature that answers with addresses.
3. MX gives no port. Mail from other servers is delivered to the standard SMTP (Simple Mail Transfer Protocol) port, and MX has no way to name a different one, so the MX hosts must accept connections on port 25.
4. SRV only helps clients that look it up. A client that does not implement SRV lookups will never discover port 5070 or 3890. Test with the real SIP and LDAP clients, and configure an explicit host and port wherever one ignores SRV.
5. Keep TTLs sensible. A TTL (time-to-live, how long caches keep an answer) of 300 seconds lets you change priorities within five minutes during an outage. All records of one RRset (same name and type) must have the same TTL (RFC 2181 section 5.2).
Trade-offs and pitfalls
- An alias as an MX or SRV target is illegal even where some software tolerates it. Give each target its own A or AAAA record.
- Weight 0 is the convention when there is no load splitting (RFC 2782 recommends it for readability), which is why the single-target LDAP priorities above use it, while the two SIP gateways at one priority carry real weights.
Design name resolution for an environment with several cloud accounts, Kubernetes clusters and an on-premises network, where internal services need private names and the same names must resolve differently outside. How do queries flow in each direction, and what would you do to avoid inconsistent answers?
Sample Answer
Direct answer
Give every zone exactly one owner, let resolvers find the owner by forwarding on the longest matching domain suffix, and generate any name that must differ inside versus outside from one source of truth. Concretely: a private root such as internal.example.com split into sub-zones (aws., k8s., onprem.), each served by the team and system that owns those hosts; on-premises resolvers forward the cloud and cluster sub-zones to the cloud resolvers' inbound endpoints; the cloud resolvers forward the on-premises sub-zone out through outbound endpoints; each Kubernetes cluster's CoreDNS (the in-cluster DNS server) answers its own names and forwards the rest to the cloud resolver. Split-horizon (same name, different answer inside and outside) is done by two zones of the same name with different audiences, and every record in both comes from one repository and is checked by an automated consistency test.
Two rules do most of the work: one owner per zone, and forwarding on the longest matching suffix. Everything else below (generated views, the checker script, TTL choices) exists to keep those two rules true over time.
Terms used below. A zone is one slice of the DNS namespace that a single set of servers is responsible for. Kubernetes terms: a cluster is a set of machines running containers; a pod is the smallest unit that runs your containers and gets its own IP address; a namespace is a named grouping of cluster objects, and kube-system is the one where Kubernetes keeps its own components; a ConfigMap is a cluster object that holds configuration text other components read (CoreDNS reads its settings from one); cluster.local is the default DNS suffix for names inside a cluster, such as my-service.default.svc.cluster.local. AWS terms: a private hosted zone is a DNS zone that only the VPCs you attach it to can see; an inbound endpoint is a set of IP addresses inside your VPC that resolvers outside it (such as on-premises ones) can send queries to; an outbound endpoint is the set of addresses from which the VPC resolver sends queries out to resolvers elsewhere; a forwarding rule says "queries for this domain go to these resolver addresses". An on-premises resolver does the same with a conditional forwarder. A stub domain in CoreDNS is a block that sends one suffix to a chosen upstream resolver. Longest-suffix match means that when several rules fit a name, the one with the most matching trailing labels wins: api.k8s.corp.lab matches both corp.lab and k8s.corp.lab, and k8s.corp.lab is chosen.
Who asks whom
| Query starts at | Asks first | What sends it onward | Answered by |
|---|---|---|---|
Pod, name ends cluster.local | CoreDNS in the cluster | nothing, CoreDNS owns it | CoreDNS |
| Pod, any other name | CoreDNS | default forward (or a stub-domain block) to the cloud resolver | cloud resolver, or the stub domain's upstream |
| Cloud workload, cloud name | VPC resolver (base plus 2) | the private hosted zone attached to that VPC | the hosted zone |
| Cloud workload, on-premises name | VPC resolver | forwarding rule, out through the outbound endpoint | on-premises resolvers |
| On-premises client, cloud or cluster name | on-premises resolver | conditional forwarder to the inbound endpoint addresses | the VPC resolver behind the endpoint |
| Anyone, public name | their resolver | normal recursion | the public zone's provider |
Query flow in each direction
The cloud side is described with Amazon Route 53's resolver (VPC Resolver, where a VPC is a Virtual Private Cloud, a private network inside a cloud account); other clouds have equivalent inbound and outbound forwarding features.
- Pod to an in-cluster name (
cluster.local): CoreDNS answers it. These names never leave the cluster. - Pod to any other name: CoreDNS forwards to the cloud resolver (in a VPC, the address at the base of the VPC range plus two, for example 10.0.0.2 for 10.0.0.0/16). Kubernetes lets you add stub domains in the
corednsConfigMap inkube-system, for example the block shown after this list, which sends one suffix to a specific upstream. CoreDNS does not support a hostname as the upstream of a stub domain, only IP addresses. - Cloud workload to a cloud name: the VPC resolver answers from the private hosted zone (a private DNS zone attached to chosen VPCs) associated with that VPC. Route 53 picks the most specific matching zone and returns NXDOMAIN (name does not exist) if that zone has no such record. It does not fall back to the public zone.
- Cloud workload to an on-premises name: a forwarding rule for
onprem.internal.example.comsends the query out of an outbound endpoint to the on-premises resolvers over the VPN or private link. If a query matches several rules the most specific one wins, and a resolver rule overrides a private hosted zone of the same name. Rules can be shared across accounts, so one set of rules serves all accounts. - On-premises client to a cloud name: the on-premises resolver has a conditional forwarder for
aws.internal.example.comandk8s.internal.example.compointing at the cloud resolver's inbound endpoint addresses. Point the forwarders at the inbound endpoint addresses, not at the VPC base-plus-two address: the inbound endpoint is the address set AWS provides for queries coming from your own network. - Anyone to a public name: resolved recursively as usual. The public zone for
example.comstays at its public provider.
The stub-domain block from flow 2 (the format Kubernetes documents for the coredns ConfigMap, whose own example uses consul.local:53 and 10.150.0.1; the suffix here is this design's):
onprem.internal.example.com:53 {
errors
cache 30
forward . 10.150.0.1
}
Reading it line by line: onprem.internal.example.com:53 { opens a block that applies only to queries for that suffix, arriving on port 53. errors makes CoreDNS log the errors it hits to standard output. cache 30 keeps answers for at most 30 seconds. forward . 10.150.0.1 sends every query in this block to the upstream resolver at 10.150.0.1 (the . means all names in the block's zone; the address is illustrative). Queries for any other suffix do not match this block and take CoreDNS's default path.
Making the longest-suffix rule concrete (executed)
Unbound configured with two forward zones, corp.lab to one authoritative server and k8s.corp.lab to another, chose by suffix. Because the lab upstreams sit on loopback, the server: block needs do-not-query-localhost: no; without it Unbound refuses to use 127.0.0.1 upstreams and answers SERVFAIL. In production the forward addresses are real resolver addresses and that line is not needed.
server:
do-not-query-localhost: no
forward-zone:
name: "corp.lab"
forward-addr: 127.0.0.1@5301
forward-zone:
name: "k8s.corp.lab"
forward-addr: 127.0.0.1@5302
The server on port 5301 held corp.lab with ldap IN A 10.0.0.20 and www IN A 10.0.0.80; the server on port 5302 held k8s.corp.lab with api IN A 10.96.0.1. dig @127.0.0.1 +short ldap.corp.lab A returned 10.0.0.20 (from the parent zone's server) and dig @127.0.0.1 +short api.k8s.corp.lab A returned 10.96.0.1 (from the cluster's server). The lab uses corp.lab; in production, build the private root under a domain you own so it cannot collide with a name someone else registers.
Split-horizon without inconsistency
The same name can legitimately differ by audience. In BIND this is a view per audience; two views of corp.lab, with the internal one matched by source address, answered www.corp.lab differently from the same server:
view "internal" { match-clients { 127.0.0.1; }; zone "corp.lab" { type primary; file "/z/db.corp.int"; }; };
view "external" { match-clients { any; }; zone "corp.lab" { type primary; file "/z/db.corp.ext"; }; };
The two zone files were identical except for one line, www IN A 10.0.0.80 in db.corp.int and www IN A 203.0.113.80 in db.corp.ext. Queried from 127.0.0.1 it returned 10.0.0.80; queried from 127.0.0.2 (dig -b 127.0.0.2 @127.0.0.2, after ip addr add 127.0.0.2/32 dev lo so that BIND listens on that address) it returned 203.0.113.80. In Route 53 the same pattern is a public and a private hosted zone with the same name. The private one wins for the VPCs it is attached to, and a name missing from it yields NXDOMAIN inside, so every public name the inside also needs must be copied into the private zone.
What I do to avoid inconsistent answers
- One owner per name. Each sub-zone has one authoritative system. A parent zone delegates with NS records (private hosted zones support NS delegation of a subdomain); nobody duplicates a sub-zone into a second place.
- Same rules everywhere. The forward rules on every resolver (on-premises, each VPC, each cluster) come from one definition, because an environment missing the
k8srule falls into the parent zone's rule and gets NXDOMAIN from a server that never held the name. - Generate both split-horizon views from one record source. Internal and external records live in the same repository and are rendered into each zone, so a name cannot exist in one view only by accident.
- Test it continuously. A checker script queries each resolver for a list of expected answers and flags any mismatch:
#!/usr/bin/env bash
# usage: dnscheck.sh expected.txt resolver [resolver...] (resolver = host or host:port)
exp=$1; shift; rc=0
while read -r name type want; do
[[ -z $name || $name == \#* ]] && continue
for r in "$@"; do
host=${r%%:*}; port=53; [[ $r == *:* ]] && port=${r##*:}
got=$(dig @"$host" -p "$port" +short +time=2 +tries=1 "$name" "$type" 2>&1 | grep -v "^;;" | sort | paste -sd, -)
if [[ $got == "$want" ]]; then s=ok; else s=MISMATCH; rc=1; fi
printf '%-9s %-22s via %-15s want=%-12s got=%s\n' "$s" "$name" "$r" "$want" "${got:-<empty>}"
done
done < "$exp"
exit $rc
The expectations file expected.txt has one name type want line per check:
# name type want
ldap.corp.lab A 10.0.0.20
api.k8s.corp.lab A 10.96.0.1
www.corp.lab A 10.0.0.80
Run as bash dnscheck.sh expected.txt 127.0.0.1:53 127.0.0.1:5354, with an on-premises resolver (the Unbound configuration above, port 53) and a cloud-side resolver (a second Unbound on port 5354 with only the corp.lab forward zone, built without the k8s rule), it printed:
ok ldap.corp.lab via 127.0.0.1:53 want=10.0.0.20 got=10.0.0.20
ok ldap.corp.lab via 127.0.0.1:5354 want=10.0.0.20 got=10.0.0.20
ok api.k8s.corp.lab via 127.0.0.1:53 want=10.96.0.1 got=10.96.0.1
MISMATCH api.k8s.corp.lab via 127.0.0.1:5354 want=10.96.0.1 got=<empty>
ok www.corp.lab via 127.0.0.1:53 want=10.0.0.80 got=10.0.0.80
ok www.corp.lab via 127.0.0.1:5354 want=10.0.0.80 got=10.0.0.80
and exited 1, so a pipeline stops. Run it from each environment against its own resolver. The MISMATCH row is the one to act on: the resolver on port 5354 returned nothing (<empty>) for the k8s name because it lacks the k8s forwarding rule.
Reading the script: exp=$1; shift; rc=0 stores the expectations file name, drops it from the argument list so "$@" holds only resolvers, and starts the exit code at 0. while read -r name type want reads each line of the file as three fields. [[ -z $name || $name == \#* ]] && continue skips blank lines and lines starting with # (comments). host=${r%%:*} cuts everything from the first : onward, leaving the host; port=53; [[ $r == *:* ]] && port=${r##*:} keeps 53 unless the resolver was written host:port, in which case ${r##*:} keeps everything after the last :. For 127.0.0.1:5354 these give 127.0.0.1 and 5354. In the dig line, +short prints only the answer data, +time=2 +tries=1 limits the wait to two seconds and one attempt, grep -v "^;;" drops dig's error lines (they start with ;;), sort orders multi-address answers, and paste -sd, - joins the lines into one comma-separated string so it can be compared with the expected value. The if compares, sets s to ok or MISMATCH and sets rc=1 on any mismatch, printf prints the aligned row, ${got:-<empty>} prints <empty> when nothing came back, and done < "$exp" feeds the file to the loop. exit $rc returns 1 if any row mismatched.
- Short TTLs on dynamic names, longer on stable ones. Cluster service names change often (the Kubernetes example above caches for 30 seconds); stable infrastructure names can be longer. Mismatched TTLs can also make two environments disagree for a few minutes even when the data is identical.
Trade-offs
- A central resolver layer is one more path to keep highly available; give each endpoint at least two addresses in different zones or sites.
- Split-horizon doubles the places a record can be wrong. If the inside and outside answers do not need to differ, publish one name and expose it only on private addresses behind access control.
- If everything is in a single cloud with no on-premises network, drop the on-premises forwarders and keep the rest.
What problem does DNSSEC solve and what does it not? Describe how a validating resolver decides an answer is genuine, and what operational risks come with turning it on.
Sample Answer
Direct answer
DNSSEC (DNS Security Extensions) adds digital signatures to DNS data so a validating resolver can prove that an answer really came from the zone's owner and was not altered in transit, including proof that a name or record type does not exist (RFC 4033: data origin authentication, data integrity and authenticated denial of existence). It does not encrypt anything, does not hide which names you query, does not stop denial-of-service attacks, and does not protect the hop between a client and its resolver unless the client validates too. A validating resolver (a resolver that checks signatures instead of trusting what it receives) decides an answer is genuine by following a chain of trust, where each zone vouches for the zone below it, starting from a trust anchor (a key the resolver is configured to trust without proof, normally the root zone's key) down through each delegation, then checking the signature on the answer. The operational risk is that a mistake does not degrade service: a broken chain makes validating resolvers return SERVFAIL, so the domain looks down. Start with the chain of trust, what SERVFAIL means when a link breaks, and the +cd test that separates chain problems from data problems; the sections after the lab run go further into proofs of non-existence and key-rollover timing.
How validation works, link by link
Four record types carry it:
DNSKEY: a zone's public keys. A key with flags 257 is a key-signing key (KSK, the key referenced from the parent); flags 256 marks a zone-signing key (ZSK, used to sign ordinary records). Two keys exist so that the key the parent has to know about changes rarely, because changing it means a request to the parent, while the key that signs the many ordinary records can be replaced often by you alone. One key can do both jobs, and the lab zone below has just one.RRSIG: a signature over one record set, naming the key tag of the signing key (a 16-bit number that labels a key so a signature can say which key made it) and a validity window (inception and expiration times).DS(delegation signer): a hash of a child zone'sDNSKEY, published in the parent zone. It is the link between levels.NSEC(next secure) orNSEC3(its hashed variant): records that prove non-existence.
The resolver's steps, in order (RFC 4035 section 5):
- Start from the trust anchor, the root's key (or its
DS), configured in the resolver. - Fetch the root
DNSKEYset and check it matches the anchor and is signed. - For the top-level domain, fetch the
DSfrom the root zone (itsRRSIGvalidates under the root key). Fetch the TLDDNSKEYset and check that a key hashes to thatDS. - Repeat for the domain:
DSin the TLD zone,DNSKEYin the domain, hash match. - Check the
RRSIGon the answer against the domain'sDNSKEY. - If every link holds, the resolver sets the AD (authenticated data) flag in its reply (RFC 4035 section 3.2.3). If a link in a signed chain fails, it must return
SERVFAIL(server failure, response code 2) unless the client set the CD (checking disabled) bit (RFC 4035 section 5.5).
A zone without a DS in its parent is an "island of security" (RFC 4033: a signed zone with no chain up to the parent): treated as unsigned, not as failed.
Lab run
Lab: BIND 9.20 with dnssec-policy default signing a made-up root, test and example.test (no real DNSSEC infrastructure is involved), Unbound 1.26 validating with the root DS as trust anchor. Each DS was copied into its parent zone with dnssec-dsfromkey.
$ dig +norecurse +noall +answer @127.0.0.10 test DS
test. 86400 IN DS 55972 13 2 6F13C6A78B77107132B3172DFC06CE03F0DBA9BDAFAA8734FB08FEEA BEF30366
$ dig +norecurse +noall +answer @127.0.0.11 example.test DS
example.test. 86400 IN DS 15733 13 2 1FF321FF44D0C88A116406914EA6B5EB1EA146D84DB07D0DDF30465A D36AAEDC
$ dig +norecurse +noall +answer @127.0.0.12 example.test DNSKEY
example.test. 3600 IN DNSKEY 257 3 13 7/8gpSIJyX8TEDF+fYb+N1qPkoKSpugsYQeX8ag+hBxiGTDrSefBZVGq imlvQv5El0O803rKKjpJONr1AL7pvQ==
$ dig +dnssec +noall +comments +answer @127.0.0.20 web.example.test A
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 20431
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags: do; udp: 1232
;; ANSWER SECTION:
web.example.test. 300 IN A 192.0.2.20
web.example.test. 300 IN RRSIG A 13 3 300 20261019193102 20261006062048 15733 example.test. GsdD5zujGnG9Vh7NQ/HntTJP6mGfhIqOKhDeSBOK+1bpM9Brw5SjrA6W ZZe72oEZCcKRD/Vb9zWnOMzdHt6PAQ==
Read each line field by field. The DS line is: owner name example.test., TTL 86400, class IN, type DS, then four fields: key tag 15733 (which key it points at), algorithm 13 (ECDSA P-256 with SHA-256, IANA registry), digest type 2 (SHA-256, the hash used) and the long hex string, which is the hash of the child's public key. The DNSKEY line is: name, TTL 3600, IN, DNSKEY, flags 257, protocol 3 (always 3), algorithm 13 and then the public key in base64. The RRSIG line is: name, TTL, IN, RRSIG, then the type it covers (A), algorithm 13, labels 3 (the number of labels in the signed name, which lets a resolver spot wildcard answers), original TTL 300, signature expiration, signature inception, key tag 15733, signer name example.test. and the base64 signature. In short: the DS for example.test has key tag 15733, algorithm 13 and digest type 2. The DNSKEY published by the domain is flagged 257, and the RRSIG over the A record names key tag 15733 as the signer: the key the parent vouched for signed the answer. The resolver's reply carries ad. The same tag, 15733, appears in the DS and the RRSIG because this zone has a single key doing both jobs. The RRSIG window runs from 20261006062048 (inception) to 20261019193102 (expiration), written year, month, day, hour, minute, second in UTC, about two weeks: signatures expire, which is why signing must be automated.
Authenticated denial of existence
A signed "no" is needed too: otherwise an attacker could answer "that name does not exist" for a real name and nobody could tell. The NSEC record is how a signed zone proves a name is absent.
$ dig +dnssec +noall +comments +authority @127.0.0.20 nope.example.test A | grep -E 'status|^;; flags|IN\s+NSEC\s'
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 52890
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 0, AUTHORITY: 6, ADDITIONAL: 1
example.test. 120 IN NSEC mail.example.test. A NS SOA MX TXT RRSIG NSEC DNSKEY TYPE65534
mail.example.test. 120 IN NSEC ns1.example.test. A RRSIG NSEC
The filter keeps only the status, flags and NSEC lines; the six authority records are the SOA, the two NSEC records, and an RRSIG for each. TYPE65534 in the apex type list is a BIND-specific bookkeeping type from its automatic signing, not part of the chain. An NSEC record says "no names exist between this owner and the next name". mail.example.test to ns1.example.test covers the alphabetical gap where nope would sort, so nope.example.test cannot exist; the apex NSEC covers the position where a wildcard would sit, ruling out a wildcard match. Here is why: a wildcard record *.example.test would answer for any name that has no record of its own, and in the sorted order of names * sorts before letters, so a wildcard would sit right after the apex name example.test, in the gap the apex NSEC says is empty (it jumps straight to mail.example.test). With nothing in that gap, no wildcard exists that could have answered for nope. The six records are signed, so the denial is itself authenticated. NSEC3 (RFC 5155) hashes the names so the chain cannot be trivially walked to list the zone (RFC 5155 notes that offline dictionary attacks against the hashes remain possible, only more expensive); RFC 9276 says to use zero extra iterations and recommends against a salt.
Operational risks
| Risk | What happens | Mitigation |
|---|---|---|
Wrong or stale DS at the parent | Validators return SERVFAIL for the whole domain | Change the DS only through a staged rollover; monitor with a validating resolver |
| Signatures expire | Same SERVFAIL once the window closes | Automatic re-signing plus an alert on days to expiry |
| Key rollover mistakes | Cached old DNSKEY or DS meets new signatures | Respect RFC 7583 timings below |
| Provider migration | New provider cannot sign for the old DS | Plan the DS change and the overlap before moving |
| Larger responses | Signed answers need EDNS and may fall back to TCP; firewalls that block it break resolution | Allow TCP/53 and EDNS |
Heavy NSEC3 | Extra CPU on resolvers | Zero iterations (RFC 9276) |
The failure, reproduced
Rebuilding the lab with the first hex digit of example.test's DS digest changed to 0 (so the parent vouches for a hash that matches no key): (the two blocks below are trimmed to the header, flags and answer lines; dig also prints the OPT pseudosection, the question section and query statistics)
$ dig @127.0.0.20 www.example.test A
;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 7632
;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1
$ dig +cd @127.0.0.20 www.example.test A
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 10309
;; flags: qr rd ra cd; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
www.example.test. 300 IN CNAME web.example.test.
web.example.test. 300 IN A 192.0.2.20
The data is intact; only validation fails. Asking with +cd and getting an answer while the plain query gets SERVFAIL is the diagnostic: the problem is the chain (DS mismatch or expired signature), not the zone content or reachability. Rolling the DS back, or removing it (which makes the zone an unsigned island), restores resolution once the parent's DS TTL (86400 seconds here) lapses in caches.
Rollover timing
RFC 7583 gives the timings. For a pre-publish ZSK rollover, the new key must sit in the DNSKEY set for Ipub = Dprp + TTLkey before it signs anything, where Dprp is the delay for the change to reach every authoritative server and TTLkey is the DNSKEY TTL. With TTLkey = 3600 s and Dprp = 300 s that is 3900 s (65 minutes). A KSK rollover by the double-DS method publishes the new DS in the parent first, waits for the parent's caches (the lab's DS TTL is 86400 s, one day) before changing the DNSKEY, then removes the old DS only after the child's DNSKEY caches expire. The reason for that order is that the parent's DS is the one link you do not control: while both DS records are there, a cache holding either one still finds a matching key, so no cache is ever left without a link. Plain-language reading of the ZSK formula: Ipub is the publication interval and Dprp the propagation delay, and a resolver may still hold the old DNSKEY set for up to TTLkey after the new one reaches every server, so the new key must have been visible for Dprp + TTLkey before any signature relies on it.
Pitfalls
- Treating DNSSEC as encryption or as a fix for a compromised registrar account: an attacker who controls the registrar can publish a new
DSas well. - Enabling it without a validating test resolver and an expiry alert. The first sign of trouble is user reports.
- Forgetting the
DS. Signing the zone without publishing theDSgives none of the protection (an island of security); publishing theDSbefore the zone is signed breaks it.
What is split-horizon DNS, why would an organisation run it, and how do conditional forwarders fit in? Sketch a design where internal-only services and public services share one domain securely.
Sample Answer
Direct answer
Split-horizon DNS (also called split-brain or split-view DNS) means the same domain name answers differently depending on who asks: internal clients get internal addresses and internal-only names, outside clients get only the public view. Organisations run it so internal users reach a service by its private address, and so that internal-only hosts such as a database name never appear in public DNS. A conditional forwarder is a rule on a resolver: "for queries inside zone X, send them to these servers instead of walking from the root". It is how the internal resolver reaches a partner's private zone, or sends public names to the right place.
My recommendation for one shared domain: keep internal-only services under their own subdomain (for example int.example.test) that is served only by internal servers, so the public servers never hold those names at all, and use split horizon only for names that must resolve differently from inside, such as www. Where an internal-only name has to live in the shared zone itself, views are the tool, and the design below handles that case. A view is a software boundary on one server; separate servers are a network boundary, and I prefer the network boundary when the estate can afford it.
The design
| Name (all inside the one shared zone) | Public view (outside) | Internal view (corporate ranges) | Why |
|---|---|---|---|
www.example.test | 192.0.2.20 | 10.0.5.20 | Same name, internal users skip the public edge |
mail.example.test | 192.0.2.30 | 192.0.2.30 | Public service, identical in both views |
db.example.test | does not exist (NXDOMAIN) | 10.0.9.5 | Internal only, never published |
Rules that make it secure:
- Recursion only for internal clients. The public side answers authoritatively (from zone data it owns) and refuses recursion (going off to look up other names for the asker); otherwise it becomes an open resolver that outsiders can abuse for amplification, meaning they send small queries with a forged source address and your server sends large answers at the victim.
- The internal view is a superset. Every public name internal users need must also exist in the internal view; if the internal view serves
example.testit is authoritative for the whole zone (it answers from its own copy and never asks anyone else about names in that zone), and a missing public record gives internal clientsNXDOMAIN("no such name") for it. - No transfers to strangers.
allow-transfer { none; }unless a named secondary needs a copy. A zone transfer hands another server a full copy of the zone, so an open transfer lets anyone download every name in it. - Do not treat DNS views as access control. A name that is hidden is not a name that is protected; the database still needs authentication and network controls.
- DNSSEC and views (applies only if the public zone is signed). DNSSEC adds cryptographic signatures to DNS data. A signed zone's parent publishes a DS record (delegation signer), a fingerprint of the zone's signing key, which tells every validating resolver "this zone is signed, so insist on valid signatures". An internal validating resolver that gets unsigned internal answers for the same name therefore cannot verify them, marks them bogus and returns
SERVFAIL(the generic "server failure" error) instead of the data. Sign the internal view with the same keys, or tell the internal validators to treat the zone as insecure (Unbound'sdomain-insecure, an option that skips validation for the named zone and passesunbound-checkconf).
BIND views, executed in a lab
This configuration passes named-checkconf and was run on BIND 9.20. The client address 127.0.0.2 stands in for a corporate client; 127.0.0.1 stands in for an outsider:
acl corp { 127.0.0.2; 10.0.0.0/8; };
options {
directory "/etc/bind";
listen-on { 127.0.0.12; };
listen-on-v6 { none; };
pid-file "/run/named/named.pid";
dnssec-validation no;
allow-transfer { none; };
};
view "internal" {
match-clients { corp; };
recursion yes;
allow-recursion { corp; };
zone "example.test" { type primary; file "/etc/bind/internal.example.test.zone"; };
zone "partner.test" { type forward; forward only; forwarders { 198.51.100.53; }; };
};
view "external" {
match-clients { any; };
recursion no;
zone "example.test" { type primary; file "/etc/bind/external.example.test.zone"; };
};
The external zone file has www A 192.0.2.20 and mail A 192.0.2.30 (plus ns1 and the SOA/NS); the internal one has www A 10.0.5.20, mail A 192.0.2.30 and db A 10.0.9.5. Reading the configuration line by line:
acl corp { ... };names a list of addresses (an access control list, ACL) that other lines can refer to ascorp.listen-on { 127.0.0.12; };makes the server answer only on that address;listen-on-v6 { none; }turns off IPv6 listening;pid-fileis where the process records its ID.dnssec-validation no;switches signature checking off, a simplification for the lab.allow-transfer { none; };refuses zone transfers to everyone.view "internal" { ... };opens a view: a named set of zones and settings that only some clients see.match-clients { corp; };selects which clients get this view;recursion yes;withallow-recursion { corp; };lets exactly those clients have the server look up other names for them.zone "example.test" { type primary; file "..."; };makes this server the primary (owner of the master copy) for the zone, read from that file. The two views use two different files for the same zone name, which is the split.zone "partner.test" { type forward; forward only; forwarders { ... }; };is the conditional forwarder, explained below.view "external"withmatch-clients { any; };is the catch-all for everyone else, withrecursion no;.
Views are evaluated in order and the first match-clients that matches wins, so the narrow internal view must come first. (In the executed lab the partner.test forwarder pointed at the server itself on a spare port; the documentation-range address above is the production shape and passed named-checkconf.)
$ dig -b 127.0.0.2 +short @127.0.0.12 www.example.test A # status: NOERROR
10.0.5.20
$ dig -b 127.0.0.2 +short @127.0.0.12 db.example.test A # status: NOERROR
10.0.9.5
$ dig -b 127.0.0.1 +short @127.0.0.12 www.example.test A # status: NOERROR
192.0.2.20
$ dig -b 127.0.0.1 +short @127.0.0.12 db.example.test A # status: NXDOMAIN
The outsider gets the public address for www and NXDOMAIN for db; the internal client gets both internal answers.
The conditional forwarder
The zone "partner.test" { type forward; forward only; ... } statement in the internal view is the BIND form: queries for partner.test go to the partner's resolver and nowhere else. forward only means no fallback to iterating from the root, so a dead forwarder produces SERVFAIL for that zone only. A lab run shows the difference with a zone-level forward to a dead address (198.51.100.53) and the answer held by a lab root server in the hints file: forward only returned status: SERVFAIL after Query time: 10001 msec, while forward first returned status: NOERROR, www.partner.test. 300 IN A 198.51.100.77, after Query time: 1204 msec, because it fell back to iterating. The Unbound form of the same rule, which passes unbound-checkconf:
forward-zone:
name: "partner.test."
forward-addr: 198.51.100.53
Trade-offs and pitfalls
- Views vs separate servers. Views are cheapest and fine for a small estate; they put internal data on a server that also answers the internet, so one mistake in an access control list (ACL) leaks the internal zone. Separate servers (a managed provider for the public zone, internal servers inside the network) remove that failure mode at the cost of keeping two zones in step.
- Drift. The same name in two zones means two places to edit. Generate both from one source and review diffs; a public record added to only the public file is missing for internal users.
- Forwarder dependencies. Each conditional forwarder is a dependency on someone else's resolver and the VPN to reach it. Monitor it, and restrict it to the specific zones needed.
- What would change the choice: if no name needs to differ between inside and outside, drop split horizon entirely and delegate an internal subdomain to internal servers; it is simpler and the public servers cannot leak what they never held.
Unlock Full Question Bank
Get access to all 39 DNS, DHCP, and Name Resolution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.