DNS, DHCP, and Name Resolution Questions
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
How can an attacker get a resolver to cache a forged answer, and what makes the attack hard today? What would you still worry about?
Sample Answer
Direct answer
An off-path attacker (one who cannot see the victim's traffic, as opposed to an on-path attacker, who sits on the route and can) tricks a recursive resolver into caching a forged answer by racing the real authoritative server: the resolver sends a query over UDP, and the attacker floods it with fake replies that must match the query's 16-bit transaction ID, its source port, its question and the server address it was sent to. A fake that arrives first and matches is accepted and cached for its TTL. The 2008 attack disclosed by Dan Kaminsky (CERT note VU#800113, 8 July 2008) made this practical against resolvers that used a fixed source port, and what makes it hard today is the combination of a random ID and a random source port on every query (RFC 5452), strict matching and bailiwick rules (a resolver accepts only data that belongs to the zone it asked about), and above all DNSSEC, which makes a forged answer fail signature validation. What I would still worry about: unsigned zones, attackers who can see or influence the path, side channels that reveal the source port, and compromises that produce valid-looking data, such as a hijacked registrar account.
The attack, in order
- The attacker makes the victim resolver ask a question whose answer it does not already hold, for instance by triggering a lookup (many resolvers answer anyone who can reach them, or a victim's own user or mail server can be induced to look the name up).
- The resolver sends a query to the authoritative server with a transaction ID and source port.
- The attacker sends many spoofed replies purporting to be from that server's address. Each carries a guessed ID and port. Only one reply per race can be accepted.
- If a forged reply matches and beats the real one, the resolver caches the attacker's data until its TTL (time-to-live) expires, and every client of that resolver is served it.
Why Kaminsky's variant mattered. Before it, attackers forged the answer for a real name such as www.example.test; after one failed race the real answer was cached and the attacker had to wait out the TTL. Kaminsky's insight was to ask for random names that cannot be cached, such as a1b2c3.example.test, so every attempt is a fresh race, and to put the poison not in the answer but in the authority and additional sections of the forged reply: an NS record for example.test and a glue A record pointing at the attacker's server. Here is an illustrative forged reply for a query about a1b2c3.example.test (203.0.113.66 is a documentation address standing in for the attacker's server):
;; QUESTION SECTION:
;a1b2c3.example.test. IN A
;; ANSWER SECTION: (empty: the name does not exist)
;; AUTHORITY SECTION:
example.test. 86400 IN NS ns.example.test.
;; ADDITIONAL SECTION:
ns.example.test. 86400 IN A 203.0.113.66
A DNS reply has an answer section (the records that answer the question), an authority section (which servers are responsible) and an additional section (extra records offered to save a lookup, such as the address of a name server). The NS record there says "example.test is served by ns.example.test", and the additional-section A record (glue, the address of a name server whose name sits inside the zone it serves) says where that server is. If the resolver caches them, every later query for any example.test name goes to the attacker. The glue is in-domain (it belongs to the zone being asked about), so a rule that rejects out-of-domain data does not stop it; what stopped it was making each race unlikely to be won.
What makes it hard: the arithmetic
RFC 5452 requires a resolver to match the query ID, the source and destination addresses, the destination port against the query's source port, and the question, and to pick the ID and source port unpredictably across the full range available (ports 53 or 1024 and above). A 16-bit ID gives 65,536 values; random ports multiply that. The ID and port randomisation is the main defence, so the arithmetic is worth doing. A pinned calculation, where the attacker lands 1000 distinct forged (ID, port) guesses in each race and the port range is Linux's default ephemeral range (the block of high port numbers the operating system hands to outgoing connections) 32768 to 60999:
import math
ids = 65536 # 16-bit DNS query ID
ports = 60999 - 32768 + 1 # Linux default ephemeral range 32768-60999
forged = 1000 # forged replies that land inside one race window
print(f"ports in range: {ports}")
print(f"ID x port space: {ids * ports} ({math.log2(ids * ports):.1f} bits)")
for label, space in (("fixed port ", ids), ("random port", ids * ports)):
p = forged / space
races = math.log(0.5) / math.log(1 - p)
print(f"{label}: p per race = {p:.3e}, races for 50% = {races:.0f}")
ports in range: 28232
ID x port space: 1850212352 (30.8 bits)
fixed port : p per race = 1.526e-02, races for 50% = 45
random port: p per race = 5.405e-07, races for 50% = 1282469
Where the races formula comes from: if one race is won with probability p, it is lost with probability 1 - p, so losing n races in a row has probability (1 - p)^n. The attacker has a 50% chance of at least one win when that equals 0.5, and solving (1 - p)^n = 0.5 for n gives n = ln(0.5) / ln(1 - p), which is the math.log(0.5) / math.log(1 - p) line. Check by running it forward: with the fixed port, 1 - (1 - 0.01526)^45 is 0.499, and with the random port 1 - (1 - 5.405e-07)^1282469 is 0.500 (both checked in Python).
With a fixed port, 1000 forged replies per race win about 1.5% of the time and about 45 races give a 50% chance, which Kaminsky's fresh-name trick makes cheap. With a random port, a race wins about 5.4 in ten million, and it takes about 1.28 million races for a 50% chance. The port range here is 28,232 ports, about 14.8 bits, so the total is 30.8 bits rather than the 32 you get from a full 64k range; narrowing the allowed range (a firewall rule, or a resolver setting such as Unbound's outgoing-port-avoid that removes ports from use) directly weakens it.
Other layers:
- Bailiwick and glue checks. A resolver should accept data only if it belongs to the zone it asked about (RFC 5452 section 6); Unbound's
harden-glue(defaultyes) trusts glue only within the server's authority. - 0x20 encoding. The resolver randomises the upper and lower case of the query name, for example asking for
eXaMpLe.TeStinstead ofexample.test; an honest server echoes the case back exactly, so a forger must guess the case pattern too, and every letter in the name adds one more bit to guess. Unbound'suse-caps-for-idis an experimental option and defaults tono. - DNSSEC. The real fix: the forged data has no valid signature under the zone's key, so a validating resolver rejects it regardless of how well the forger guessed (a chain of signatures from the root key down to the zone: each level vouches for the key of the level below).
What I would still worry about
| Residual risk | Why randomisation does not cover it |
|---|---|
| Zones and resolvers without DNSSEC | The race is only hard, not impossible; at 30.8 bits an attacker with a fast path still gets there eventually |
| On-path attackers | They see the ID and port, so there is nothing to guess; only DNSSEC (or an encrypted resolver-to-server channel) helps |
| Side channels that leak the port | A side channel is information that leaks through a side effect rather than the reply itself. The 2020 SAD DNS research (UC Riverside) used the global rate limit on outgoing ICMP (Internet Control Message Protocol, the error-message protocol, for example "port unreachable") messages: by probing and watching how many ICMP errors came back, the attacker could tell which source port was open, then guessed only the ID. Listed mitigations: randomise or disable that ICMP behaviour, add secrets such as 0x20 or DNS cookies (a random value the client and server exchange and echo, RFC 7873, which off-path forgers cannot know), deploy DNSSEC, and shorten the window a query stays outstanding |
| Port randomisation undone en route | A NAT or firewall that remaps source ports predictably gives back the entropy (the unpredictability, measured in bits); confirm with a capture that ports seen leaving the network still vary |
| Valid-looking bad data | A registrar or DNS-provider account takeover, or a hijacked route to an authoritative server, produces correctly signed or simply authoritative answers; DNSSEC does not stop a registrar compromise because the attacker can change the DS too |
| Shared resolvers | One poisoned cache serves everyone behind it, so open or lightly protected forwarders raise the blast radius |
What I would do
Run a validating resolver (with DNSSEC on), restrict who may query it, keep the outgoing port range wide, leave harden-glue on, sign my own zones, lock down registrar accounts with multi-factor authentication and registry locks where offered (a registry lock is an extra step in which the registry itself, not just the registrar, must approve changes to a domain's delegation), and capture traffic once to verify source ports actually vary leaving the network.
Your site has about 1,000 clients and a single DHCP server that has just become a single point of failure. What are your options for making DHCP redundant, how do they differ in lease state and failure behavior, and which would you pick and why?
Sample Answer
Direct answer
Use DHCP failover: two servers that replicate lease state, so either can renew any client and the remaining server can issue from the whole pool. At one site with about 1,000 clients, run it in load-balance mode (both answer, so a failure is a non-event and the standby is never untested). If the second server is on the far side of a WAN, run hot-standby (one answers, the other waits). Do not choose a split scope, which gives each server a separate slice and no shared state, and do not count VM-level high availability as DHCP redundancy: it restarts the same single server.
DHCP (Dynamic Host Configuration Protocol) hands out an IP address on a lease: a time-limited right to use it. A client first asks its server to renew at T1, which defaults to half the lease; if that fails it broadcasts to any server at T2, which defaults to 7/8 of the lease (RFC 2131, section 4.4.5). With an 8-hour lease that is T1 = 4 hours and T2 = 7 hours. So a short outage only hurts clients that are new or rebooting, not everyone at once, which is why the single server often goes unnoticed until the day it matters.
A relay (usually the router on a VLAN) turns a client's broadcast request into a unicast to a server on another network. Sites with one DHCP server for many VLANs rely on relays, and with a redundant pair every relay must be told about both servers.
The options compared
| Option | Lease state | When one server dies | Main weakness |
|---|---|---|---|
| Split scope: two independent servers, each owns a slice (Microsoft cites 70% / 30% as typical) | None shared; two lease databases | The remaining server can only hand out its own slice, and clients cannot keep their old address | Needs each slice to cover the whole peak; unusable when the scope is already highly used (Microsoft's own caveat) |
| Failover, load-balance | Replicated between the two | The remaining server answers everyone once it knows the partner is down | Two servers must stay in sync and keep clocks within a minute (Windows checks this when creating the relationship) |
| Failover, hot-standby | Replicated | Kea: the standby stays silent until it declares the partner down, then answers. Windows: the standby renews existing clients and gives new clients only its reserve addresses, and uses the whole pool only after the MCLT | Kea: a new-client gap while it decides (shown below). Windows: the reserve limits what it can give out early |
| Windows failover cluster | One database on shared storage | The cluster restarts the role on another node | The shared storage becomes the single point of failure (Microsoft's description) |
| VM-level HA (hypervisor restarts the VM) | One database, one server | The same server boots elsewhere | Does not detect a hung or crashed DHCP service while the VM is up, unless you add application monitoring; one configuration, one patch level, one database to corrupt |
Failover is supported between exactly two servers, and Windows DHCP failover covers IPv4 scopes only.
Lease state and the failover states
A normal pair looks like this. Both servers hold the same lease table and talk over a control channel, a TCP connection between the two servers that carries lease updates and heartbeats (short "I am alive" messages). In load-balance mode either server answers a new client, 50/50 by default. In hot-standby mode only the active server answers and the standby keeps its copy of the table current. The vocabulary below describes what happens when that conversation stops.
Windows adds a Maximum Client Lead Time (MCLT). Microsoft describes it as the maximum amount of time one server can extend a lease for a client beyond the time known by the partner server. In numbers (illustrative): server B's copy says a client's lease ends at 10:00. With an MCLT of 1 hour, server A may not promise that client more than 11:00 until B has been told about the longer lease. If A dies at that moment, B only has to wait until 11:00 to be sure no client can still be using an address it believed was free, so it never hands one address to two clients. Its default is 1 hour (Add-DhcpServerv4Failover -MaxClientLeadTime).
A silent partner looks the same whether it crashed or only the link between the two servers broke. If each server simply assumed the other was dead, both would issue addresses from the whole pool and could give the same address to two clients. So when the servers cannot talk, each moves to COMMUNICATION INTERRUPTED (it keeps serving, with limits); PARTNER DOWN is the state, entered by an explicit decision or a timer, in which the remaining server takes over. Windows has -AutoStateTransition (default false, so no timer triggers it unless you enable it) and -StateSwitchInterval for an automatic move from COMMUNICATION INTERRUPTED to PARTNER DOWN. In hot-standby, -ReservePercent (default 5) is the share of free addresses the standby keeps to lease to new clients after a failure. The reserve applies to hot-standby mode only. Take a /21 (10.20.0.0 to 10.20.7.255, 8 x 256 = 2,048 addresses) with the pool 10.20.0.50 to 10.20.7.200. The pool leaves out 50 addresses at the bottom (10.20.0.0 to 10.20.0.49) and 55 at the top (10.20.7.201 to 10.20.7.255), so it holds 2,048 - 50 - 55 = 1,943 addresses. With 1,000 clients, 943 are free, and 5% of 943 is 47.15, about 47. The standby may give new clients only that reserve while it is in COMMUNICATION INTERRUPTED, and after PARTNER DOWN it still waits one MCLT (1 hour by default) before it may use the whole pool (Microsoft Learn, DHCP failover overview). So the reserve must last through the switch interval plus the MCLT. If more than about 47 new clients arrive in that window, the standby runs out of addresses to give and refuses new clients while it keeps renewing existing ones. How long 47 addresses last depends on the arrival rate (illustrative rates): at 5 new clients per minute 47 / 5 = 9.4 minutes, and at 20 per minute 47 / 20 = 2.35 minutes. Even the slow rate would need 5 x 60 = 300 addresses to cover a 1 hour MCLT, so on a busy site raise -ReservePercent or shorten the MCLT; a shorter switch interval alone does not close the gap.
What failure behavior looks like (executed)
Kea (the ISC DHCP server) has a high-availability hook with the same modes. Two Kea servers in hot-standby, primary 172.31.1.11 and standby 172.31.1.12, on one network, lab timers shortened (heartbeat-delay 1000, max-response-delay 4000 and max-ack-delay 1000 milliseconds, max-unacked-clients 1; Kea's documented defaults for max-response-delay, max-ack-delay and max-unacked-clients are 60000, 10000 and 10). The relevant block in each server's config, with this-server-name set to dhcp-a or dhcp-b:
"high-availability": [ {
"this-server-name": "dhcp-a", "mode": "hot-standby",
"heartbeat-delay": 1000, "max-response-delay": 4000,
"max-ack-delay": 1000, "max-unacked-clients": 1,
"peers": [
{ "name": "dhcp-a", "url": "http://172.31.1.11:8000/", "role": "primary" },
{ "name": "dhcp-b", "url": "http://172.31.1.12:8000/", "role": "standby" } ] } ]
What each setting means (defaults from the Kea documentation): this-server-name says which entry in peers is this server; mode picks hot-standby; heartbeat-delay is how often, in milliseconds, a server sends the partner a heartbeat (default 10000); max-response-delay is how long without a successful exchange with the partner before this server decides communication is interrupted (default 60000); max-ack-delay is how long a client may keep trying without its request being answered before this server counts it as an "unacked" client (default 10000); max-unacked-clients is how many such clients are tolerated before this server assumes the partner is offline (default 10). In peers, each partner has a name, the url of its control channel, and a role.
Observed with a client run one at a time. udhcpc is BusyBox's small DHCP client, and no lease, failing is what it prints when it sent its DISCOVER requests and no server made an OFFER:
both up:
udhcpc: lease of 172.31.1.100 obtained from 172.31.1.11, lease time 3600
(172.31.1.100 is listed in the standby's lease table: replicated)
primary stopped:
udhcpc: no lease, failing <- standby silent, partner not yet declared down
udhcpc: no lease, failing <- same
HA_STATE_TRANSITION dhcp-b: server transitions from HOT-STANDBY to PARTNER-DOWN state, partner state is UNDEFINED <- standby's own log, a moment before the client's next retry
udhcpc: lease of 172.31.1.101 obtained from 172.31.1.12, lease time 3600
Two things to carry over to production (the first is Kea's behaviour; a Windows standby in COMMUNICATION INTERRUPTED would lease new clients from its reserve instead of staying silent). The standby gave nothing to the first two clients because it waits for evidence that the primary is gone (unanswered client requests plus lost communication), so new clients see a gap. And the standby already held the lease record for 172.31.1.100, which is what lets it answer that client's renewal. With Kea's default values the standby waits at least 60 seconds of silence and more than 10 clients whose requests the partner did not answer within 10 seconds, so plan on a longer gap than in this lab. Note that kea-dhcp4 -t only checks syntax: it accepts a nonsense role, so a passing config test does not prove the HA settings are valid.
My pick
For a single site, load-balance failover. On Windows:
Add-DhcpServerv4Failover -ComputerName 'dhcp-a.corp.example.com' `
-Name 'hq-failover' -PartnerServer 'dhcp-b.corp.example.com' `
-ScopeId 10.20.0.0 -LoadBalancePercent 50 -MaxClientLeadTime 1:00:00 `
-AutoStateTransition $true -StateSwitchInterval 1:00:00 -SharedSecret 'replace-me'
This was checked for syntax with the PowerShell parser and each parameter against Microsoft Learn; it was not run, since the DhcpServer module needs a Windows server. The 50% split and 1-hour MCLT are the documented defaults; I would start the switch interval at 1 hour, long enough not to flip during a reboot and patch cycle and short enough that new clients are not limited for long. The shared secret turns on message authentication between the partners.
What would change the answer: servers at different sites, so use hot-standby with the remote server as standby; stateful IPv6 (DHCPv6) or three or more servers, which Windows failover does not cover; a requirement that the standby never answer unless needed (hot-standby again).
Pitfalls
Putting both servers on the same host, switch or power feed makes redundancy cosmetic. Relays (routers that forward client broadcasts to a server) must point at both servers, or the standby never sees requests. Test by stopping the active server in a maintenance window and measuring how long a new client waits.
Tell me about a time you debugged a production DNS outage. Use the STAR method: describe the situation, the action(s) you took to identify and mitigate the fault, the tools you used, how you communicated with stakeholders, and what permanent changes you implemented to prevent recurrence.
Sample Answer
Direct answer
Pick a real outage where you personally drove the diagnosis, and tell it in the five parts the question names: the situation, the actions you took to identify and mitigate the fault, the tools you used, how you communicated with stakeholders, and the permanent changes that stopped a repeat. A strong version shows an ordered diagnosis (what you ruled out and why), a mitigation chosen for speed and confirmed on every server before you call it fixed, a communication rhythm, and a fix that changes a process or a check, not just "we were more careful". The example below is a model of the shape, with details you would replace by your own.
Story skeleton, with a worked example
Situation. A company's zone is served by three public name servers at different providers, all fed from a hidden primary (the server where the zone is edited, not listed to the public). The three are secondaries: copies that fetch the zone from the primary by zone transfer. Mid-afternoon, a subset of customers report that the application resolves to an address of a load balancer that was decommissioned earlier that week. Support tickets arrive from several countries, and engineers inside the office see the right address because the office resolver is configured to ask the primary and the up-to-date servers, while customers' resolvers pick among all three public servers, one of which is stale.
Task. As the on-call engineer you own restoring correct resolution for external customers and keeping stakeholders informed. State what "restored" means: every authoritative server returns the current address.
Actions, in the order they were taken.
- Scope the problem. Reproduce from outside using several public resolvers and then ask each authoritative server directly, with recursion off:
dig @ns-a.example.com api.example.com A +norecurse +short
dig @ns-b.example.com api.example.com A +norecurse +short
dig @ns-c.example.com api.example.com A +norecurse +short
Two returned the new address and the third still returned the old one. That separates "DNS is broken everywhere" from "one server is stale". The stale answer explains why only some customers were affected, because resolvers pick among the name servers. To rule out the load balancer and application, test the new address directly, bypassing DNS:
$ curl -s -o /dev/null -w "%{http_code}\n" --resolve api.example.com:8081:127.0.0.1 http://api.example.com:8081/health
200
The 200 (in the lab, against a local test server standing in for the new load balancer) showed the new address served the application, so the fault was in what DNS handed out.
2. Compare the SOA serials on the three servers: dig +short @ns-c.example.com example.com SOA. The serial is the version number inside the zone's SOA record, and a secondary fetches the zone only when the primary's serial is higher. The stale server's serial was lower than the primary's (in a lab, 6 against 7). Its refresh had failed.
3. Find why it was stale. TSIG is a shared-secret signature that the primary checks on every transfer request. The stale secondary's log showed transfers being rejected with a TSIG BADSIG error (bad signature: the secret on the two sides no longer matched), after a key rotation that had updated the primary and two of the three secondaries. A lab reproduction with a deliberately wrong key printed these lines:
primary: request has invalid signature: TSIG k: tsig verify failure (BADSIG)
secondary: refresh: failure trying primary 127.0.0.2#53 (source 0.0.0.0#0): tsig indicates error
The primary's line says it refused the request, the secondary's line says why its refresh failed, and both name the key. Testing a manual transfer with the key from that server's host reproduced the failure.
4. Choose the fastest safe fix. Taking the stale server out of the delegation (the NS records held by the parent zone, here the .com registry, which tell the world which servers answer for the domain) would be slow, because resolvers may cache those NS records for up to 48 hours (a live query for a .com domain returned 172800 seconds). So correct the key on the stale secondary, force a retransfer (rndc retransfer, the BIND command that makes a secondary fetch the zone again now), confirm the serial matches the primary, then re-query all three servers.
5. Handle residual caching. Resolvers that had already cached the old answer would keep it for the record's TTL. With a TTL (time-to-live, how long caches may keep an answer) of 3600 seconds, the worst case for a cache that stored it just before the fix was one hour. That arithmetic is what you tell stakeholders, rather than "fixed", so they know when the last customer recovers.
In your own words. This is a spoken version of about a minute. Replace the details with yours.
"Last year I was the on-call engineer when customers in several countries started reporting that our API resolved to a load balancer we had retired that week. Our own staff saw nothing, which told me it was not the application. I asked each of our three public name servers directly and two returned the new address while one still returned the old one, so I knew it was one stale server and not a global DNS fault. Its zone version number was behind the primary's, and its log showed transfers failing with a bad-signature error after a key rotation. I fixed the key on that server, forced a transfer, checked all three agreed, and then told support that some customers would keep the old answer for up to an hour because of the cache lifetime. We recovered fully within that hour. Afterwards I added an alert for any server whose version number lags, and I put a verification step into the key rotation checklist. What I would do differently is automate the rotation so a human never edits three servers by hand."
First status update. A model of the first message in the incident channel and on the status page:
"[14:20] Investigating. Some customers in several regions are reaching an old address for the API. Internal tools are not affected. We have identified one DNS server serving outdated data and are correcting it. Next update by 14:50."
Communication. Open an incident channel and a status page entry within the first minutes. Post a first update that states what is known, what is not, and when the next update comes. Update at a fixed interval, for example every 30 minutes, even when the update is "no change". Tell support exactly what to say: a customer can recover sooner by flushing their resolver's cache or by waiting up to the TTL. When it is resolved, say so and give the time the last cached copy can expire.
Tools. dig (with +norecurse, +short, +tcp), the name server's logs, the transfer test with a TSIG key, rndc retransfer on BIND to force a transfer, external probes from several regions, and the status page.
Permanent changes.
- A monitor that compares SOA serials across every authoritative server and pages when they differ for longer than the refresh interval. This would have caught the fault before customers did.
- Key rotation turned into a runbook step with a verification command per server, and an expiry alert on each key.
- A synthetic check that resolves the application name through several public resolvers every minute.
- A rule to lower TTLs a day before planned infrastructure changes, so an old answer expires quickly.
- A blameless post-mortem with a timeline, listing the detection delay as a finding.
Result. State what you can honestly measure from your own incident: the time from first report to a correct answer on every authoritative server, and the number of customers or tickets affected. Do not invent figures. If you do not remember them, say how you would find them (ticket system timestamps, resolver logs).
What you would do differently. For example: the rotation was done by hand across servers without a verification step, and the first detection was customer reports. You would automate the rotation and put serial-parity alerts in before the next key change.
Trade-offs and pitfalls in telling the story
- Say "I" for your actions and "we" for team outcomes. An answer that never says what you did sounds like observation.
- Show the dead ends you ruled out (the resolver, the load balancer, the application) and what ruled them out. This shows method.
- Show that restoring correct answers came before the full explanation: the fix was verified on every server first, and why the rotation was missed was left to the post-mortem.
- Avoid blaming a person. The permanent changes should address the process gap.
- Choose a story where DNS was actually the fault, because a story that ends in "it was the application" does not answer this question.
A secondary DNS server has stopped picking up changes made on the primary. How do zone transfers work, what are the usual reasons they fail, and how would you debug and then lock transfers down?
Sample Answer
Direct answer
A secondary (also called slave) keeps a copy of a zone by comparing the serial number in the zone's SOA record (start of authority, the zone's header record) with the primary's, then pulling changes over TCP with a zone transfer. A secondary that stops updating is almost always one of four things: the serial did not increase, the secondary never heard or accepted the NOTIFY (the primary's "something changed" message), TCP port 53 is blocked, or the transfer is rejected by an access list or TSIG (transaction signature) check. Debug by comparing serials on both servers, then testing a transfer by hand from the secondary's own host, and lock transfers down by allowing them only to holders of a TSIG key.
How a zone transfer works
- Serial comparison. The SOA record holds SERIAL, REFRESH, RETRY and EXPIRE (RFC 1035 section 3.3.13). The secondary asks the primary for the SOA every REFRESH seconds. If the primary's serial is greater (compared with sequence-space arithmetic on an unsigned 32-bit number, not plain subtraction; the next paragraph shows what that means), the secondary pulls the zone. A failed refresh is retried after RETRY seconds. If nothing succeeds for EXPIRE seconds, the secondary stops answering for the zone.
What "greater" means (RFC 1982). The serial is a 32-bit counter, so it runs from 0 to 4294967295 and then wraps to 0. A seriali1counts as newer thani2wheni1 > i2andi1 - i2is less than 2^31 (2147483648), or wheni1 < i2andi2 - i1is more than 2^31. Normal case: the primary holds 2024010103 and the secondary 2024010102; the difference is 1, so the primary's is newer and the secondary pulls. Wrap case: the primary's counter has passed 4294967295 and wrapped to 0 while the secondary holds 4294967295; plain comparison says 0 is smaller, but the rule above says 0 is newer (the gap, 4294967295, is more than 2^31). Backwards case: a primary serial of 2023123101 against a secondary's 2024010102 differs by 887001, which is less than 2^31, so the primary's counts as older and is ignored. Jumps larger than 2^31 in one step (for example from 100 to 3000000000) also count as older. All of these were computed with Python against the RFC 1982 rule. This is why a date-counter serial (YYYYMMDDnn) can only be increased by small steps and a lowered serial is silently ignored. - NOTIFY. After a change, the primary sends a NOTIFY message (RFC 1996) to the servers in its notify set (the zone's NS servers plus any
also-notifyaddresses) so they do not wait a full REFRESH. A secondary that receives NOTIFY from an address that is not a known primary for the zone (not one of the addresses in that zone'sprimarieslist on the secondary) ignores it and logs an error (RFC 1996 section 3.10); it does not query the sender, so it learns of the change only when its own REFRESH timer fires. A NOTIFY from a known primary makes the secondary check that primary's SOA at once instead of waiting. NOTIFY is a shortcut. The refresh timer is the safety net. - Transfer. AXFR (full zone transfer) is defined over TCP only (RFC 5936), starts and ends with the zone's SOA record, and a general-purpose implementation is recommended to use TSIG to control it. IXFR (incremental transfer, RFC 1995) sends only the differences since the secondary's serial and falls back to a full transfer when the primary lacks that history.
- Authentication. TSIG (RFC 8945) signs each message with a shared secret and a timestamp, so the primary can accept transfers by key instead of by source address.
The lab used for the evidence below
Primary at 127.0.0.2, secondary at 127.0.0.3, BIND 9.20, zone example.com. The primary has allow-transfer { key xfer-key; }, also-notify { 127.0.0.3; }; the secondary has primaries { 127.0.0.2; } and server 127.0.0.2 { keys { xfer-key; }; }, with a key made by tsig-keygen -a hmac-sha256 xfer-key.
Debug in this order
1. Compare serials on both servers. This tells you at once whether the secondary is behind.
$ dig +short @127.0.0.2 example.com SOA | awk '{print "primary serial",$3}'
$ dig +short @127.0.0.3 example.com SOA | awk '{print "secondary serial",$3}'
dig +short prints just the SOA data, seven space-separated fields: primary server, contact mailbox, serial, refresh, retry, expire, minimum. For example ns1.example.com. hostmaster.example.com. 2024010103 7200 3600 1209600 300. The serial is the third field, which is why awk prints $3. A healthy pair prints the same number twice. With the primary at 2024010103 and a secondary that has not caught up, the lab printed:
primary serial 2024010103
secondary serial 2024010102
A lower number on the secondary means it is behind; the other steps find out why.
Equal serials but a missing record means the change was made without bumping the serial. In the lab, adding late IN A 192.0.2.30 to the primary and reloading without changing the serial left the secondary answering status: NXDOMAIN for late.example.com, because from its point of view nothing had changed. Fix: increment the serial on every change, or generate it (a date-counter such as 2024010102 or a Unix timestamp) in the deploy pipeline.
A serial that moves backwards is ignored. When the primary's serial was set to 2023123101 while the secondary held 2024010102, the secondary kept 2024010102 and never received the new record. Recovery is to raise the serial above the old value, or to follow the documented serial-reset procedure for your server software.
2. Is NOTIFY arriving, and from where? Check the secondary's log. A multi-homed primary (one with more than one IP address) may send NOTIFY from a different address than the one the secondary lists. Before the fix, the lab primary sent from 127.0.0.1 and the secondary logged:
received notify for zone 'example.com'
zone example.com/IN: refused notify from non-primary: 127.0.0.1#50034
Without NOTIFY the secondary only learns of the change at its next REFRESH, which is why this fault looks like "works, but slowly". Fix on the primary: notify-source 127.0.0.2; (and transfer-source 127.0.0.3; on the secondary for its own transfers). After the fix the log shows notify from 127.0.0.2#56137: serial 2024010103 and the transfer starts at once.
3. Is TCP/53 open both ways? UDP/53 may work while TCP is dropped, so SOA checks pass and NOTIFY arrives but the transfer hangs. In the lab, after iptables -A INPUT -p tcp --dport 53 -j DROP on the primary and a serial bump:
primary serial 2024010103
secondary serial 2024010102
$ dig +tcp +time=2 +tries=1 @127.0.0.2 example.com SOA
;; Connection to 127.0.0.2#53(127.0.0.2) for example.com failed: timed out.
The secondary logs notify from 127.0.0.2#56137: serial 2024010103 then Transfer started. and nothing further. Test from the secondary's host with dig +tcp @primary example.com SOA.
4. Is the transfer authorised? Request a full transfer by hand from the secondary's host, with the key:
$ dig @127.0.0.2 example.com AXFR +noall +comments # no key
;; ->>HEADER<<- opcode: QUERY, status: REFUSED, id: 3605
(flags and OPT lines omitted)
; Transfer failed.
$ dig @127.0.0.2 example.com AXFR -y hmac-sha256:xfer-key:<secret>
Without a key the lab primary returned REFUSED (a plain dig ... AXFR without +comments prints only ; Transfer failed., so add +noall +comments to see the status). With a wrong secret the transfer failed and the response carried TSIG ... BADSIG. Other TSIG failures to recognise (RFC 8945): BADKEY (key name or algorithm unknown to the server) and BADTIME (the timestamp is outside the allowed "fudge" window, the number of seconds the two servers' clocks may differ before a signature is rejected; 300 seconds is the recommended value, so the cause is clock drift between the two servers). BADSIG means the signature did not verify (usually a wrong secret). Fix clock sync with NTP (network time protocol), or correct the key name, algorithm and secret on both sides.
5. Force the retry once the cause is fixed. On BIND, rndc retransfer example.com on the secondary pulls the zone immediately. In the lab it exited with status 0 and the secondary moved from 2024010101 to 2024010105. Then re-check serials.
Lock transfers down
The first two items are the controls that decide whether a stranger can pull the zone: refuse transfers by default, and allow them only with a key. The remaining items add layers around those two.
- Set a global
allow-transfer { none; };inoptions, and open it per zone:allow-transfer { key xfer-key; };. In the lab an unsigned request was REFUSED while a signed one returned the zone, with the SOA record appearing as first and last record. - Use a key per secondary,
hmac-sha256or stronger (RFC 8945 discourages SHA-1), and rotate keys on a schedule. - Add a source-address allow-list at the network layer as well: TCP/53 from the secondaries only, and a firewall rule that lets the secondaries' source addresses reach the primary. The network rule limits who can reach the port at all, and the key limits who may pull the zone once they can, so each covers a gap the other leaves.
- Keep the primary hidden (not listed in the zone's NS records, so resolvers never query it and only the secondaries face the internet; the secondaries are the servers you publish).
- Monitor SOA serial parity across all name servers and alert when they differ for longer than REFRESH plus RETRY.
Why open transfers matter
An open AXFR hands out every hostname in the zone, including internal-looking names (vpn, jenkins, staging-db), which saves an attacker the guesswork of enumeration.
Pages load slowly for users on certain ISPs, and you have narrowed it to DNS lookups taking too long. How do you find out whether the delay is in caching, the resolver, the path or the authoritative side?
Sample Answer
Direct answer
Split the lookup path into four hops and measure each one on its own, always from inside an affected ISP (an internet service provider) network: a volunteer user, a probe host on that network, or a cloud VM in the same autonomous system (the numbered network an ISP announces). Throughout, "resolver" means the recursive resolver that does lookups for users, an "authoritative server" is one that holds the zone's records, and an NS record names a zone's authoritative servers. The four hops are the resolver's cache (is the answer already stored?), the resolver itself (is that server slow or overloaded?), the path between resolver and authoritative servers (loss, truncation, fragmentation, a bad route), and the authoritative side (the servers that hold the zone). You can only see inside your own resolver's cache, so for the ISP's cache you infer state from the TTL (time-to-live, the seconds an answer may be reused) that comes back.
First confirm the narrowing. In the browser, the Navigation Timing API (the browser's built-in record of when each stage of a page load began and ended) has two fields, domainLookupStart and domainLookupEnd, whose difference is the DNS time for that page load. Group them by ISP network. If the slow group is one or two ISPs, the cause sits at or behind those ISPs' resolvers, not in your zone's content.
The four checks, in order
| Hop | Command from the affected network | What the result tells you |
|---|---|---|
| Cache | dig @ISP_RESOLVER www.example.com, run twice a few seconds apart | TTL falls between runs: served from cache. TTL resets to the full value each time: a miss every time (several resolver nodes behind one address, each with its own cache, or a TTL clamp, meaning the resolver caps the TTL it honours at a small value, so it refetches more often than the zone asked). |
| Resolver | Same query to a second resolver from the same network (dig @OTHER_RESOLVER ...) and to the ISP resolver for a different, popular name | Only the ISP resolver is slow, and for all names: the resolver is the problem. Slow only for your name: look further out. |
| Path | dig @AUTH_NS www.example.com +norecurse and the same with +tcp; compare | UDP times out or is lossy but TCP is fine: packet loss, a UDP size problem or filtering. Both slow: routing or congestion toward that name server. |
| Authoritative | Query every NS of the zone directly, dig @ns1 ... +norecurse, then ns2 | One name server slow or unreachable from that ISP: resolvers there keep retrying it, adding delay. All fast: the zone side is cleared. |
Why the TCP comparison matters, in plain words. A UDP DNS reply travels inside one IP packet, and every link has an MTU (maximum transmission unit, the largest packet it carries). If the reply is bigger than allowed there are two outcomes: the server truncates it (sends a short reply with the TC flag set, and the client repeats the question over TCP), or the IP layer fragments it (cuts it into pieces that the receiver must reassemble, and if one piece is lost or a firewall drops fragments, nothing arrives at all). EDNS is the DNS extension that lets a client state, in an extra record, how large a UDP reply it will accept. An advertised payload of 1232 bytes is chosen like this: IPv6 requires every link to carry at least 1,280-byte packets; the IPv6 header takes 40 bytes and the UDP header 8, so 1280 - 40 - 8 = 1232 bytes of DNS data always fit in one packet with no fragmentation anywhere. A smaller advertised size therefore avoids fragmentation by making bigger answers truncate and move to TCP instead. Some ISPs drop fragments, so large signed (DNSSEC, the DNS signing extensions) responses can stall for them and nobody else.
A lab that separates the hops
The lab is an authoritative server (nsd, listening on 127.0.0.2) and a caching resolver (Unbound, 127.0.0.1) that sends queries for example.com to it. Everything needed to rerun it, with what each line is for:
# example.com.zone
$ORIGIN example.com.
$TTL 300
@ IN SOA ns1.example.com. hostmaster.example.com. 2024010101 3600 600 604800 300
@ IN NS ns1.example.com.
ns1 IN A 127.0.0.2
www IN A 203.0.113.10
api IN A 203.0.113.11
# nsd.conf
server:
ip-address: 127.0.0.2
username: ""
zonesdir: "/lab"
pidfile: "/tmp/nsd.pid"
database: ""
zone:
name: "example.com"
zonefile: "example.com.zone"
# unbound.conf
server:
interface: 127.0.0.1
access-control: 127.0.0.0/8 allow
do-not-query-localhost: no
domain-insecure: "example.com"
username: ""
extended-statistics: yes
remote-control:
control-enable: yes
control-interface: 127.0.0.1
control-use-cert: no
stub-zone:
name: "example.com"
stub-addr: 127.0.0.2
What the lines do:
$TTL 300gives every record a 5 minute TTL by default, which is why the first answer shows 300.- The
SOAline is2024010101 3600 600 604800 300: serial, refresh (1 hour), retry (10 minutes), expire (7 days), negative-caching TTL (5 minutes). nsd needs a valid one to load the zone; the values are illustrative. nsd.conf:ip-addressis where nsd listens;username: ""skips dropping privileges (fine in a throwaway lab);zonesdirandzonefilesay where the zone file is;pidfilewhere nsd records its process ID;database: ""stops nsd from writing a database file; thezone:block names the zone to serve.extended-statistics: yesmakesunbound-control stats_noresetprint the per-answer-code counters such asnum.answer.rcode.NOERRORandnum.answer.rcode.SERVFAIL; without it those lines are absent and the finalgrepshows only the threetotal.numlines.unbound.conf:interfaceandaccess-controllet only loopback clients ask.do-not-query-localhost: nois needed because the authoritative server sits on a loopback address (127.0.0.2) and Unbound refuses to send queries to loopback by default.domain-insecuretells Unbound not to demand DNSSEC signatures for this unsigned lab zone. Theremote-controlblock letsunbound-controltalk to the running Unbound on 127.0.0.1 without certificates (lab only). Thestub-zonesays: forexample.com, ask the server at 127.0.0.2 directly instead of walking down from the root.
Start both (nsd -c /lab/nsd.conf -d & and unbound -c /lab/unbound.conf -d &). Three results carry the evidence: the TTL counting down from 300 (a cache hit), the aa (authoritative answer) flag on the direct query (the authoritative side agrees), and SERVFAIL for an uncached name once the authoritative side is gone. Run:
$ unbound-control -c /lab/unbound.conf dump_cache | grep www # before any query
$ dig @127.0.0.1 www.example.com +noall +answer +comments | grep -E "status|^www"
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 26737
www.example.com. 300 IN A 203.0.113.10
$ sleep 3; dig @127.0.0.1 www.example.com +noall +answer
www.example.com. 297 IN A 203.0.113.10
$ unbound-control -c /lab/unbound.conf dump_cache | grep -B1 -A3 www
;rrset 297 1 0 8 3
www.example.com. 297 IN A 203.0.113.10
END_RRSET_CACHE
START_MSG_CACHE
msg www.example.com. IN A 32896 1 297 3 1 0 0 -1
www.example.com. IN A 0
END_MSG_CACHE
EOF
$ dig @127.0.0.2 www.example.com +norecurse +noall +comments +answer | grep -E "flags|^www"
;; flags: qr aa; QUERY: 1, ANSWER: 1, AUTHORITY: 1, ADDITIONAL: 2
; EDNS: version: 0, flags:; udp: 1232
www.example.com. 300 IN A 203.0.113.10
The ; EDNS line in the last result also matches the flags pattern of the grep; it is the OPT pseudo-record in which the server states the UDP payload size it will use (1232 here, the value derived above).
Reading the cache dump: Unbound stores two things. The ;rrset line starts a cached record set: 297 seconds left and 1 record in it, followed by the record itself. The msg line is the cached complete answer to the question "www.example.com A": name, class, type, then the DNS flags as a number (32896 is hex 8080: the response and recursion-available bits), the question count (1), the seconds left (297), a validation status number, and the counts of answer, authority and additional records (1, 0, 0). The line under msg, www.example.com. IN A 0, is that answer's pointer to the record set held in the rrset cache above it, so both caches agree. The dig TTL and the dump TTL (297) match because both read the same stored entry.
Reading the three dig results: the first query populates the cache and returns the full TTL of 300; three seconds later the same name returns 297, which is a cache hit counting down. The authoritative server answers directly with the aa flag (authoritative answer) and the full 300, so the two sides agree. Now take the authoritative side away and ask for a name that was never cached, then for the cached one:
$ pkill nsd; sleep 1
$ dig @127.0.0.1 api.example.com +noall +comments | grep status # takes about 10 seconds
;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 34452
$ dig @127.0.0.1 www.example.com +noall +answer
www.example.com. 286 IN A 203.0.113.10
$ dig @127.0.0.2 api.example.com +tries=1 +time=2 +noall +comments
;; communications error to 127.0.0.2#53: connection refused
;; no servers could be reached
$ unbound-control -c /lab/unbound.conf stats_noreset | grep -E "^total.num.(queries|cachehits|cachemiss)=|rcode.(NOERROR|SERVFAIL)="
total.num.queries=6
total.num.cachehits=3
total.num.cachemiss=3
num.answer.rcode.NOERROR=3
num.answer.rcode.SERVFAIL=1
The cached name still resolves (its TTL kept counting down to 286) while the uncached one returns SERVFAIL (the resolver's "I could not get an answer" code). That is the signature to look for in the field: popular names fine, cold names slow or failing points at the path or authoritative side; everything slow points at the resolver. The counters read 6 queries although we typed only 4 lookups through the resolver (www, www, api, www). The reason is that dig resends a query that gets no answer within its default 5 second timeout, up to 3 attempts in all, and the api lookup took about 10 seconds: dig printed two "timed out" lines before the SERVFAIL arrived, so Unbound received the api question three times (the cached name's TTL fell from 297 to 286 meanwhile). Measured counters at each step: after the first two www lookups, queries=2 cachehits=1 cachemiss=1; after the api lookup, 5 / 2 / 3; after the final www lookup, 6 / 3 / 3. So the three api packets added 3 queries (1 hit, 2 misses), and the 3 hits in total are: the second www, the final www, and one of the api packets (I did not trace which). The rcode counters count replies by result: NOERROR=3 is the three www answers and SERVFAIL=1 the one api reply dig received, so 3 + 1 = 4 replies against 6 queries received. The two api packets dig gave up on are counted as queries only. Hits versus misses is the resolver-side view of the same split, and a falling hit ratio is an early sign of trouble.
Choosing the fix once a hop is isolated
- Cache: raise the TTL on records that change rarely. The trade-off is slower rollback, so keep a short TTL only on records you really switch.
- Resolver: you cannot fix the ISP's server, so give affected users a documented alternative resolver, or place an anycast resolver of your own where it matters, and report the evidence to the ISP.
- Path: lower the advertised EDNS size to 1232, make sure TCP/53 is reachable, and consider unsigned or smaller responses for the affected names.
- Authoritative: add geographically distinct name servers or an anycast provider, and remove any server that is slow from the affected ISP.
Pitfalls
Testing from your office proves nothing about an ISP you are not on. Averaging latency hides this problem, because 5% of users on one ISP barely moves a mean: use a per-ISP 95th percentile. Do not "fix" it by flushing caches you do not control, and do not conclude "the ISP is at fault" from one dig.
Unlock Full Question Bank
Get access to all 40 DNS, DHCP, and Name Resolution interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.