DNS, DHCP, and Name Resolution Questions
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
How can an attacker get a resolver to cache a forged answer, and what makes the attack hard today? What would you still worry about?
Sample Answer
Direct answer
An off-path attacker (one who cannot see the victim's traffic, as opposed to an on-path attacker, who sits on the route and can) tricks a recursive resolver into caching a forged answer by racing the real authoritative server: the resolver sends a query over UDP, and the attacker floods it with fake replies that must match the query's 16-bit transaction ID, its source port, its question and the server address it was sent to. A fake that arrives first and matches is accepted and cached for its TTL. The 2008 attack disclosed by Dan Kaminsky (CERT note VU#800113, 8 July 2008) made this practical against resolvers that used a fixed source port, and what makes it hard today is the combination of a random ID and a random source port on every query (RFC 5452), strict matching and bailiwick rules (a resolver accepts only data that belongs to the zone it asked about), and above all DNSSEC, which makes a forged answer fail signature validation. What I would still worry about: unsigned zones, attackers who can see or influence the path, side channels that reveal the source port, and compromises that produce valid-looking data, such as a hijacked registrar account.
The attack, in order
- The attacker makes the victim resolver ask a question whose answer it does not already hold, for instance by triggering a lookup (many resolvers answer anyone who can reach them, or a victim's own user or mail server can be induced to look the name up).
- The resolver sends a query to the authoritative server with a transaction ID and source port.
- The attacker sends many spoofed replies purporting to be from that server's address. Each carries a guessed ID and port. Only one reply per race can be accepted.
- If a forged reply matches and beats the real one, the resolver caches the attacker's data until its TTL (time-to-live) expires, and every client of that resolver is served it.
Why Kaminsky's variant mattered. Before it, attackers forged the answer for a real name such as www.example.test; after one failed race the real answer was cached and the attacker had to wait out the TTL. Kaminsky's insight was to ask for random names that cannot be cached, such as a1b2c3.example.test, so every attempt is a fresh race, and to put the poison not in the answer but in the authority and additional sections of the forged reply: an NS record for example.test and a glue A record pointing at the attacker's server. Here is an illustrative forged reply for a query about a1b2c3.example.test (203.0.113.66 is a documentation address standing in for the attacker's server):
;; QUESTION SECTION:
;a1b2c3.example.test. IN A
;; ANSWER SECTION: (empty: the name does not exist)
;; AUTHORITY SECTION:
example.test. 86400 IN NS ns.example.test.
;; ADDITIONAL SECTION:
ns.example.test. 86400 IN A 203.0.113.66
A DNS reply has an answer section (the records that answer the question), an authority section (which servers are responsible) and an additional section (extra records offered to save a lookup, such as the address of a name server). The NS record there says "example.test is served by ns.example.test", and the additional-section A record (glue, the address of a name server whose name sits inside the zone it serves) says where that server is. If the resolver caches them, every later query for any example.test name goes to the attacker. The glue is in-domain (it belongs to the zone being asked about), so a rule that rejects out-of-domain data does not stop it; what stopped it was making each race unlikely to be won.
What makes it hard: the arithmetic
RFC 5452 requires a resolver to match the query ID, the source and destination addresses, the destination port against the query's source port, and the question, and to pick the ID and source port unpredictably across the full range available (ports 53 or 1024 and above). A 16-bit ID gives 65,536 values; random ports multiply that. The ID and port randomisation is the main defence, so the arithmetic is worth doing. A pinned calculation, where the attacker lands 1000 distinct forged (ID, port) guesses in each race and the port range is Linux's default ephemeral range (the block of high port numbers the operating system hands to outgoing connections) 32768 to 60999:
import math
ids = 65536 # 16-bit DNS query ID
ports = 60999 - 32768 + 1 # Linux default ephemeral range 32768-60999
forged = 1000 # forged replies that land inside one race window
print(f"ports in range: {ports}")
print(f"ID x port space: {ids * ports} ({math.log2(ids * ports):.1f} bits)")
for label, space in (("fixed port ", ids), ("random port", ids * ports)):
p = forged / space
races = math.log(0.5) / math.log(1 - p)
print(f"{label}: p per race = {p:.3e}, races for 50% = {races:.0f}")
ports in range: 28232
ID x port space: 1850212352 (30.8 bits)
fixed port : p per race = 1.526e-02, races for 50% = 45
random port: p per race = 5.405e-07, races for 50% = 1282469
Where the races formula comes from: if one race is won with probability p, it is lost with probability 1 - p, so losing n races in a row has probability (1 - p)^n. The attacker has a 50% chance of at least one win when that equals 0.5, and solving (1 - p)^n = 0.5 for n gives n = ln(0.5) / ln(1 - p), which is the math.log(0.5) / math.log(1 - p) line. Check by running it forward: with the fixed port, 1 - (1 - 0.01526)^45 is 0.499, and with the random port 1 - (1 - 5.405e-07)^1282469 is 0.500 (both checked in Python).
With a fixed port, 1000 forged replies per race win about 1.5% of the time and about 45 races give a 50% chance, which Kaminsky's fresh-name trick makes cheap. With a random port, a race wins about 5.4 in ten million, and it takes about 1.28 million races for a 50% chance. The port range here is 28,232 ports, about 14.8 bits, so the total is 30.8 bits rather than the 32 you get from a full 64k range; narrowing the allowed range (a firewall rule, or a resolver setting such as Unbound's outgoing-port-avoid that removes ports from use) directly weakens it.
Other layers:
- Bailiwick and glue checks. A resolver should accept data only if it belongs to the zone it asked about (RFC 5452 section 6); Unbound's
harden-glue(defaultyes) trusts glue only within the server's authority. - 0x20 encoding. The resolver randomises the upper and lower case of the query name, for example asking for
eXaMpLe.TeStinstead ofexample.test; an honest server echoes the case back exactly, so a forger must guess the case pattern too, and every letter in the name adds one more bit to guess. Unbound'suse-caps-for-idis an experimental option and defaults tono. - DNSSEC. The real fix: the forged data has no valid signature under the zone's key, so a validating resolver rejects it regardless of how well the forger guessed (a chain of signatures from the root key down to the zone: each level vouches for the key of the level below).
What I would still worry about
| Residual risk | Why randomisation does not cover it |
|---|---|
| Zones and resolvers without DNSSEC | The race is only hard, not impossible; at 30.8 bits an attacker with a fast path still gets there eventually |
| On-path attackers | They see the ID and port, so there is nothing to guess; only DNSSEC (or an encrypted resolver-to-server channel) helps |
| Side channels that leak the port | A side channel is information that leaks through a side effect rather than the reply itself. The 2020 SAD DNS research (UC Riverside) used the global rate limit on outgoing ICMP (Internet Control Message Protocol, the error-message protocol, for example "port unreachable") messages: by probing and watching how many ICMP errors came back, the attacker could tell which source port was open, then guessed only the ID. Listed mitigations: randomise or disable that ICMP behaviour, add secrets such as 0x20 or DNS cookies (a random value the client and server exchange and echo, RFC 7873, which off-path forgers cannot know), deploy DNSSEC, and shorten the window a query stays outstanding |
| Port randomisation undone en route | A NAT or firewall that remaps source ports predictably gives back the entropy (the unpredictability, measured in bits); confirm with a capture that ports seen leaving the network still vary |
| Valid-looking bad data | A registrar or DNS-provider account takeover, or a hijacked route to an authoritative server, produces correctly signed or simply authoritative answers; DNSSEC does not stop a registrar compromise because the attacker can change the DS too |
| Shared resolvers | One poisoned cache serves everyone behind it, so open or lightly protected forwarders raise the blast radius |
What I would do
Run a validating resolver (with DNSSEC on), restrict who may query it, keep the outgoing port range wide, leave harden-glue on, sign my own zones, lock down registrar accounts with multi-factor authentication and registry locks where offered (a registry lock is an extra step in which the registry itself, not just the registrar, must approve changes to a domain's delegation), and capture traffic once to verify source ports actually vary leaving the network.
Explain how DNS can be abused to flood a third party with traffic, and what you would change on authoritative and recursive servers to avoid being part of it.
Sample Answer
Direct answer
DNS amplification is a reflection attack. DNS usually runs over UDP, which has no handshake, so an attacker can send a small query with the victim's address forged as the source. The server answers the victim, and if the answer is much bigger than the query, the attacker multiplies their bandwidth. Any server that answers such queries can be used: open recursive resolvers (they answer anyone) and authoritative servers that return large answers. An authoritative server holds the records for particular zones and answers for them from its own data. A recursive resolver does lookups on behalf of its clients by asking other servers, and caches what it learns. An open recursive resolver does this for any address on the internet, which is what makes it reusable by an attacker. To avoid being a participant, close recursion to your own clients, shrink answers, rate-limit identical responses on authoritative servers, and filter spoofed source addresses at the network edge. Two controls matter most: rate limiting (and restricting recursion on resolvers), which stops your server from being a useful reflector, and ingress filtering at the network edge, which removes the forged addresses at their source. The answer-size controls reduce how much damage each forged query can do.
The mechanism in order
- The attacker, on a network that does not filter forged sources, sends UDP query packets whose source address is the victim's.
- The query asks for something with a large answer: a long TXT record (free-form text stored in DNS), a DNSSEC-signed record set (records returned together with the cryptographic signature records that DNSSEC adds, which makes them much larger), or ANY (a query type that asks for every record type held at a name).
- The server replies to the source address on the packet, which is the victim. The victim never asked, so it simply receives unwanted traffic, and the real attacker's address never appears.
- Thousands of servers doing this at once produce a flood the victim cannot filter by source, because the traffic comes from legitimate, well-known servers.
Measured amplification
The lab authoritative server (BIND 9.20 with max-udp-size 4096; set so that it will send a large UDP reply; the 9.20 default is 1232) holds a TXT record made of six 200-character strings. A query for it with an EDNS buffer of 4096 bytes and no cookie, captured on the wire, was 44 bytes of DNS data. In that command, +bufsize=4096 tells the server the client will accept UDP replies up to 4096 bytes, and +nocookie leaves out the EDNS cookie, a small random token that a client can attach to a query so a server can later recognise it (a query without one looks like a first contact from an unverified address, which is what a forged query looks like):
$ dig @127.0.0.2 big.example.com TXT +bufsize=4096 +nocookie | grep 'MSG SIZE'
;; MSG SIZE rcvd: 1262
IP 127.0.0.1.52285 > 127.0.0.2.53: 38531+ [1au] TXT? big.example.com. (44)
The ratio on DNS payload is 1262 / 44 = 28.7. Counting the 28 bytes of IPv4 plus UDP headers on each side, it is (1262 + 28) / (44 + 28) = 17.9 times the packet size. A 100 Mbit/s spoofed stream becomes roughly 1.8 Gbit/s at the victim on the packet-size ratio. Real attackers choose names with far bigger answers, which is why response size is the lever to pull. The same query against the same server left at its default max-udp-size was not reflected at that size: tcpdump showed a 44-byte UDP reply with the TC (truncated) bit set, so dig retried over TCP, and the UDP amplification was 1.0 instead of 17.9. That is the 1232-byte cap in the table below doing its job.
What to change on the authoritative servers
| Control | Why it helps | Where it sits |
|---|---|---|
Turn recursion off (recursion no;) | An authoritative server that will not resolve arbitrary names cannot be asked for other people's data (RFC 5358: authoritative-only servers should turn recursion off) | Server option |
| Response rate limiting (RRL) | Caps how many identical responses go to one client prefix per second, so a spoofed victim stops receiving them | Server option |
| Small EDNS (extension mechanisms for DNS, which lets a client announce the largest UDP reply it accepts) buffer, set to 1232 bytes | The DNS Flag Day 2020 recommendation avoids IP fragmentation on nearly all networks (1280-byte IPv6 minimum MTU minus 48 bytes of headers), and also bounds the biggest UDP answer | Server option |
Minimal answers for ANY (RFC 8482) and minimal-responses | Removes the cheapest way to ask for everything at once and trims the additional section | Server option |
| Truncation (TC bit) for oversized answers | A real client retries over TCP, where the handshake makes forging the source address impractical; a spoofed victim gets only a tiny truncated packet | Server behaviour |
Terms in the table. IP fragmentation is what happens when a UDP packet is larger than a link's maximum packet size (the MTU, commonly 1500 bytes): the network splits it into pieces, and many firewalls drop fragments, so large answers can fail to arrive. The TC bit (truncated) is a flag in the DNS reply header that tells the client "the reply did not fit, ask again over TCP".
A configuration that passes named-checkconf on BIND 9.20:
options {
recursion no;
minimal-responses yes;
minimal-any yes;
max-udp-size 1232;
edns-udp-size 1232;
rate-limit {
responses-per-second 20;
slip 2;
window 5;
exempt-clients { 10.0.0.0/8; };
};
};
Reading the configuration line by line: recursion no; refuses to resolve names the server is not authoritative for. minimal-responses yes; leaves out optional extra records in the additional section. minimal-any yes; answers an ANY query over UDP with only one record set instead of all of them. max-udp-size 1232; caps the size of any UDP reply the server sends, so bigger answers are truncated (TC bit) and retried over TCP. edns-udp-size 1232; is the buffer size the server advertises when it sends queries to other servers, which matters when it also acts as a resolver; per the BIND reference both sizes default to 1232 in 9.20, and setting them explicitly records the intent. Inside rate-limit: responses-per-second 20; allows about 20 identical responses per second to one client network; window 5; sets the length in seconds of the rolling window over which responses are counted (the BIND default is 15); slip 2; lets every second rate-limited response out as a small truncated reply instead of dropping it (explained next); exempt-clients { 10.0.0.0/8; }; is a list of addresses that are never rate limited, such as your own monitoring.
Slip, explained first because the test below depends on it. When a client exceeds the limit, the server drops the surplus responses. With slip set to N, every Nth rate-limited response is instead sent as a tiny truncated reply (TC bit), so a real client whose address is being forged by an attacker can still get through by retrying over TCP, while the victim receives only a few small packets. slip 0 means no small replies are sent, so everything over the limit is dropped; slip 2 is the BIND default.
RRL measured. The lab server ran with responses-per-second 5; window 5; slip 0;. Thirty identical queries were fired in parallel from one source address with a one-second wait and one try each. At 5 responses per second the first 5 were answered; the other 25 went over the limit and, with slip 0, were dropped without even a truncated reply, so those clients timed out:
answered: 5 dropped: 25
When the same thirty queries came from an address listed in exempt-clients, the result was answered: 30 dropped: 0. Slip is the knob that keeps legitimate clients working: according to the BIND reference, slip 2 (the default) answers every other rate-limited UDP request that lacks a valid server cookie (the server's half of the EDNS cookie exchange, which a client that has talked to this server before can present) with a small truncated response, so a real client can retry over TCP, while slip 0 drops all of them. Choose slip 2 for production, slip 0 only when you accept that real clients behind the same prefix will see timeouts. RRL needs an exempt list for your own monitoring and for large shared resolvers that front many real users, because their queries all arrive from one address.
What to change on the recursive servers
An open resolver is the worst case, because it answers anyone and can be pointed at any large record, including one in a zone the attacker controls and fills with oversized data. In the lab, a recursive server with allow-recursion { 127.0.0.4; }; behaved this way:
$ dig -b 127.0.0.1 @127.0.0.4 www.example.com A +noall +comments | grep -o 'status: [A-Z]*'
status: REFUSED
$ dig -b 127.0.0.4 @127.0.0.4 www.example.com A +short
192.0.2.10
- Restrict recursion and cache access to your own clients with
allow-recursionandallow-query-cache(BIND), oraccess-controlentries (Unbound). RFC 5358 says nameservers should not offer recursive service to external networks by default. - Keep recursive and authoritative roles on separate servers where practical.
- Apply the same EDNS size cap, and the same RRL where the software supports it.
- Test from outside your network:
dig @your-server-ip example.org A +recursefrom a host on the internet must return REFUSED.
Network layer
Response size controls how much damage a spoofed query does, but the root cause is the forged source address. Ingress filtering (BCP 38, RFC 2827) means an edge router drops packets whose source address does not belong to the prefixes legitimately reachable behind that interface. Apply it at your own network edge so that you are not the source of forged packets, and ask your upstream to do so.
Trade-offs and pitfalls
- RRL slows legitimate bursts from one address. Size the limit above the busiest legitimate client and watch the dropped and truncated counters.
- Truncation pushes more clients onto TCP, so confirm that TCP/53 is open and that the server handles the connection load.
- A 1232-byte buffer makes large DNSSEC answers truncate more often. That is the intended trade-off against fragmentation.
- RRL does not stop an attack by itself if many different servers each answer a small number of queries. Ingress filtering at the source remains the real fix.
Your DNS logs show repeated zone transfer requests from many unknown addresses. How do you decide whether data left the building, what do you do right now, and what changes keep it from happening again?
Sample Answer
Direct answer
A zone transfer (AXFR, the full-zone transfer request a secondary server uses to copy a zone) hands the requester every record in the zone in one TCP session. Whether data left depends on one thing in your logs: for each requesting address, did the server log a refusal or did it log a transfer that started and ended? Contain first by restricting transfers to your own secondaries with a shared-secret signature, then prove the fix from outside, then add the controls that stop it recurring.
Step 1: decide whether data left (ordered checks)
- Confirm the server logs transfers at all. In BIND, logging is grouped into categories. Completed or started outbound transfers go to the
xfer-outcategory, and refused requests go to thesecuritycategory. If neither is going to a file or the system log with retention, you cannot prove a negative, and the fallback evidence is network flow records for TCP port 53 with large byte counts out of the server. - Separate refusals from completed transfers. This lab log shows both, from a server that requires a TSIG key (transaction signature: a shared-secret code, an HMAC, which is a keyed hash that only holders of the secret can compute, attached to a DNS message to prove who sent it) for transfers. The three addresses 127.0.0.1, 127.0.0.2 and 127.0.0.3 are loopback aliases standing in for outside hosts, and each address asked twice in a loop (127.0.0.1 appears a third time, which is the plain
digfrom Step 2). The lab server wrote both log categories to one file with this logging block (in production they may go to different files, so tally across both), listened on port 5353, and served this zone, all TTLs 300. The zone has an SOA, one NS and three address records, which is why a full transfer is 6 records (the SOA appears twice, at the start and as the closing marker):
logging { channel lf { file "/var/log/named/named.log" versions 3 size 1m; print-time yes; }; category xfer-out { lf; }; category security { lf; }; category default { null; }; };
$TTL 300
@ IN SOA ns1.example.com. hostmaster.example.com. 1 7200 900 1209600 300
@ IN NS ns1.example.com.
ns1 IN A 192.0.2.53
www IN A 192.0.2.80
staging IN A 192.0.2.81
The log:
06-Oct-2026 08:24:52.894 client @0xffff93b00000 127.0.0.1#55405 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.899 client @0xffff93880000 127.0.0.1#55411 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.903 client @0xffff93b00000 127.0.0.2#50758 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.908 client @0xffff93600000 127.0.0.2#50766 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.913 client @0xffff93380000 127.0.0.3#45162 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.917 client @0xffff93880000 127.0.0.3#45168 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.921 client @0xffff93600000 127.0.0.1#55417 (example.com): zone transfer 'example.com/AXFR/IN' denied
06-Oct-2026 08:24:52.927 client @0xffff93100000 127.0.0.1#55424/key xfer-key (example.com): transfer of 'example.com/IN': AXFR started: TSIG xfer-key (serial 1)
06-Oct-2026 08:24:52.927 client @0xffff93100000 127.0.0.1#55424/key xfer-key (example.com): transfer of 'example.com/IN': AXFR ended: 1 messages, 6 records, 310 bytes, 0.001 secs (310000 bytes/sec) (serial 1)
Read one line left to right: date, time, the word client, an internal handle starting with @0x, then the requester as address#source-port, then the zone name in brackets, then what happened. denied means nothing left. AXFR started followed by AXFR ended means the whole zone left, and the ended line gives the record and byte counts so you know exactly how much (here 6 records, 310 bytes). The /key xfer-key after the address shows the request was signed with the TSIG key.
3. Tally by source address. The address is the fifth whitespace-separated field of each line above (date, time, client, handle, address), so split($5,a,"#") cuts 127.0.0.1#55405 at the # and keeps the address. This command counts denied lines per address:
grep -E "zone transfer .* denied" named.log | awk '{split($5,a,"#"); n[a[1]]++} END{for(k in n) print n[k], k}' | sort -k2
The grep keeps only the refusal lines, awk adds one to a counter for each address, and the END block prints each count and address. On the log above it printed:
3 127.0.0.1
2 127.0.0.2
2 127.0.0.3
That is 7 denied attempts from 3 addresses (the extra attempt from 127.0.0.1 is the plain dig in Step 2). If your syslog adds a hostname or process name in front, the address moves to a later field, so print one line with awk '{for(i=1;i<=NF;i++) print i, $i}' and pick the right field number.
4. List every completed transfer whose source is not one of your secondaries. grep 'AXFR started' named.log and compare each client address with the secondaries list. One hit is a disclosure; zero hits with the logging confirmed in step 1 means no zone data left by transfer.
5. Classify what the zone contains. A zone served to the public already exposes each name that someone can guess and query, but a transfer also exposes names nobody linked to, staging and admin hosts, and every TXT record. A zone from an internal view (a BIND view, which gives different answers depending on who asks) or split-horizon setup (different answers for inside and outside clients) exposes the internal network map, which is the serious case.
Many unknown sources with all lines denied is reconnaissance, not a breach. Record it and watch for follow-on activity against the hostnames in the zone.
Step 2: what to do right now
- Restrict transfers to the secondaries and require TSIG. Do not rely on a default: in this lab a BIND 9.20 zone with no
allow-transferrefused a loopback client, but other servers and older configurations may allow any host. The only reliable check is running a transfer from outside against every authoritative address.
key "xfer-key" { algorithm hmac-sha256; secret "bXktdHJhbnNmZXItc2VjcmV0LTMyLWJ5dGVzLWxvbmchISE="; }; # lab-only secret
zone "example.com" { type primary; file "/etc/bind/z/example.com.zone"; allow-transfer { key xfer-key; }; };
- Verify from a host outside the allowed list, and with the key:
$ dig -p 5353 @127.0.0.1 example.com AXFR +noall
; Transfer failed.
$ dig -p 5353 @127.0.0.1 example.com AXFR +noall +answer -y hmac-sha256:xfer-key:<secret> # the secondary's view
example.com. 300 IN SOA ns1.example.com. hostmaster.example.com. 1 7200 900 1209600 300
example.com. 300 IN NS ns1.example.com.
ns1.example.com. 300 IN A 192.0.2.53
staging.example.com. 300 IN A 192.0.2.81
www.example.com. 300 IN A 192.0.2.80
example.com. 300 IN SOA ns1.example.com. hostmaster.example.com. 1 7200 900 1209600 300
The signed transfer returns the SOA, the zone's records and the SOA again as a closing marker: those six records are the "6 records" in the ended log line, and note that the unlinked staging host is in it.
- Do not block TCP port 53 at the firewall to stop this. RFC 9210 says DNS servers MUST support TCP, because responses larger than a UDP packet fall back to it, so blocking TCP breaks normal answers. The control is
allow-transfer, not the port. - Preserve evidence: copy the server log and any packet capture before log rotation, and note the time of the first logged attempt.
- If step 1 found a completed transfer to an unknown address, treat the zone contents as disclosed: review whether internal hostnames, IP plans or credentials-bearing TXT records were in it, and notify your security lead.
Step 3: changes that keep it from happening again
- Transfers to secondaries only, authenticated with TSIG (RFC 5936 section 5 says a DNS implementation SHOULD provide means to restrict AXFR sessions to specific clients, accepts address-based grants, and RECOMMENDS access control based on TSIG or SIG(0) as the stronger option).
- Put the zone's master on a hidden primary (a primary server that is not listed in the zone's NS records) that only the secondaries can reach, so the servers that face the internet never accept transfer requests from anyone else.
- Keep internal zones in an internal view or on internal-only servers; never publish them on internet-facing authoritative servers.
- Alert on any
AXFR startedfrom a non-secondary address, and on a denied-attempt rate above your normal (a baseline of zero is typical). - Rotate the TSIG key on a schedule and when staff with access leave.
Trade-offs
Restricting by source IP alone is simple but proves only an address, not the requester (TCP makes blind spoofing hard, but any host in an allowed range passes), and it breaks when a secondary's address changes; TSIG costs key distribution but authenticates the requester. Logging every denied attempt costs log volume during a scan, so aggregate by source before alerting.
What problem does DNSSEC solve and what does it not? Describe how a validating resolver decides an answer is genuine, and what operational risks come with turning it on.
Sample Answer
Direct answer
DNSSEC (DNS Security Extensions) adds digital signatures to DNS data so a validating resolver can prove that an answer really came from the zone's owner and was not altered in transit, including proof that a name or record type does not exist (RFC 4033: data origin authentication, data integrity and authenticated denial of existence). It does not encrypt anything, does not hide which names you query, does not stop denial-of-service attacks, and does not protect the hop between a client and its resolver unless the client validates too. A validating resolver (a resolver that checks signatures instead of trusting what it receives) decides an answer is genuine by following a chain of trust, where each zone vouches for the zone below it, starting from a trust anchor (a key the resolver is configured to trust without proof, normally the root zone's key) down through each delegation, then checking the signature on the answer. The operational risk is that a mistake does not degrade service: a broken chain makes validating resolvers return SERVFAIL, so the domain looks down. Start with the chain of trust, what SERVFAIL means when a link breaks, and the +cd test that separates chain problems from data problems; the sections after the lab run go further into proofs of non-existence and key-rollover timing.
How validation works, link by link
Four record types carry it:
DNSKEY: a zone's public keys. A key with flags 257 is a key-signing key (KSK, the key referenced from the parent); flags 256 marks a zone-signing key (ZSK, used to sign ordinary records). Two keys exist so that the key the parent has to know about changes rarely, because changing it means a request to the parent, while the key that signs the many ordinary records can be replaced often by you alone. One key can do both jobs, and the lab zone below has just one.RRSIG: a signature over one record set, naming the key tag of the signing key (a 16-bit number that labels a key so a signature can say which key made it) and a validity window (inception and expiration times).DS(delegation signer): a hash of a child zone'sDNSKEY, published in the parent zone. It is the link between levels.NSEC(next secure) orNSEC3(its hashed variant): records that prove non-existence.
The resolver's steps, in order (RFC 4035 section 5):
- Start from the trust anchor, the root's key (or its
DS), configured in the resolver. - Fetch the root
DNSKEYset and check it matches the anchor and is signed. - For the top-level domain, fetch the
DSfrom the root zone (itsRRSIGvalidates under the root key). Fetch the TLDDNSKEYset and check that a key hashes to thatDS. - Repeat for the domain:
DSin the TLD zone,DNSKEYin the domain, hash match. - Check the
RRSIGon the answer against the domain'sDNSKEY. - If every link holds, the resolver sets the AD (authenticated data) flag in its reply (RFC 4035 section 3.2.3). If a link in a signed chain fails, it must return
SERVFAIL(server failure, response code 2) unless the client set the CD (checking disabled) bit (RFC 4035 section 5.5).
A zone without a DS in its parent is an "island of security" (RFC 4033: a signed zone with no chain up to the parent): treated as unsigned, not as failed.
Lab run
Lab: BIND 9.20 with dnssec-policy default signing a made-up root, test and example.test (no real DNSSEC infrastructure is involved), Unbound 1.26 validating with the root DS as trust anchor. Each DS was copied into its parent zone with dnssec-dsfromkey.
$ dig +norecurse +noall +answer @127.0.0.10 test DS
test. 86400 IN DS 55972 13 2 6F13C6A78B77107132B3172DFC06CE03F0DBA9BDAFAA8734FB08FEEA BEF30366
$ dig +norecurse +noall +answer @127.0.0.11 example.test DS
example.test. 86400 IN DS 15733 13 2 1FF321FF44D0C88A116406914EA6B5EB1EA146D84DB07D0DDF30465A D36AAEDC
$ dig +norecurse +noall +answer @127.0.0.12 example.test DNSKEY
example.test. 3600 IN DNSKEY 257 3 13 7/8gpSIJyX8TEDF+fYb+N1qPkoKSpugsYQeX8ag+hBxiGTDrSefBZVGq imlvQv5El0O803rKKjpJONr1AL7pvQ==
$ dig +dnssec +noall +comments +answer @127.0.0.20 web.example.test A
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 20431
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags: do; udp: 1232
;; ANSWER SECTION:
web.example.test. 300 IN A 192.0.2.20
web.example.test. 300 IN RRSIG A 13 3 300 20261019193102 20261006062048 15733 example.test. GsdD5zujGnG9Vh7NQ/HntTJP6mGfhIqOKhDeSBOK+1bpM9Brw5SjrA6W ZZe72oEZCcKRD/Vb9zWnOMzdHt6PAQ==
Read each line field by field. The DS line is: owner name example.test., TTL 86400, class IN, type DS, then four fields: key tag 15733 (which key it points at), algorithm 13 (ECDSA P-256 with SHA-256, IANA registry), digest type 2 (SHA-256, the hash used) and the long hex string, which is the hash of the child's public key. The DNSKEY line is: name, TTL 3600, IN, DNSKEY, flags 257, protocol 3 (always 3), algorithm 13 and then the public key in base64. The RRSIG line is: name, TTL, IN, RRSIG, then the type it covers (A), algorithm 13, labels 3 (the number of labels in the signed name, which lets a resolver spot wildcard answers), original TTL 300, signature expiration, signature inception, key tag 15733, signer name example.test. and the base64 signature. In short: the DS for example.test has key tag 15733, algorithm 13 and digest type 2. The DNSKEY published by the domain is flagged 257, and the RRSIG over the A record names key tag 15733 as the signer: the key the parent vouched for signed the answer. The resolver's reply carries ad. The same tag, 15733, appears in the DS and the RRSIG because this zone has a single key doing both jobs. The RRSIG window runs from 20261006062048 (inception) to 20261019193102 (expiration), written year, month, day, hour, minute, second in UTC, about two weeks: signatures expire, which is why signing must be automated.
Authenticated denial of existence
A signed "no" is needed too: otherwise an attacker could answer "that name does not exist" for a real name and nobody could tell. The NSEC record is how a signed zone proves a name is absent.
$ dig +dnssec +noall +comments +authority @127.0.0.20 nope.example.test A | grep -E 'status|^;; flags|IN\s+NSEC\s'
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 52890
;; flags: qr rd ra ad; QUERY: 1, ANSWER: 0, AUTHORITY: 6, ADDITIONAL: 1
example.test. 120 IN NSEC mail.example.test. A NS SOA MX TXT RRSIG NSEC DNSKEY TYPE65534
mail.example.test. 120 IN NSEC ns1.example.test. A RRSIG NSEC
The filter keeps only the status, flags and NSEC lines; the six authority records are the SOA, the two NSEC records, and an RRSIG for each. TYPE65534 in the apex type list is a BIND-specific bookkeeping type from its automatic signing, not part of the chain. An NSEC record says "no names exist between this owner and the next name". mail.example.test to ns1.example.test covers the alphabetical gap where nope would sort, so nope.example.test cannot exist; the apex NSEC covers the position where a wildcard would sit, ruling out a wildcard match. Here is why: a wildcard record *.example.test would answer for any name that has no record of its own, and in the sorted order of names * sorts before letters, so a wildcard would sit right after the apex name example.test, in the gap the apex NSEC says is empty (it jumps straight to mail.example.test). With nothing in that gap, no wildcard exists that could have answered for nope. The six records are signed, so the denial is itself authenticated. NSEC3 (RFC 5155) hashes the names so the chain cannot be trivially walked to list the zone (RFC 5155 notes that offline dictionary attacks against the hashes remain possible, only more expensive); RFC 9276 says to use zero extra iterations and recommends against a salt.
Operational risks
| Risk | What happens | Mitigation |
|---|---|---|
Wrong or stale DS at the parent | Validators return SERVFAIL for the whole domain | Change the DS only through a staged rollover; monitor with a validating resolver |
| Signatures expire | Same SERVFAIL once the window closes | Automatic re-signing plus an alert on days to expiry |
| Key rollover mistakes | Cached old DNSKEY or DS meets new signatures | Respect RFC 7583 timings below |
| Provider migration | New provider cannot sign for the old DS | Plan the DS change and the overlap before moving |
| Larger responses | Signed answers need EDNS and may fall back to TCP; firewalls that block it break resolution | Allow TCP/53 and EDNS |
Heavy NSEC3 | Extra CPU on resolvers | Zero iterations (RFC 9276) |
The failure, reproduced
Rebuilding the lab with the first hex digit of example.test's DS digest changed to 0 (so the parent vouches for a hash that matches no key): (the two blocks below are trimmed to the header, flags and answer lines; dig also prints the OPT pseudosection, the question section and query statistics)
$ dig @127.0.0.20 www.example.test A
;; ->>HEADER<<- opcode: QUERY, status: SERVFAIL, id: 7632
;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 0, ADDITIONAL: 1
$ dig +cd @127.0.0.20 www.example.test A
;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 10309
;; flags: qr rd ra cd; QUERY: 1, ANSWER: 2, AUTHORITY: 0, ADDITIONAL: 1
www.example.test. 300 IN CNAME web.example.test.
web.example.test. 300 IN A 192.0.2.20
The data is intact; only validation fails. Asking with +cd and getting an answer while the plain query gets SERVFAIL is the diagnostic: the problem is the chain (DS mismatch or expired signature), not the zone content or reachability. Rolling the DS back, or removing it (which makes the zone an unsigned island), restores resolution once the parent's DS TTL (86400 seconds here) lapses in caches.
Rollover timing
RFC 7583 gives the timings. For a pre-publish ZSK rollover, the new key must sit in the DNSKEY set for Ipub = Dprp + TTLkey before it signs anything, where Dprp is the delay for the change to reach every authoritative server and TTLkey is the DNSKEY TTL. With TTLkey = 3600 s and Dprp = 300 s that is 3900 s (65 minutes). A KSK rollover by the double-DS method publishes the new DS in the parent first, waits for the parent's caches (the lab's DS TTL is 86400 s, one day) before changing the DNSKEY, then removes the old DS only after the child's DNSKEY caches expire. The reason for that order is that the parent's DS is the one link you do not control: while both DS records are there, a cache holding either one still finds a matching key, so no cache is ever left without a link. Plain-language reading of the ZSK formula: Ipub is the publication interval and Dprp the propagation delay, and a resolver may still hold the old DNSKEY set for up to TTLkey after the new one reaches every server, so the new key must have been visible for Dprp + TTLkey before any signature relies on it.
Pitfalls
- Treating DNSSEC as encryption or as a fix for a compromised registrar account: an attacker who controls the registrar can publish a new
DSas well. - Enabling it without a validating test resolver and an expiry alert. The first sign of trouble is user reports.
- Forgetting the
DS. Signing the zone without publishing theDSgives none of the protection (an island of security); publishing theDSbefore the zone is signed breaks it.
A security team wants to block clients from using public encrypted DNS, while users argue it protects their privacy. How do DoH and DoT change what an enterprise can see and enforce, and what would you propose?
Sample Answer
Direct answer
Encrypted DNS protects the query from everyone on the path, and that includes the enterprise. DoT (DNS over TLS) wraps DNS in TLS on its own TCP port, 853 (RFC 7858). DoH (DNS over HTTPS) carries DNS messages inside ordinary HTTPS requests (RFC 8484). Once a device talks to a public encrypted resolver, the company can no longer see the names it looks up, apply DNS-based blocking, or log queries per client. I would not fight this with a blanket ban alone. I would run an enterprise resolver that itself speaks encrypted DNS, point managed devices at it by policy (so users get privacy from the local network and the company keeps visibility on its own devices), block the unmanaged routes at the network edge, and be explicit with staff about what the company can see.
What each side sees and can enforce
| Plain DNS (UDP/TCP 53) | DoT (TCP 853) | DoH (TCP 443) | |
|---|---|---|---|
| Query names visible to the network | Yes | No (encrypted) | No (encrypted) |
| Easy to identify as DNS | Yes, by port | Yes, by port 853 | No, looks like web traffic to a firewall that only reads ports |
| Block with a firewall rule | Port rule works | One rule: deny outbound TCP 853 | Port rule would block all web traffic; needs destination lists or TLS inspection |
| Enterprise DNS filtering, logging, response policy (rewriting or refusing the answer for chosen names) | Works | Works only if devices use the enterprise resolver | Works only if devices use the enterprise resolver |
| What remains visible | Everything | Destination IP, volume, timing | Destination IP, volume, timing, and usually the server name the client announces in the TLS handshake (SNI, server name indication, normally sent unencrypted so the server knows which site is wanted) |
RFC 8484 states this plainly: filtering or inspection systems that rely on unsecured transport of DNS will not function in a DoH environment. The protocol encrypts and authenticates the server, but it does not tell the resolver who is a managed device.
The two positions are both partly right
- Security team: DNS logs and DNS firewalls (blocking known-bad domains, spotting command-and-control beaconing (malware contacting its operator's server at regular intervals), alerting on newly registered domains) only work if the queries pass through resolvers the company controls. A browser that silently sends its lookups to a public DoH service bypasses all of it.
- Users: On a coffee-shop or home network, plain DNS is readable and forgeable by anyone on the path, and encrypted DNS fixes that. The privacy gain is real against the local network and the ISP. It does not hide queries from the resolver operator, so it moves trust to a different party rather than removing it.
What I would propose
Steps 1 to 3 do the enforcing. Steps 4 to 6 make the setup smoother for clients and defensible to staff.
- Run an enterprise resolver that supports encrypted DNS. Clients on the corporate network and on VPN use it, over DoT or DoH, so the connection is protected from the LAN while the company keeps logging and filtering at the resolver. The resolver should also forward upstream over TLS. An Unbound example that passes
unbound-checkconf(the forwarder address is a documentation placeholder, which stands in for the vetted upstream you choose):
server:
interface: 127.0.0.4
tls-cert-bundle: /etc/ssl/certs/ca-certificates.crt
forward-zone:
name: "."
forward-tls-upstream: yes
forward-addr: 192.0.2.53@853#dns-upstream.example.net
Reading the block, line by line:
server:opens the global section.interface: 127.0.0.4is the address Unbound listens on for client queries (a loopback placeholder here; on a real resolver it is the address your clients are pointed at).tls-cert-bundlenames the file of trusted certificate authorities (CAs) that Unbound uses to check the upstream's TLS certificate. Without it the upstream cannot be authenticated.forward-zone:withname: "."means "the root", so every name that is not answered locally is forwarded rather than resolved by walking down from the root servers.forward-tls-upstream: yesmakes Unbound speak DoT (TLS) to the upstream instead of plain port 53.forward-addr: 192.0.2.53@853#dns-upstream.example.nethas three parts: the upstream's IP address,@853the port to connect to, and#dns-upstream.example.netthe name that the upstream's certificate must carry. The#namepart is what stops someone else answering on that IP: Unbound completes the TLS handshake only if the certificate is valid for that name.
unbound-checkconf printed unbound-checkconf: no errors in <file>, naming the file it checked. It checks syntax only; it does not contact the upstream.
- Configure managed endpoints by policy, and lock the settings. Firefox has a
DNSOverHTTPSenterprise policy withEnabled,ProviderURLandLockedkeys, so you can point it at the enterprise DoH endpoint and stop users changing it. Chrome'sDnsOverHttpsModepolicy takesoff,automatic(DoH first, with fallback to insecure) orsecure(DoH only, with no fallback, so resolution fails if the DoH server is unreachable), sosecurewith the enterprise template is the managed-device setting. An operating-system or browser setting that the user can change is not enforcement. - Block the unmanaged routes at the network edge. Deny outbound TCP 853 except from your resolvers (one firewall rule). For DoH, deny outbound connections to the public DoH providers you identify by address list or secure web gateway category, and refuse their hostnames at the enterprise resolver so a client that uses the system resolver to find them fails. A secure web gateway is a proxy that filters web traffic by site category, which is how "public DoH provider" can be a category you block. In the lab, an Unbound resolver configured with
local-zone: "dns.google." always_refuseand the same forcloudflare-dns.com.returned:
dns.google: status: REFUSED
cloudflare-dns.com: status: REFUSED
- Let clients discover the enterprise's own encrypted resolver. DDR (Discovery of Designated Resolvers, RFC 9462) lets a client that knows only the plain resolver address ask that resolver for the name
_dns.resolver.arpa.with record type SVCB (service binding, a record that says how to reach a service: target host name, protocol and port). The answer points to the same operator's encrypted endpoint. With a local-data entry in Unbound, the lab returned:
$ dig @127.0.0.1 _dns.resolver.arpa. SVCB +noall +answer
_dns.resolver.arpa. 300 IN SVCB 1 dns.corp.example. alpn="dot" port=853 ipv4hint=192.0.2.53
Read it as: priority 1, target host dns.corp.example., alpn="dot" the protocol (DoT), port=853, and ipv4hint an address to try without another lookup (all values are illustrative). The security rule is that for verified discovery the client must check that the encrypted resolver's certificate lists the IP address of the plain resolver it started from (RFC 9462 requires an iPAddress entry in the certificate's subjectAltName). The reason is that the plain query can be forged by anyone on the path, and a certificate naming that IP shows the owner of that address designated the encrypted endpoint. Advertising DDR means clients upgrade to your resolver rather than to a public one.
5. Give guests and personal devices a separate network with no promise of DNS visibility, so you are not trying to inspect devices you do not manage.
6. Write down the privacy rules. Say what the resolver logs, how long they are kept, who may query them, and that encrypted DNS is on for staff against outside observers. This turns a dispute into a documented trade-off.
Why not just block everything
Blocking 853 is cheap and reliable. Blocking DoH is a moving target: any HTTPS server can answer DNS queries, so address lists age, and TLS inspection (a gateway that decrypts and re-encrypts traffic using a certificate the company's devices trust) breaks pinned applications (apps that accept only their own expected certificate and so reject the gateway's) and adds cost and legal exposure. Block what is cheap (853, known provider hostnames and addresses), and rely on managed-endpoint policy for the rest. Some devices, such as unmanaged phones or applications that ship their own resolver, will still get through, and the answer to that gap is endpoint management, not more DNS rules.
Performance and privacy trade-offs
- Performance: A TLS or HTTPS connection costs extra round trips once, then queries reuse the connection, so steady-state latency is close to plain DNS. Enterprise resolver capacity must cover many long-lived TLS sessions. Cache hits at the enterprise resolver are faster than a cold lookup to a public service.
- Privacy: Employees' lookups are visible to the company on managed devices, which is a policy and legal decision, not just a technical one. Public encrypted DNS hides queries from the local network but hands them to the provider.
- Failure mode: With Chrome's
securemode, an outage of the enterprise resolver means no resolution at all, whileautomaticfalls back to insecure DNS. Choose consciously and monitor the resolver.
That is every published DNS, DHCP, and Name Resolution question for Cybersecurity Engineer so far. Browse the other topics in this category, or practice this one interactively.