DNS, DHCP, and Name Resolution Questions
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
On a Linux server, an application calls a hostname. Before any DNS packet leaves the box, what decides where the answer comes from, and how would you work out why two processes on the same host resolve the same name differently?
Sample Answer
Direct answer
For almost every ordinary Linux program the answer is chosen by the C library's (glibc, the GNU C library) getaddrinfo() (the standard function a program calls to turn a name into addresses), which reads /etc/nsswitch.conf (the Name Service Switch, NSS). The hosts: line lists sources in order. On many servers that is files dns: first /etc/hosts, and only if that has no entry does the DNS stub resolver send packets to the servers in /etc/resolv.conf. Two processes on one host can disagree because they do not all take that path, or do not share the same inputs: a different resolver library (for example musl, a small alternative C library common in Alpine images), an environment variable, another mount or network namespace (a private copy of the network configuration, files and routes that the kernel gives each container), a process that started before a config change, or an application-level cache.
What happens before any DNS packet leaves
- The program calls
getaddrinfo("api.corp.test"). - glibc consults
nsswitch.confforhosts. Withfiles dns,/etc/hostsis searched first. A hit ends the search for that address family (getent hostsasks for IPv6 first, so in the lab a DNS query for the IPv6 record type still left even though the address printed came from the hosts file; the strace output in the diagnosis list shows it). Withdns filesthe order flips. - If
dnsis reached, the stub resolver applies/etc/resolv.conf: up to threenameserverlines (MAXNS, the library's built-in limit, is 3), thesearchlist, andoptions. A name with fewer dots thanndots(default 1) is tried with the search suffixes appended before being tried as given. Defaults fortimeout(5 seconds) andattempts(2) apply. - On a host running
systemd-resolved,resolv.confusually points at the local stub listener 127.0.0.53, ornss-resolvereplaces thednssource. Then the local caching resolver, not your upstream, produces the answer.
The proof, run in a Debian container with a local DNS server
The zone served locally says api.corp.test is 10.1.1.1; /etc/hosts says 10.9.9.9; resolv.conf has search corp.test.
$ grep '^hosts' /etc/nsswitch.conf
hosts: files dns
$ getent hosts api.corp.test # goes through NSS, like most programs
10.9.9.9 api.corp.test
$ dig +short api.corp.test # talks to DNS only, ignores /etc/hosts
10.1.1.1
$ getent hosts api # search list applied
10.1.1.1 api.corp.test
$ dig +short api # dig does not use the search list by default
$ LOCALDOMAIN=other.test getent hosts api ; echo "exit=$?"
exit=2
$ sed -i 's/^hosts:.*/hosts: dns files/' /etc/nsswitch.conf
$ getent hosts api.corp.test
10.1.1.1 api.corp.test
Three different answers from one machine: the hosts file wins under files dns, dig skips NSS entirely, and one process with LOCALDOMAIN in its environment cannot expand the short name. getent hosts exits with status 2 when the key is not found.
Why exit=2: LOCALDOMAIN is an environment variable that replaces the search line of resolv.conf for that one process. With LOCALDOMAIN=other.test the short name api is tried as api.other.test, which does not exist, and then as plain api, which does not exist either, so the lookup fails. The same command with a search domain that contains the record works:
$ LOCALDOMAIN=corp.test getent hosts api ; echo "exit=$?"
10.1.1.1 api.corp.test
exit=0
So two processes that differ only in their environment can disagree about a short name while both are behaving correctly.
Diagnosing two processes that disagree (in this order)
Steps 1 and 2 settle the common cases, so run them first; steps 3 to 7 are for when the files agree but the processes still differ. In steps 3 to 7, PID is the process ID of the application (find it with pgrep or ps), and /proc/PID is the kernel's live view of that process.
- Reproduce with a tool that follows the same path as the application.
getent hosts NAMEuses the NSS path.dig,hostandnslookuptalk to DNS directly and ignore/etc/hostsand NSS order. Ifgetentanddigdiffer, the cause is local (hosts file, NSS order, search list), not the DNS servers. - Read the static inputs.
grep '^hosts' /etc/nsswitch.conf,grep NAME /etc/hosts, andcat /etc/resolv.conf(alsols -l /etc/resolv.confto see whether it is a symlink managed bysystemd-resolved). On such hosts useresolvectl statusto see the real servers per interface andresolvectl flush-cachesto clear its cache. - Compare the two processes' environments.
tr '\0' '\n' < /proc/PID/environ | grep -E 'LOCALDOMAIN|RES_OPTIONS|HOSTALIASES'.RES_OPTIONScan amend theoptionsline per process, andLOCALDOMAINreplaces the search list. The file holds the process's variables separated by NUL bytes, whichtrturns into lines. For a process started withLOCALDOMAIN=other.test RES_OPTIONS=ndots:3the lab printedRES_OPTIONS=ndots:3andLOCALDOMAIN=other.test. No output means the process has none of them. - Compare their world.
readlink /proc/PID/ns/netandreadlink /proc/PID/rootshow whether a process sits in a different network namespace or chroot (a changed root directory, so the process sees a different/etc). The first prints an identifier such asnet:[4026533483]: compare it withreadlink /proc/$$/ns/netfor your own shell, and equal numbers mean the same network namespace (in the lab both printednet:[4026533483]). The second prints/for a process with the normal root. A containerised process reads the container's own/etc/hostsand/etc/resolv.conf, not the host's. - Check which resolver library the program uses.
file $(readlink -f /proc/PID/exe)andlddreveal a static binary or a musl-based image. In the labfileprintedELF 64-bit LSB pie executable, ARM aarch64, ... dynamically linked, interpreter /lib/ld-linux-aarch64.so.1, andlddlistedlibc.so.6 => /lib/aarch64-linux-gnu/libc.so.6: a dynamic binary using glibc, so it followsnsswitch.conf. A musl image would name an interpreter such as/lib/ld-musl-x86_64.so.1, and a statically linked binary would printstatically linkedand have nolibc.so.6line inldd. Go programs use a pure-Go resolver by default;GODEBUG=netdns=go+1ornetdns=cgo+1prints which one is in use and forces either. Java and some other runtimes keep their own address cache with their own lifetime. - Check the timeline. glibc 2.33 and later reload
nsswitch.confwhen it changes; before 2.33 the file was read once per process, so a long-running daemon keeps the old order until restarted. A process started before an/etc/hostsorresolv.confedit and one started after can legitimately differ if it caches. - Trace the actual calls when it is still unclear.
strace -f -e trace=openat,connect,sendto -p PIDshows which files are opened and which servers receive packets (straceprints every system call a process makes;-p PIDattaches to a running process, and the same options in front of a command run it under strace instead). On the lab'sgetent hosts api.corp.test,strace -f -e trace=openat,connectprinted these lines of interest:
openat(AT_FDCWD, "/etc/resolv.conf", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/nsswitch.conf", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/hosts", O_RDONLY|O_CLOEXEC) = 3
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
Read it in order: the resolver config was read, then the switch file, then /etc/hosts, then a connection to port 53 of 127.0.0.1 shows a DNS query was also sent. With dns files in the switch file the /etc/hosts line moves after the connect lines.
The fix is to remove the divergence: put one record in one place, align NSS order across hosts through configuration management, and restart the stale process.
Trade-offs and pitfalls
- A stale
/etc/hostsentry added for a migration is the most common cause of "works in dig, fails in the app". Search hosts files first whendigand the application disagree. - Do not diagnose application resolution with
digalone, and do not assumenslookupshows what the program sees. - A forgotten
searchdomain makes short names behave differently inside containers, because the container image or the orchestrator writes its ownresolv.conf.
An application reports intermittent name lookup failures while command-line checks work fine. How would you observe what the process itself is doing, and how would you decide between local resolver configuration, the DNS server, and an application bug?
Sample Answer
Direct answer
dig is not the application's resolver, so "dig works" proves little. dig reads no /etc/hosts, ignores /etc/nsswitch.conf, and by default skips the search list. A program on Linux usually calls getaddrinfo(), the C library function that turns a name into addresses, which does all three (it consults /etc/hosts, follows /etc/nsswitch.conf, the file that lists where names are looked up and in what order, and applies the search list, the domain suffixes appended to short names). The method is: reproduce the application's own lookup path with getent (getent ahosts NAME calls getaddrinfo() itself, while getent hosts NAME, used in the traces below, calls the older gethostbyname2(); both go through nsswitch.conf, /etc/hosts, the search list and resolv.conf, so they read the same files and servers, but hosts asks for the IPv6 record and then the IPv4 record one after the other), watch what the real process does with strace and a packet capture, then use the evidence to choose among three suspects: local resolver configuration, the DNS server, or the application itself.
Step 1: see the difference between dig and the library
A tiny lab: a local resolver on 127.0.0.1 that knows one name, plus an /etc/hosts entry and a search domain. Setup (Linux, run as root in a throwaway container). The dnsmasq flags: --keep-in-foreground stay in the terminal instead of forking away; --port=53 --listen-address=127.0.0.1 --bind-interfaces listen only on loopback; --no-resolv do not read resolv.conf for upstream servers and --no-hosts do not load /etc/hosts, so the lab knows only what we tell it; --address=/www.example.com/203.0.113.10 answer that name with that address (an illustrative documentation address).
dnsmasq --keep-in-foreground --port=53 --listen-address=127.0.0.1 --bind-interfaces \
--no-resolv --no-hosts --address=/www.example.com/203.0.113.10 \
--log-queries --log-facility=/lab-dnsmasq.log &
echo '10.9.9.9 legacy-app' >> /etc/hosts
printf 'search example.com\nnameserver 127.0.0.1\n' > /etc/resolv.conf
$ getent hosts legacy-app
10.9.9.9 legacy-app
$ dig +short legacy-app
$ getent hosts www
203.0.113.10 www.example.com
$ dig +short www
$ dig +search +short www
203.0.113.10
getent finds the name from /etc/hosts and expands the short name www through the search list; plain dig returns nothing for both. A check that only uses dig can be green while the program is red, or the reverse.
Step 2: watch what the process itself does
strace prints every system call a process makes: a system call (syscall) is a request from a program to the operating system kernel, such as opening a file, connecting a socket or sending a packet. Attach to the process, or start it under strace, and follow children (-f). Here the same trace on getent (library-loading lines left out):
$ strace -f -e trace=openat,connect,sendto -s 60 getent hosts www
openat(AT_FDCWD, "/etc/host.conf", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/resolv.conf", O_RDONLY|O_CLOEXEC) = 3
connect(3, {sa_family=AF_UNIX, sun_path="/var/run/nscd/socket"}, 110) = -1 ENOENT (No such file or directory)
connect(3, {sa_family=AF_UNIX, sun_path="/var/run/nscd/socket"}, 110) = -1 ENOENT (No such file or directory)
openat(AT_FDCWD, "/etc/nsswitch.conf", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/hosts", O_RDONLY|O_CLOEXEC) = 3
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
sendto(3, ")\260\1\0\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\34\0\1", 33, MSG_NOSIGNAL, NULL, 0) = 33
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
sendto(3, ")\260\1\0\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\34\0\1", 33, MSG_NOSIGNAL, NULL, 0) = 33
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
sendto(3, "\254e\1\0\0\1\0\0\0\0\0\0\3www\0\0\34\0\1", 21, MSG_NOSIGNAL, NULL, 0) = 21
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
sendto(3, "\254e\1\0\0\1\0\0\0\0\0\0\3www\0\0\34\0\1", 21, MSG_NOSIGNAL, NULL, 0) = 21
openat(AT_FDCWD, "/etc/hosts", O_RDONLY|O_CLOEXEC) = 3
connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
sendto(3, "\332$\1\0\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\1\0\1", 33, MSG_NOSIGNAL, NULL, 0) = 33
Reading it in order: the library reads its settings (host.conf, resolv.conf), tries the name-service cache daemon nscd (absent here, so ENOENT, "no such file", which is harmless), reads nsswitch.conf and /etc/hosts, and only then sends DNS packets, always to the nameserver in resolv.conf. In this run each query appears twice in a row with the same ID.
The sendto payload is the raw DNS packet, written with C-style escapes (\1 is byte 1, \260 is octal for byte 0xB0, and a printable character stands for itself). Decoding the first one: )\260 is the 2-byte query ID; \1\0 is the flags, with only the recursion-desired bit set; \0\1 says 1 question; the next six zero bytes are zero answer, authority and additional records; then the name, as length-prefixed labels (\3www is 3 then "www", \7example, \3com, and \0 ends the name); \0\34 is the type: \34 is octal for 28, which is AAAA (the IPv6 address record); \0\1 is class IN. That is 12 + 17 + 4 = 33 bytes, matching the 33 in the call. The three distinct queries are: AAAA for www.example.com (the search domain applied), AAAA for the bare www (type \34), and finally type 1, A (the IPv4 record), for www.example.com.
In a failing process, add recvfrom (receive a reply) and poll to the traced calls. Compare a healthy path with a lab where the first nameserver is dead (192.0.2.1, a documentation address that nothing answers) and the second is healthy, with options timeout:1 attempts:1. In this trace /etc/resolv.conf holds only those two nameserver lines and the options line, with no search line (with search example.com kept, the lab's dnsmasq answers the AAAA question with an error code, so the library also tries www.example.com.example.com, and the same lookup measured three queries and about three seconds):
$ strace -f -tt -e trace=sendto,recvfrom,connect -s 60 getent hosts www.example.com
08:32:15.133527 connect(3, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("192.0.2.1")}, 16) = 0
08:32:15.133550 sendto(3, "(\365\1\0\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\34\0\1", 33, MSG_NOSIGNAL, NULL, 0) = 33
08:32:16.137738 connect(4, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, 16) = 0
08:32:16.137847 sendto(4, "(\365\1\0\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\34\0\1", 33, MSG_NOSIGNAL, NULL, 0) = 33
08:32:16.138034 recvfrom(4, "(\365\201\205\0\1\0\0\0\0\0\0\3www\7example\3com\0\0\34\0\1", 1024, 0, {sa_family=AF_INET, sin_port=htons(53), sin_addr=inet_addr("127.0.0.1")}, [28 => 16]) = 33
The line that proves a failure is the one that is missing: a sendto to 192.0.2.1 at .133550 with no recvfrom from that server, followed 1.004 seconds later (the timeout:1) by the same question sent to the next server, which does answer. Only the first query (the AAAA) is shown. The A query that follows repeats the pattern: a send to 192.0.2.1, one second of silence, then a send to 127.0.0.1 and its reply. So this one lookup took about two seconds, which is the symptom the application sees. A trace of a healthy path has a recvfrom shortly after every sendto.
Questions to answer from the trace: which resolv.conf did it open (a container or a chroot, which gives a process a different root directory, may see a different file than your shell; a container's mount namespace is its own private view of the filesystem, and cat /proc/PID/root/etc/resolv.conf shows the file as that process sees it), which server did it contact, and in what order, and did a reply ever come back (recvfrom lines) before the program gave up.
Add a packet capture beside it, tcpdump -ni any port 53, and match the query IDs: a query with no reply is the server or path; a reply with an error code (SERVFAIL, NXDOMAIN) is the server's answer; no query at all while the program reports a failure is the application, a cached failure, or a different resolver such as a built-in one.
Step 3: separate the three suspects
| Evidence | Points to |
|---|---|
getent hosts NAME in a loop also fails sometimes | Local configuration or the DNS server, not the program |
getent always passes, the program still fails, and the trace shows no query for that name | Application: its own resolver, an in-process cache that holds a failure, or a connection pool using a stale address |
Queries leave for the first nameserver and some get no reply, while the second server answers | Dead or unreachable first server in resolv.conf (local configuration) |
dig @EACH_SERVER NAME in a loop: one server returns SERVFAIL or times out | That DNS server |
The strace shows a different resolv.conf or connect target than the shell uses | Container or namespace resolver configuration |
Loop to catch the intermittent case without waiting for users: for i in $(seq 1 100); do getent hosts NAME >/dev/null || echo "fail $i"; done, and the same with dig @SERVER +tries=1 +time=2 NAME +short once per server.
Search list and ndots, with a measured example
search example.com in resolv.conf is the search list: suffixes the library appends to a name. options ndots:N decides the order: a name with fewer than N dots is tried with each search suffix first and as typed last; a name with at least N dots is tried as typed first (default N is 1; resolv.conf(5)). Each failed attempt is a full extra query, so a slow or failing resolver multiplies the cost of a lookup. With the lab's dnsmasq logging queries (the --log-queries --log-facility flags above), looking up a.b.c (2 dots) gave this order of A queries:
-- ndots:1, lookup of a.b.c (2 dots):
a.b.c
a.b.c.example.com
-- ndots:5, lookup of a.b.c (2 dots):
a.b.c.example.com
a.b.c
and getent hosts www (0 dots) asked only for www.example.com among the A queries, because that name answered. Its AAAA queries went first to www.example.com and then to the bare www, as the first trace in this answer shows, because the AAAA answer for the suffixed name was empty. The log was read with grep "query\[A\]" /lab-dnsmasq.log | awk '{print $6}' | awk '!s[$0]++': keep the A queries, print the name column, and drop repeats while keeping order.
Why one dead server looks "intermittent"
The C library queries the servers in order of the nameserver lines (at most three; MAXNS is the library's constant for that limit), waits timeout seconds (default 5) and retries attempts times (default 2) as described in resolv.conf(5). The rotate option makes it alternate servers. To see the effect, one process makes 20 lookups, with 127.0.0.1 healthy and 192.0.2.1 (a documentation address where nothing answers) first:
# loop.py
import socket, time
times = []
for i in range(20):
t = time.perf_counter()
socket.getaddrinfo("www.example.com", 80, socket.AF_INET)
times.append(time.perf_counter() - t)
slow = sum(1 for t in times if t > 0.5)
print("".join("S" if t > 0.5 else "." for t in times), "(S = slow lookup)")
print(f"{slow} of 20 lookups in ONE process took longer than 0.5 s")
time.perf_counter() is a clock, and each lookup is marked S (slower than 0.5 s) or . (fast) in the first output line:
--- options timeout:1 attempts:1
SSSSSSSSSSSSSSSSSSSS (S = slow lookup)
20 of 20 lookups in ONE process took longer than 0.5 s
--- options timeout:1 attempts:1 rotate
S.S.S.S.S.S.S.S.S.S. (S = slow lookup)
10 of 20 lookups in ONE process took longer than 0.5 s
(/etc/resolv.conf was nameserver 192.0.2.1, nameserver 127.0.0.1, then the options line shown.) Without rotate, every lookup pays the dead server's timeout first and then succeeds. That is a steady slowdown, not yet an intermittent fault. With rotate the library starts each lookup at the next server in turn (round-robin), so lookups alternate between starting at the dead server (slow) and starting at the healthy one (fast), which is the alternating pattern: exactly half of 20 is 10. Which phase it starts in varies from run to run: the block above starts with a slow lookup, and re-runs on the same lab printed .S.S.S. starting with a fast one. An application with a 500 ms deadline then fails every other lookup, which is what a user calls intermittent. The same arithmetic produces intermittent failures whenever lookups are spread over a dead and a healthy server: a load balancer in front of several resolver nodes of which one is down, or a fleet of hosts whose first nameserver differs. A single getent run sends only a few queries, so one run can look fine, or only one or two seconds slow, while a long-running process fails half its lookups against a 500 ms deadline. In a lab run with rotate, ten separate getent runs took between 1.0 and 2.0 seconds each, depending on which server each of the run's two queries (AAAA and A) started at, so loop the check and count, as in the loop above.
Fixes by suspect
- Local configuration: remove or replace the dead server, order servers best first, set
timeout:1 attempts:2where the application deadline is short, checkndotsand the search list for extra queries (see the measured example above: a lowerndotsor fully qualified names with a trailing dot skip the suffix attempts). - DNS server: fix or take out of rotation the failing server, then verify with the loop above.
- Application: use the system resolver or configure its cache and timeouts, and retry failed lookups with a short backoff.
Pitfalls
Do not trust a single dig or a single getent; intermittent faults need a loop and counts. Do not restart the application before capturing a trace, since the restart clears the state you want to look at.
That is every published DNS, DHCP, and Name Resolution question for Systems Engineer so far. Browse the other topics in this category, or practice this one interactively.