Network Troubleshooting and Diagnostics Questions
Systematically diagnosing network problems: a layered troubleshooting methodology, diagnostic tools and commands (ping, traceroute, tcpdump, packet capture), root-cause analysis, and connectivity and user-access issues. Covers isolating faults across the stack, reading packet-level evidence, and driving from symptom to root cause. The diagnostic discipline that spans all networking layers.
A high-throughput web server is showing thousands of sockets in TIME_WAIT state and is approaching ephemeral port exhaustion. Using netstat/ss and tcpdump, describe how you would analyze the cause, then propose both a kernel-level and an application-level change you might make, explaining the safety considerations of each.
Sample Answer
Direct answer
Thousands of sockets in TIME_WAIT approaching ephemeral-port exhaustion on a high-throughput server means connections are being closed and cycled faster than TIME_WAIT's mandatory hold period can clear them. Analyze the cause with netstat/ss for the socket-state view, tcpdump for the wire-level closing behavior, and then propose kernel-level and application-level fixes, each with a real safety trade-off.
Structured elaboration
- Confirm the scale of the problem:
netstat -ant | grep TIME_WAIT | wc -l(orss -ant state time-wait | wc -l) against the total ephemeral port range (cat /proc/sys/net/ipv4/ip_local_port_range) shows how close to exhaustion the server actually is. - Identify which side is actively closing, using tcpdump: TIME_WAIT is held by whichever side sends the FIRST FIN (the active closer);
tcpdump -n -i eth0 host <peer-ip> and port <port> and 'tcp[tcpflags] & (tcp-fin|tcp-rst) != 0'captured during normal traffic shows directly whether this server or its clients/downstream peers are the ones initiating the close. If the SERVER is consistently sending the first FIN (common for a server that opens short-lived outbound connections to a downstream dependency, or that proactively closes idle client connections), the server itself is generating the TIME_WAIT accumulation; if the CLIENT is sending the first FIN, the TIME_WAIT burden sits on the client side instead and the server-side count is coming from somewhere else (like many short-lived server-to-downstream connections). - Understand why TIME_WAIT exists before tuning around it: it exists to ensure any delayed, in-flight packets from a just-closed connection aren't misinterpreted as belonging to a NEW connection that reuses the same 4-tuple shortly after; shortening or bypassing it carelessly reintroduces the exact correctness problem it was designed to prevent.
- Kernel tuning options and their safety considerations:
net.ipv4.tcp_tw_reuseallows reusing a TIME_WAIT socket for a new OUTGOING connection under safe conditions (using TCP timestamps to distinguish old from new), generally safe for outbound connections specifically;SO_REUSEADDRat the application level allows quickly rebinding a LISTENING socket after restart without waiting out TIME_WAIT, a narrower, different use case than tcp_tw_reuse and shouldn't be conflated with it. - Application-level changes, often more durable: shifting from short-lived, one-request-per-connection behavior to persistent, reused (keep-alive) connections directly reduces the rate of connection churn generating new TIME_WAIT entries, addressing the root cause the tcpdump evidence identified rather than just tuning around the symptom.
- Widen the ephemeral port range as an additional, complementary lever, giving more headroom before exhaustion without reducing the underlying churn rate.
Worked example
ss -ant | grep TIME_WAIT | wc -l shows roughly 28,000 sockets in TIME_WAIT against an ephemeral range of about 28,231 total ports, confirming the server is genuinely near exhaustion. tcpdump -n host <downstream-ip> and port 443 and 'tcp[tcpflags] & tcp-fin != 0' during normal operation shows THIS server consistently sending the first FIN on connections to a downstream dependency, confirming the server itself (acting as a client to that dependency) is the active closer generating the churn, rather than inbound client connections being closed by remote peers. Investigating further reveals the application opens a fresh outbound connection to that downstream dependency for every single request rather than reusing a connection pool. Implementing connection pooling for that downstream dependency reduces new-connection churn dramatically; widening the ephemeral port range is applied as an additional, complementary safety margin, not a substitute for fixing the churn.
Trade-offs & pitfalls
Applying tcp_tw_reuse without understanding what specific safety property it depends on (accurate TCP timestamps to distinguish genuinely old from new connections) risks reintroducing the exact correctness problem TIME_WAIT exists to prevent, in edge cases. Skipping the tcpdump step and relying on netstat/ss counts alone tells you HOW MANY sockets are stuck in TIME_WAIT but not WHICH SIDE is actively closing the connections that create them, which is exactly the fact that determines whether the fix belongs on this server (its own outbound connection pattern) or is really a client-behavior problem outside this server's control; always confirm the active-closer side with a capture before choosing where to apply the fix.
Users report DNS resolution succeeds (dig/nslookup returns valid answers) but a web application can't connect to the service by hostname. Describe the layer-by-layer checks you would perform to determine whether the issue is DNS, TCP/connectivity, firewall/ACL, or application binding. Include commands such as 'dig', 'nslookup', 'ss'/'netstat', 'curl' with IP and hostname, and advice on how to isolate name resolution from transport and application errors.
Sample Answer
Direct answer
When DNS resolves correctly but the application still can't connect by hostname, the fault is downstream of name resolution: most often the TCP connection itself, a firewall/security-group rule that's scoped to IP but not matching what DNS returned, or the application binding to the wrong interface or expecting a different hostname (TLS SNI or virtual-host routing).
Structured elaboration
- Confirm exactly what DNS returned:
dig(ornslookup) the hostname and note the resolved IP precisely. This matters because a common trap is DNS returning a valid but STALE or unexpected IP (an old load-balancer IP, or a different region than intended) that is technically "a valid answer" while being the wrong one for what you actually want to reach. - Test connectivity to that exact IP directly:
curlor a raw TCP test (nc) against the IP that DNS returned, on the target port. If this fails while the same test to a different, known-good IP succeeds, the problem has nothing to do with DNS: it's firewall, security-group, or routing scoped to that specific IP. - Test connectivity by hostname versus by IP, side by side: if connecting by IP succeeds but by hostname fails on an otherwise identical request, suspect something hostname-specific in the path: TLS SNI-based routing at a load balancer or reverse proxy that requires the correct hostname to route the connection at all, or an HTTP Host-header-based virtual-host that rejects requests presenting the "wrong" name.
- Check
ss/netstatand firewall/ACL state: confirm the client isn't blocked by a local host firewall, and that any network ACL or security group is scoped correctly (an ACL written against an old IP that DNS no longer returns, after the backend moved, is a common cause).
Worked example
dig myservice.internal returns 10.4.4.20. curl directly to https://10.4.4.20 times out; curl to 10.4.4.19 (an old, previously-working IP for the same service, still reachable) succeeds. This tells you immediately that the problem isn't the application, DNS, or TLS: it's that 10.4.4.20 specifically isn't reachable, most likely because a firewall rule or security-group entry was written against the old IP 10.4.4.19 and never updated when the service moved to a new IP as part of some other change.
Trade-offs & pitfalls
The most common mistake is treating "DNS resolves" as proof that DNS is not the problem; DNS can resolve successfully to the WRONG or STALE IP, which is a DNS-adjacent problem (caching, a record that wasn't updated) manifesting as a connectivity failure elsewhere. Always confirm what IP was actually returned and test that specific IP directly before concluding the fault lies purely in firewall or routing.
A service seems to be listening but a client can't connect. Using ss (or netstat) on the host, walk through how you'd confirm what's actually listening, on which address and port, and how you'd distinguish a loopback-only bind from one that's reachable externally, a process-ownership problem, and other local blockers (host firewall, SELinux, network namespace) from an actual network-path problem.
Sample Answer
Direct answer
Use ss (or the older netstat) to see exactly what's listening, on which address and port, and confirm whether the bind is scoped to loopback only or to all interfaces, since that single distinction explains a large share of it's-listening-but-nothing-external-can-connect reports. Beyond the bind address and process ownership, a separate class of local blockers, the host firewall, SELinux, and network namespaces, can each independently prevent an otherwise healthy listener from being reached, and each needs its own specific check rather than being lumped in with the network.
Structured elaboration
- List listening sockets with process ownership: ss -ltnp (listening, TCP, numeric, show process) lists every listening TCP socket along with the PID and process name holding it; this immediately answers whether anything is actually listening on this port, and whether it's the process you expect.
- Read the local address field carefully: 0.0.0.0:8080 means the process is listening on all interfaces and is reachable from outside the host (firewall permitting); 127.0.0.1:8080 means it is bound only to loopback and is fundamentally unreachable from any other host, no matter what firewall rules say, because the OS never even considers external interfaces for that socket.
- Check established connections and their state with ss -tn (without the -l) to see active connections; a large number stuck in SYN-RECV suggests the three-way handshake is not completing (possibly a firewall dropping the client's ACK, or a backlog queue issue on the server), while a large number in TIME_WAIT on a busy server is often benign churn rather than a problem, unless it is approaching ephemeral port exhaustion.
- Check the host firewall explicitly, as its own distinct layer from anything upstream: on iptables-based hosts,
iptables -L -n -v(oriptables -S) shows whether a rule is dropping or rejecting the port in question, and on firewalld-based hosts,firewall-cmd --list-allshows the active zone's allowed services and ports; a socket can be correctly listening on 0.0.0.0 and still be unreachable purely because the host's own firewall drops the inbound SYN before it ever reaches that socket. - Check SELinux (on systems that enforce it) as a separate, non-firewall local blocker:
getenforceconfirms whether SELinux is enforcing at all, and if it is, a service listening on a non-standard port that was never labeled for that port's SELinux port-type will be denied at the kernel security-module level even though the bind and firewall are both correct;ausearch -m avc -ts recent(or sealert on systems that have it) surfaces the specific denial, andsemanage port -l | grep <port>shows what port-type is currently associated with that port, withsemanage port -a -t <type> -p tcp <port>as the fix once the correct type is identified. - Check network namespaces on containerized or namespace-isolated hosts: a process can be listening perfectly well, but inside a different network namespace than the one you are inspecting from (for example inside a container's own namespace rather than the host's default namespace), so ss run in the wrong namespace will show nothing at all, not a loopback-only bind or a blocked port.
ip netns listenumerates namespaces on the host, andip netns exec <ns> ss -ltnp(ornsenter --net=<path> ss -ltnp) runs the same check inside the namespace that actually owns the socket; a 0.0.0.0 bind inside a container's namespace is still only reachable from outside according to whatever port-publishing or bridging rule (for example Docker's own iptables-based NAT rules) connects that namespace to the host and the network beyond it. - Distinguish not-listening-at-all from listening-but-blocked-further-out: if ss -ltnp shows nothing on the expected port in the correct namespace, the application itself never started or crashed, a fact entirely independent of firewalls, SELinux, or networking; if it does show a correct, externally-bound listener and external clients still cannot connect, the fault has moved to one of the local blockers above, or to something outside the host entirely.
Worked example
ss -ltnp shows LISTEN 0 128 127.0.0.1:8080 0.0.0.0:* users:(("myapp",pid=4521,fd=6)). The process is confirmed running and listening, but bound specifically to 127.0.0.1, loopback only; changing the bind address to 0.0.0.0 is the first fix, before any firewall or SELinux check is even relevant. Suppose instead ss -ltnp shows the same process correctly bound to 0.0.0.0:8080, and external clients still cannot connect. iptables -L -n -v on the host shows no DROP or REJECT rule referencing port 8080, ruling out the host firewall. getenforce reports Enforcing, and ausearch -m avc -ts recent shows an AVC denial for that process attempting to bind to port 8080, because the application was reconfigured to use a nonstandard port that was never added to SELinux's port-type list for that daemon; semanage port -a -t http_port_t -p tcp 8080 (or the appropriate type for the service) resolves it. In a third case, the process runs inside a container; ss -ltnp on the host shows nothing at all for port 8080, not because the process is not listening but because it is listening inside the container's own network namespace; docker exec <container> ss -ltnp (or ip netns exec <ns> ss -ltnp) confirms the process is listening correctly inside its namespace, and the actual question becomes whether the container's port-publishing rule correctly maps the host port to it.
Trade-offs & pitfalls
It is easy to see the process is listening in ss output and conclude the service is correctly exposed, without reading the bind address carefully enough to notice it is loopback-only, or without checking that you are even looking in the right network namespace. Also do not confuse ss -ltnp's absence of a listener with a firewall or SELinux problem; if nothing is listening in the namespace you are inspecting, no firewall rule or SELinux policy anywhere will make the connection succeed, and time spent checking those first is wasted until is-it-listening-and-where is ruled out. SELinux denials in particular are easy to miss because the application logs may show nothing useful at all (the bind or connection attempt is blocked below the application's own visibility), so always check ausearch or the audit log specifically once a correctly-bound, correctly-firewalled listener still is not reachable.
A DNS record was updated as part of a failover, but many clients still resolve the old IP. Explain the end-to-end places where DNS caching can occur (OS resolver, local DNS forwarder, CDN, browser), how you would identify which cache is serving stale results, and commands to force-cache checks and clears. Include how dig or nslookup could help isolate the layer serving stale data.
Sample Answer
Direct answer
DNS caching happens at several independent layers between a client and the authoritative record, and a failover's DNS update only takes effect for a given client once EVERY cache layer between them and the authoritative server has expired its old entry, so "many clients still resolve the old IP" usually means one specific layer's cache outlived the TTL you expected.
Structured elaboration
- Operating-system resolver cache: many OSes cache DNS answers locally, sometimes for the record's stated TTL, sometimes with their own separate minimum caching behavior that can outlast the record's actual TTL.
- Local DNS forwarder or corporate resolver: many networks route through an internal recursive resolver (a corporate DNS forwarder) that caches independently of any individual client, meaning a whole office or site can be affected by one shared, stale cache entry even if every individual laptop's own cache is clean.
- CDN or edge-caching layers: if the service sits behind a CDN, the CDN's own DNS resolution or its edge nodes' cached upstream mappings can lag independently of client-side DNS entirely.
- Browser-level DNS cache: many browsers maintain their own short-lived DNS cache separate from the OS, which can occasionally outlast a very short record TTL due to a browser-specific minimum caching floor.
- How to identify which layer is serving stale data:
digdirectly against the authoritative server (dig @<authoritative-ns> <name>) confirms what the CORRECT, current answer is;dig @<specific-resolver-ip> <name>against the specific resolver a client actually uses (matching what that client would get) reveals whether that resolver's cache is the one serving stale data; checking the TTL remaining on the STALE answer (rather than just its value) tells you roughly how much longer that specific layer's cache will keep serving it, which helps set expectations rather than guessing.nslookup <name> <resolver-ip>gives the same isolation for engineers more familiar with that tool. - Force-checking and clearing at each layer, with concrete commands per layer:
- Force-check a specific resolver:
dig @<resolver-ip> <name>ornslookup <name> <resolver-ip>. - Clear the OS resolver cache on macOS:
sudo dscacheutil -flushcache; sudo killall -HUP mDNSResponder. - Clear the OS resolver cache on Linux with systemd-resolved:
sudo resolvectl flush-caches(older systemd:sudo systemd-resolve --flush-caches); on systems using nscd:sudo systemctl restart nscd. - Clear the OS resolver cache on Windows:
ipconfig /flushdns. - Clear a browser's internal DNS cache: in Chrome/Edge, visit
chrome://net-internals/#dns(oredge://net-internals/#dns) and click "Clear host cache"; Firefox exposes an equivalent underabout:networking#dns. - For a corporate forwarder you don't administer directly, you typically cannot force a clear yourself; instead confirm its remaining TTL via
dig @<forwarder-ip>and escalate to whoever manages it if it's ignoring the record's stated TTL. - Confirm the record's TTL was actually set low enough BEFORE the failover (not just after), since a record kept at a long TTL right up until the failover means every cache that fetched it in the preceding window will hold the stale answer for that record's full original TTL, failover or not.
- Force-check a specific resolver:
Worked example
The DNS record was updated to point at the new IP as part of a failover. dig @<authoritative-ns> <name> correctly shows the new IP. A specific affected client resolves via a corporate DNS forwarder; dig @<forwarder-ip> <name> for the same name returns the OLD IP with 240 seconds of TTL remaining. This isolates the stale answer specifically to that corporate forwarder's cache, which fetched the old answer before the failover and, per its TTL, won't naturally refresh for another 4 minutes. On the affected client itself, running sudo dscacheutil -flushcache; sudo killall -HUP mDNSResponder (macOS) or ipconfig /flushdns (Windows) clears any locally cached copy immediately, though it cannot force the upstream corporate forwarder to drop its own entry before that entry's TTL naturally expires; other clients using a different resolver that happened to query AFTER the failover already see the new IP correctly.
Trade-offs & pitfalls
The common mistake is treating "the DNS record is updated" as equivalent to "clients see the new IP," without accounting for every caching layer between the client and the authoritative server independently holding onto the old answer for up to its own TTL. Clearing the OS or browser cache on an individual client only fixes that one client; it does nothing for a shared corporate forwarder or CDN edge node serving many other clients from the same stale entry, so identify which specific layer is actually responsible (via dig @<layer>) before assuming a local flush will resolve the wider complaint. For any service expecting to fail over, set the record's TTL LOW well BEFORE the planned failover so the actual cutover propagates as fast as the shortened TTL allows.
Your packet capture contains many IPv6 Neighbor Discovery issues. Explain how ND differs from ARP, which ICMPv6 messages are important for ND troubleshooting, and how you would use 'ip -6 neigh', tcpdump, and router advertisements to diagnose IPv6 neighbor reachability and autoconfiguration problems.
Sample Answer
Direct answer
IPv6 Neighbor Discovery (ND) does the same fundamental job as ARP, mapping a Layer 3 address to a Layer 2 (MAC) address, but it's built on ICMPv6 messages rather than a separate Layer 2 broadcast protocol, and it additionally handles router discovery and address autoconfiguration, tasks ARP itself never covered at all.
Structured elaboration
- Core mechanism difference: ARP uses its own dedicated Layer 2 broadcast protocol (ARP Request/Reply) with no relationship to ICMP; ND is built entirely on ICMPv6 message types, meaning the same general protocol family used for IPv6 error reporting and diagnostics also carries neighbor-resolution and router-discovery functions.
- Key ICMPv6 messages for ND specifically: Neighbor Solicitation (the ND equivalent of an ARP Request, asking "who has this address") and Neighbor Advertisement (the equivalent of an ARP Reply); additionally, Router Solicitation and Router Advertisement, which have NO ARP equivalent at all, since ARP never handled router or prefix discovery; Redirect messages exist in ND too, analogous to ICMPv4 Redirect.
- What ND additionally provides that ARP never did: Router Advertisements carry prefix information that hosts use for STATELESS address autoconfiguration (SLAAC), letting a host construct its own IPv6 address from an advertised prefix without needing DHCP at all; ARP has no equivalent concept, since IPv4 address assignment is handled entirely separately (statically or via DHCP), with ARP purely doing address-to-MAC resolution afterward.
- Diagnosing ND issues with tooling:
ip -6 neigh(the IPv6 analog ofip neigh/arp -n) shows the neighbor cache directly, with states analogous to ARP's (REACHABLE, STALE, FAILED, INCOMPLETE);tcpdump/Wireshark filtered foricmp6isolates the relevant ND traffic for capture-based analysis; watching Router Advertisements specifically confirms whether autoconfiguration information is being correctly and consistently advertised. - A specific ND-related failure mode worth checking for: inconsistent or conflicting Router Advertisements from MULTIPLE sources on the same segment (perhaps an unintended, unauthorized router or a misconfigured device also sending RAs) can cause hosts to receive conflicting prefix or default-router information, producing intermittent or inconsistent reachability that has no direct ARP-world analog, since ARP never involved anything resembling "which router should I use" in the first place.
Worked example
ip -6 neigh on an affected host shows an entry for the expected default gateway's link-local address in state FAILED, directly analogous to how an ARP entry in FAILED state would indicate the equivalent IPv4 problem. A capture filtered for icmp6 reveals TWO different devices on the segment sending Router Advertisements with conflicting prefix information, one legitimate, one from a misconfigured device that shouldn't be advertising at all; hosts intermittently picking up the illegitimate RA experience autoconfiguration and default-router inconsistency that manifests as exactly the kind of neighbor-reachability problem being investigated, even though the ROOT cause is upstream of neighbor discovery itself, in router advertisement, a category of problem ARP-based IPv4 networks simply don't have.
Trade-offs & pitfalls
Approaching an ND problem purely with ARP-troubleshooting habits (checking only the neighbor cache, as you would for ARP) can miss failure modes specific to ND's ADDITIONAL responsibilities, like conflicting Router Advertisements, since ARP never had anything analogous to diagnose in the first place; explicitly check Router Advertisement consistency as part of any thorough ND investigation, not just the neighbor cache entries themselves.
Unlock Full Question Bank
Get access to all 26 Network Troubleshooting and Diagnostics interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.