Network Troubleshooting and Diagnostics Questions
Systematically diagnosing network problems: a layered troubleshooting methodology, diagnostic tools and commands (ping, traceroute, tcpdump, packet capture), root-cause analysis, and connectivity and user-access issues. Covers isolating faults across the stack, reading packet-level evidence, and driving from symptom to root cause. The diagnostic discipline that spans all networking layers.
You need to process a set of packet captures (or a live traffic feed) to answer a concrete question, e.g. which flows are retransmitting the most, or what the handshake latency looks like per flow, without opening each one by hand. Describe how you'd script this (naming the library or tool you'd reach for), what fields you'd extract, and what you'd have to be careful of, like out-of-order packets or multiple capture points, for the numbers to be trustworthy.
Sample Answer
Direct answer
For a concrete question like "which flows are retransmitting the most" or "what's the per-flow handshake latency," script it with a pcap-parsing library (Python's scapy, or shelling out to tshark's field-extraction mode) rather than opening captures by hand; the key engineering concerns are correctness (handling out-of-order and duplicate packets) and scale (not loading an entire large capture into memory at once).
Structured elaboration
- Choose the extraction path:
tshark -r file.pcap -T fields -e ip.src -e ip.dst -e frame.lenstreams field values without building Python objects for every packet, and scales best for very large captures; scapy'sPcapReader(as opposed tordpcap) reads packets one at a time, which is the right choice when you need custom logic per packet in Python. - Aggregate incrementally: keep running totals in a dictionary keyed by the flow tuple (source IP, destination IP, and for TCP also ports), rather than storing every packet, so memory use stays flat regardless of capture size.
- Handle multiple capture points and out-of-order packets explicitly: if flows are captured at more than one point in the path, do not assume packet order in the file matches wall-clock order; sort by timestamp within each flow before computing anything that depends on ordering (like handshake latency, which needs to match a SYN to its SYN-ACK, not just the Nth and N+1th packet).
- Validate against a known-answer capture before trusting the script's output on real production data.
Worked example
Here is a minimal top-talkers-by-bytes script, executed against a synthetic 60-packet capture (3 flows) built with scapy for this validation:
from collections import defaultdict
from scapy.all import PcapReader, IP
def top_talkers(path, n=10):
counts = defaultdict(int)
byte_totals = defaultdict(int)
with PcapReader(path) as reader:
for pkt in reader:
if IP in pkt:
key = (pkt[IP].src, pkt[IP].dst)
counts[key] += 1
byte_totals[key] += len(pkt)
return sorted(byte_totals.items(), key=lambda kv: kv[1], reverse=True)[:n]
Run against the test capture, this correctly reported two flows: 10.1.1.10 -> 10.2.2.20 with 40 packets totaling 28,700 bytes, and 10.1.1.11 -> 10.2.2.21 with 20 packets totaling 11,700 bytes, matching the known composition of the synthetic file exactly. For handshake-latency specifically, the same pattern applies but keyed by the full 4-tuple, matching each SYN to the SYN-ACK sharing its source/destination ports and computing the timestamp delta.
Trade-offs & pitfalls
Loading an entire multi-gigabyte capture with rdpcap (which reads the whole file into memory as a list) is a common and expensive mistake; use a streaming reader instead. When captures come from multiple points, do not assume a packet appearing in one capture but not another means it was dropped; capture loss at the tap/collection point itself is a real, separate failure mode from network packet loss, and conflating the two produces wrong conclusions.
You have a pcap file from a failing request and need to analyze it with Wireshark to find the root cause. Explain how you would use 'Follow TCP Stream', filtering, and common Wireshark statistics (IO graph, Expert Info, and TCP flow graph) to identify whether the failure is due to application error, TCP retransmission, or malformed packets. Describe the sequence of checks you would perform after opening the capture.
Sample Answer
Direct answer
Given a pcap from a failing request, work from the general to the specific in Wireshark: first use a display filter to isolate just the relevant conversation out of what may be a much larger capture, then use "Follow TCP Stream" to see the whole conversation as a human-readable transcript, then Expert Info to surface anything Wireshark itself flags as anomalous, then the IO graph and TCP flow graph if timing (not just content) is the suspected issue.
Structured elaboration
- Filter down to the relevant stream first: before doing anything else, use a display filter such as
ip.addr == <client-or-server-ip> && tcp.port == <port>, or right-click a packet belonging to the failing exchange and choose "Conversation Filter", then "TCP", to reduce a large capture to just the one stream involved; working against a filtered view avoids wasting time scrolling through unrelated traffic and makes every later step (Follow Stream, Expert Info, the graphs) operate on the right data instead of the whole file. If the specific stream isn't obvious yet, filter more broadly first (for exampletcp.flags.reset == 1to find resets, orhttp.response.code >= 400for application errors) to locate it. - Follow TCP Stream: right-click any packet in the filtered flow and choose Follow, then TCP Stream; this reconstructs the full request and response as a readable transcript (for plaintext protocols; for TLS, you'd only see this post-decryption if you have the session keys), which is usually the fastest way to see WHAT was actually exchanged and where in the conversation things went wrong, without manually piecing together individual packets.
- Expert Info: Analyze menu, then Expert Information; Wireshark automatically flags things like retransmissions, duplicate ACKs, zero windows, and malformed packets across the entire capture (or just the filtered subset, if you apply the same display filter first), giving you a prioritized starting list of anomalies rather than requiring you to scroll through every packet manually looking for something suspicious.
- IO graph: a visual plot of packets or bytes over time, which can also be restricted to the filtered stream; useful for spotting a sudden drop-off (suggesting the connection stalled at a specific moment) or a burst pattern that might correlate with the failure's timing.
- TCP flow / time-sequence graph: plots sequence numbers over time for a specific stream; a flat line (no forward progress in sequence numbers) during a stall visually confirms data genuinely stopped moving, while a step pattern with occasional gaps visually distinguishes steady progress with periodic retransmission from a complete stall.
- Sequence of checks for this specific scenario: filter down to the one relevant stream first; then use Follow Stream to see what was actually asked for and what response (if any) came back; if the transcript looks incomplete or cuts off, check Expert Info for retransmissions or resets around that point; if the transcript looks complete but slow, use the TCP flow graph to see whether the delay was in the network (gaps between segments) or entirely at the application layer (data flowing steadily, then a large gap before the response actually starts).
Worked example
A large capture is filtered down to tcp.port == 443 && ip.addr == 10.0.4.22 to isolate just the failing request's stream. Follow TCP Stream on that filtered result shows a complete HTTP request sent, followed by nothing at all on the response side, cutting off abruptly. Expert Info (run against the same filtered stream) flags multiple retransmissions and a final RST (reset) from the server side around the same timestamp. The TCP flow graph confirms: data flows normally in the request direction, then the connection simply stops advancing before a reset appears. Together, this points at the server actively terminating the connection partway through processing (a RST, not a clean close, and not a timeout), which narrows the investigation toward the server side (an application crash, a proxy timeout terminating the backend connection) rather than a pure network-path issue.
Trade-offs & pitfalls
Skipping the filtering step and reading raw packet-by-packet output across an entire large capture before narrowing scope wastes time these features exist specifically to save; filter to the relevant stream, then start broad within it (transcript, then automated anomaly flags) before manually inspecting individual packets. Also remember Follow TCP Stream shows plaintext content only; for TLS traffic without the session keys, you'll see encrypted bytes, and the analysis has to rely on the handshake metadata and Expert Info's flow-level observations instead.
Your network team suspects a failing optical link causing CRC errors and packet corruption. Describe how to detect and confirm physical-layer errors from host-level observations (interface counters, 'ethtool -S', dmesg), what capture patterns indicate corruption (e.g., malformed packets, checksum offload artifacts), and how to distinguish real on-wire corruption from checksum offloading at the host.
Sample Answer
Direct answer
Confirm a failing optical link from host-level observations before assuming it's the fiber itself: interface error counters and OS-level tooling can distinguish real on-wire corruption from an artifact of checksum offloading, and the specific error PATTERN (not just its presence) points toward an optical problem versus something else.
Structured elaboration
- Check interface error counters for the specific signature: rising CRC (cyclic redundancy check) errors combined with input errors, without a corresponding rise in collisions or duplex-mismatch-typical symptoms, points toward physical-layer corruption rather than a Layer 2 contention issue.
- Use
ethtool -S(or the vendor equivalent) for detailed, per-interface statistics beyond the basic counters, including optical-specific diagnostics on many NICs and switches (transmit/receive power levels via DOM/DDM) that directly indicate a degrading optical link, such as receive power trending down over time even before errors become severe. - Check
dmesgfor kernel-level link-flap or hardware-error messages that might correlate with the timing of reported corruption, since a marginal optical link can produce brief link resets that a simple counter check might miss if not actively watching in real time. - Distinguish real corruption from a checksum-offload artifact: modern NICs often perform checksum calculation in hardware and can, in some capture configurations, report a packet as having a "bad checksum" in a tool like Wireshark simply because the capture was taken BEFORE the NIC's hardware checksum offload actually computed the real value (the OS hands the packet to the capture mechanism with a placeholder checksum, expecting the NIC to fill it in later); this produces what LOOKS like universal corruption in a capture, while the actual on-wire packets are fine. The key distinguishing test: if EVERY packet from a specific host shows a "bad checksum" with otherwise normal-looking, well-formed content, and the application actually works some of the time despite the capture claiming universal corruption, suspect an offload artifact in how the capture was taken rather than believing every single packet is genuinely corrupted; genuine physical-layer corruption, by contrast, typically produces a mix of some healthy and some visibly malformed packets, correlated with rising CRC counters on the actual interface.
Worked example
A capture on a host shows nearly every outbound packet flagged "bad checksum" by Wireshark. ethtool -S on that host's interface shows a modest, non-zero, non-escalating CRC error count and no other unusual counters. Because the "bad checksum" flag appears on essentially ALL packets uniformly (not correlated with the smaller number of real CRC errors reported by the interface itself), and the application is functioning normally, this is consistent with a checksum-offload artifact in the capture (the NIC computes the real checksum after the capture point, in hardware), not genuine corruption; disabling checksum offload temporarily for a diagnostic capture, or checking the "checksum validation" setting in the capture tool, would confirm this by making the flag disappear once the capture sees the NIC's actual final checksum.
Trade-offs & pitfalls
Treating every "bad checksum" flag in a capture as proof of real corruption is a very common and avoidable misdiagnosis, especially on modern hosts where checksum offload is the default; always cross-check against the interface's own CRC error counters (which reflect genuine on-wire corruption) before concluding the wire itself is at fault. Conversely, don't dismiss a genuine, rising CRC counter as "probably just offload," since that pattern specifically indicates a real physical-layer problem that offload artifacts do not produce.
Explain the differences between ping, traceroute, and mtr. For each tool, describe what types of symptoms they help identify, what their outputs mean (ICMP TTL expiry vs ICMP unreachable vs UDP/TCP-based traceroute), and how you would use their results to progress your troubleshooting.
Sample Answer
Direct answer
Ping, traceroute, and mtr answer three different questions: ping tells you whether a destination is reachable and roughly how long a round trip takes; traceroute tells you the path packets take and where along that path something stops responding; mtr combines both, continuously, so you can see loss and latency per hop over time rather than a single snapshot. Reading their outputs correctly also means knowing which ICMP message each tool relies on and how the underlying probe type (ICMP, UDP, or TCP) changes what success looks like.
Structured elaboration
- Ping: sends ICMP echo requests and waits for ICMP echo replies; a successful reply confirms basic reachability and round-trip time; ping tells you nothing about where along the path a problem is, only whether the whole round trip succeeded.
- Traceroute: sends probes with increasing TTL (starting at 1) so each successive router along the path replies with an ICMP time-exceeded message as its TTL hits zero, revealing the path hop by hop. Classic Unix/Linux traceroute does this with UDP probes to a high, deliberately unlikely-to-be-open destination port; when a probe finally reaches the real destination, that host replies with ICMP destination-unreachable, port-unreachable, which is the normal, expected signal that tells traceroute it has reached the end of the path, not an error. Windows' tracert instead uses ICMP echo probes throughout, so the final hop returns a normal ICMP echo reply rather than a port-unreachable message. A TCP-based traceroute variant (tcptraceroute, or traceroute/mtr run in TCP mode) sends TCP SYN segments with increasing TTL instead; this is useful specifically when UDP or ICMP probes are filtered somewhere in the path (a common firewall policy) but the actual TCP port you care about is open, since a SYN-based probe follows the exact path and protocol handling that real application traffic would experience.
- mtr: repeatedly sends traceroute-style probes to every hop continuously and aggregates loss percentage and latency statistics per hop over the run, rather than a single traceroute's one-shot snapshot; this makes it far better at catching intermittent loss at a specific hop, which a single traceroute would likely miss entirely by chance. mtr can typically be run in ICMP, UDP, or TCP probe mode, inheriting the same completion-signal differences described above depending on which mode is selected.
- How to progress your troubleshooting using their results: start with ping to confirm there is a problem at all and get a rough sense of loss/latency; if there is a problem, use mtr (not just one traceroute) to localize which hop it correlates with, since a single traceroute run can easily miss an intermittent issue or be misled by ECMP; and if UDP- or ICMP-based probes show the path is blocked somewhere but you suspect that is only true for those protocols, rerun in TCP mode against the actual port in question to see whether the real application traffic's path differs.
Worked example
A user reports intermittent slowness to a remote service. A single ping shows occasional but not consistent high latency. A single UDP-based traceroute shows all hops responding normally with a final ICMP port-unreachable from the destination, confirming it reached the end of the path successfully; this happened to sample during a good moment, so it does not show the intermittent issue. Running mtr for several minutes reveals hop 6 specifically showing 15% loss and elevated latency, while every other hop shows 0% loss consistently, precisely the kind of transient, hop-specific pattern a one-shot traceroute is likely to miss by chance but that mtr's continuous sampling reliably catches. Separately, a different host appears completely unreachable by traceroute (every hop past a firewall shows only asterisks), while the actual web service on that host works fine in a browser; switching to a TCP-mode traceroute on port 443 succeeds cleanly, confirming that ICMP and UDP probes are being filtered by policy while the real TCP path is healthy.
Trade-offs & pitfalls
Treating a single clean traceroute as proof there is no path problem is a common mistake when the issue is intermittent; mtr's continuous sampling exists specifically to catch what a one-shot traceroute would miss. Also remember that ICMP responses (used by ping and, often, traceroute/mtr) can be deprioritized or rate-limited by routers relative to real data traffic, so a hop showing loss in these tools does not always mean real user traffic experiences the same loss; corroborate with a protocol-appropriate test (like iperf, curl, or a TCP-mode trace against the real port) when precision matters. A related and easy-to-miss mistake is reading a fully starred-out (blocked-looking) UDP or ICMP traceroute as proof the destination itself is unreachable, when it may only mean those specific probe types are filtered; a TCP-based trace against the actual service port is often what resolves that ambiguity.
A host can reach other systems but cannot reach its default gateway. Describe how to use ARP and interface commands (for example 'ip link', 'ip addr', 'ip neigh' or 'arp -n') to determine whether the issue is a Layer 2 problem (missing ARP entry, bad link) or a Layer 3 routing problem. Explain the outputs you expect in each case.
Sample Answer
Direct answer
When a host can reach other systems but not its default gateway, the fault is almost always local: either the host never learned the gateway's MAC address (a Layer 2 / ARP problem) or the routing table itself is wrong (a Layer 3 problem). The fastest way to tell these apart is to check the ARP cache before touching routing at all.
Structured elaboration
- Check the ARP/neighbor table first:
ip neigh show <gateway-ip>(orarp -non older systems). An entry in stateREACHABLEorSTALEwith a real MAC address means Layer 2 is fine and the problem is elsewhere (routing, or the gateway itself is down). An entry in stateFAILEDorINCOMPLETE, or no entry at all, means the host sent an ARP request and got no reply: a Layer 2 problem (bad cable/port, VLAN mismatch, the gateway's interface down, or a switch not forwarding broadcast/ARP traffic between the host and gateway). - Check the interface and link state:
ip link showshould showstate UPand a carrier. A down link explains missing ARP replies trivially, and rules out anything more subtle. - Check the routing table:
ip route showand confirm the host's subnet mask and default route point at the correct gateway IP, on the correct interface. A wrong subnet mask is a classic cause: the host computes the gateway as "on-link" when it is not (or vice versa), so it never even sends the ARP request that would establish reachability. - Force a fresh ARP resolution:
ip neigh flush <gateway-ip>followed by a ping will re-trigger ARP and let you watch the outcome cleanly, useful when a stale, wrong MAC is cached (for example after the gateway's NIC was replaced).
Worked example
Say ip route show reports default via 10.0.0.1 dev eth0 and ip addr show eth0 reports 10.0.0.55/25. The /25 mask covers 10.0.0.0-10.0.0.127, so .1 is correctly on-link and the host should ARP for it directly. Now ip neigh show 10.0.0.1 returns nothing, or FAILED. Because the route calculation is fine, this narrows the problem to Layer 2: check ip link show eth0 for carrier, and if that's up, suspect the switch port (wrong VLAN, port down, or a security policy dropping ARP) rather than anything on the host's IP configuration.
Trade-offs & pitfalls
The order matters: many engineers immediately start a packet capture, which is unnecessary overhead for what a single ip neigh command already answers. The most common mistake is assuming "can't reach the gateway" is a routing problem and jumping straight to route tables, missing that an incorrect subnet mask is itself a routing misconfiguration that manifests as an ARP failure. On a switch, also confirm the gateway's own interface and VLAN membership; a host-side fix cannot repair a gateway that has silently gone down or been moved to the wrong VLAN.
Unlock Full Question Bank
Get access to all 41 Network Troubleshooting and Diagnostics interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.