Network Troubleshooting and Diagnostics Questions
Systematically diagnosing network problems: a layered troubleshooting methodology, diagnostic tools and commands (ping, traceroute, tcpdump, packet capture), root-cause analysis, and connectivity and user-access issues. Covers isolating faults across the stack, reading packet-level evidence, and driving from symptom to root cause. The diagnostic discipline that spans all networking layers.
A production OSPF network shows repeated SPF recalculations causing transient reachability issues in area 1. Outline a troubleshooting approach with specific commands/telemetry to collect, likely root causes (LSA flaps, flapping links, unstable adjacencies), and mitigation steps (timers, interface dampening, BFD). Include how you would validate your fix without risking broader outages.
Sample Answer
Direct answer
Repeated SPF (Shortest Path First) recalculations causing transient reachability issues in one OSPF area point at something in that area generating frequent LSAs (Link-State Advertisements), each one forcing every router in the area to rerun SPF. Find the specific source of instability generating those LSAs using named commands and telemetry, rather than treating the symptom (frequent SPF runs) as itself the thing to fix.
Structured elaboration
- Confirm SPF recalculation frequency and correlate with LSA activity:
show ip ospf statisticsshows the SPF run history and the trigger reason for each run;show ip ospf databaseshows the current LSAs and their sequence numbers/age, and running it repeatedly (or diffing successive outputs) reveals which specific LSA is being regenerated frequently. - Check for flapping links specifically:
show interfaces(checking the "last input/output," "carrier transitions," and error counters) andshow logging | include %LINK|%LINEPROTOon the suspect interface reveal repeated up/down transitions directly. - Check for unstable adjacencies short of full flapping:
show ip ospf neighborrepeated over time, orshow logging | include OSPF-5-ADJCHG, reveals an adjacency repeatedly transitioning through intermediate states without ever fully going down. - Tune SPF-throttling timers as a mitigation, understanding its trade-off:
timers throttle spf <start-ms> <hold-ms> <max-ms>(for exampletimers throttle spf 200 1000 5000) delays how quickly SPF reruns after an LSA and backs off further if LSAs keep arriving, reducing the immediate CPU/instability impact; this is appropriate as a stopgap while investigating, not as a substitute for fixing the actual source. - Apply interface dampening to the specific flapping interface:
dampeningunder interface configuration (or the platform-equivalent, e.g.interface GigabitEthernet0/1thendampening) suppresses further LSA generation for a penalty period after repeated flaps on that one link, scoped to the identified source rather than the whole area. - Consider BFD (Bidirectional Forwarding Detection) as a complementary tool:
bfd interval 300 min_rx 300 multiplier 3under the interface, combined withip ospf bfdto bind OSPF to that BFD session, provides fast, deterministic failure detection so a genuinely failing link fails over cleanly once to a stable backup path rather than repeatedly flapping on a marginal primary one. - Validate any fix without risking a broader outage: apply dampening or timer changes to the specific identified flapping link/interface first, in a controlled way, and confirm via
show ip ospf statisticsthat SPF recalculation frequency actually drops as a direct result, rather than applying an area-wide timer change that could mask a real, future instability elsewhere.
Worked example
show ip ospf database run at short intervals shows one specific link's LSA (between two routers on the edge of area 1) incrementing its sequence number roughly every 90 seconds, correlated precisely with the reported SPF recalculation frequency shown in show ip ospf statistics. show interfaces on that link shows a nonzero, climbing "carrier transitions" counter, consistent with a marginal physical connection. Applying dampening to that specific interface immediately reduces SPF recalculation frequency across the whole area (confirmed again via show ip ospf statistics), while the physical link itself is scheduled for replacement as the durable fix; bfd interval 300 min_rx 300 multiplier 3 is also configured on the replacement link so any future failure is detected and failed over quickly rather than flapping.
Trade-offs & pitfalls
Tuning SPF timers or dampening broadly, across the whole area, without first identifying the specific flapping source via show ip ospf database and interface counters, treats the symptom everywhere rather than the cause at its actual location, and risks making the area's response to a genuinely new, real future instability slower than it should be. Identify the specific source first with the commands above, apply a targeted fix there, and reserve area-wide timer changes for cases where the instability is genuinely diffuse rather than traceable to one link.
You find a fiber/copper link between two switches reporting 'down' in a datacenter. Describe the step-by-step physical layer troubleshooting you would perform (LEDs, cable type and pinout, SFP/optic compatibility and transceiver diagnostics, polarity on duplex fiber, ethtool/ifconfig output on hosts, vendor 'show' commands on switches, and simple loopback tests). Include what a VFL or OTDR would show and when to escalate to cabling team.
Sample Answer
Direct answer
A physical link reporting "down" between two switches is troubleshot bottom-up, literally: confirm the cable and connector first, then the optic/transceiver, then the port configuration, escalating to specialized tools only once the simple checks are exhausted.
Structured elaboration
- Check the physical indicators first: link LEDs on both ends (a dark or amber LED where a solid or blinking green is expected is often diagnostic on its own), and confirm the cable is actually seated and undamaged; the cheapest checks come first.
- Confirm cable type and pinout match the link's expectations: a straight-through cable used where crossover is needed (on older equipment without auto-MDI/MDX), or a cable rated for a lower category than the link speed requires, can produce exactly this symptom.
- Check SFP/optic compatibility and diagnostics: many switches expose transceiver diagnostics (DOM/DDM, Digital Diagnostics Monitoring) reporting the optic's actual transmit/receive power levels; a receive power reading far below the optic's rated sensitivity threshold points at a dirty or degraded fiber connection, or a mismatched optic (wrong wavelength, wrong distance rating) for the fiber type in use.
- Check host-side and switch-side software state:
ethtoolorifconfigon a host, and vendor "show interface" commands on switches, to confirm the port isn't administratively disabled and to check reported speed/duplex; a duplex mismatch specifically can cause a link to come up but perform terribly, which is a related but distinct symptom from a link that won't come up at all. - Check polarity on duplex fiber: swapped transmit/receive fibers on a duplex connection is a classic, easy-to-overlook installation mistake that produces exactly a "link down" or "link flapping" symptom despite every other check looking fine.
- Use specialized tools when simple checks don't resolve it: a Visual Fault Locator (VFL, a visible laser) can reveal a physical break or bad connector in fiber by eye; an OTDR (Optical Time-Domain Reflectometer) can precisely locate a break, excessive bend, or connector loss along a longer fiber run that a simple loopback test can't pinpoint. A basic loopback test (looping a link's transmit back to its own receive, where supported) confirms whether the local port's transceiver and circuitry are functioning at all, independent of anything downstream.
- Escalate to the cabling team once you've confirmed the fault is genuinely physical (not port configuration) and beyond what DOM/DDM readings and a loopback test can resolve from the network side.
Worked example
Both switches show the port LED dark. DOM/DDM readings on one switch show the local optic transmitting at expected power, but receiving essentially nothing. A loopback test on that same port (transmit looped to receive) shows the local transceiver and switch circuitry functioning correctly, isolating the fault to the fiber path itself, not the equipment on either end. An OTDR run against that fiber run pinpoints a break at a specific distance, consistent with recent construction work reported near that cable path, and the finding is escalated to the cabling team with a precise location rather than "the fiber seems bad somewhere."
Trade-offs & pitfalls
Jumping straight to an OTDR or escalating to the cabling team before confirming the simple things (cable seated, correct optic, port not admin-down) wastes the cabling team's time on problems that were actually configuration; conversely, spending too long on network-side checks when DOM/DDM and loopback tests have already isolated a physical fault delays a fix that's outside your control anyway. Work bottom-up, but escalate promptly once you've genuinely isolated the fault to the physical layer.
Users report packet loss but interface counters on involved devices show no errors or drops. Describe advanced areas to investigate: per-queue egress drops/tail drops, microbursts leading to transient drops, QoS shaping/policing, bufferbloat and large buffers increasing latency, hardware offload masking counters, and how to gather high-resolution telemetry (ASIC counters, per-queue stats) to find the root cause.
Sample Answer
Direct answer
When packet loss is reported but interface counters on the involved devices show no errors, look above and below where standard counters measure: transient microbursts and per-queue tail drops that come and go faster than a counter's polling interval can capture, QoS shaping or policing discarding traffic by policy rather than by fault, and hardware-level buffering behavior (bufferbloat, offload features) that hides the real picture from a simple errors/drops counter.
Structured elaboration
- Understand what standard interface counters actually measure, and their blind spot: most polled counters (SNMP or show interface) sample at intervals of seconds; a microburst that fills a queue and causes a tail drop for a few milliseconds, then clears, can produce zero visible increment in a counter polled every 30 or 60 seconds, even though real packets were genuinely dropped.
- Look for per-queue, high-resolution telemetry instead: many modern switch ASICs expose per-queue drop counters (as opposed to aggregate interface-level counters) at a finer resolution; if available, these can reveal drops on a specific priority queue that never surface in the aggregate interface statistics.
- Check whether QoS shaping or policing is discarding traffic by design, not by fault: a policer enforces a committed rate and deliberately discards or remarks traffic above that rate at ingress, usually incrementing a policy-specific counter (a conform/exceed/violate counter on the policy-map) rather than the generic interface error/drop counter an engineer checks first; a shaper, by contrast, delays and queues excess traffic rather than dropping it outright, so it manifests as added latency and jitter rather than loss, unless its own buffer also overflows. A recent QoS policy change (a lowered committed rate, or a class reclassified into a stricter policer) is a common, entirely policy-driven cause of loss that will never show up as an interface error.
- Consider bufferbloat as a related but distinct pattern: an oversized buffer does not drop packets outright, but holds them long enough to inflate latency dramatically under load; this can look like loss to an application with a tight timeout (the packet was never actually dropped, but arrived too late to be useful), so distinguish true loss from excessive queuing delay using timestamps, not just counters.
- Check for hardware offload masking the real picture: some NICs and switch ASICs handle certain processing (checksums, some queueing decisions) in hardware in ways that are not reflected in the counters the OS or standard management interface exposes; a discrepancy between what the application experiences and what standard counters report can be a sign that the relevant activity is happening below where those counters look.
- Correlate timing precisely: gather the highest-resolution telemetry available (ASIC-level counters, per-queue stats, or policy-map conform/exceed/violate counters if accessible) and correlate the exact timestamps of reported application-level loss against any spike in queue depth, utilization, or policing activity at that same moment, even a spike too brief for a standard 30-second poll to register.
Worked example
An application reports occasional lost requests. Standard show interface counters on every device in the path show zero errors or drops over the reporting period. The policy-map attached to that egress interface, however, shows a nonzero and growing exceed counter under a QoS policer applied to this traffic class; a recent change lowered the committed rate for that class as part of a broader capacity reallocation. Enabling per-queue statistics on the relevant egress interface (polled every 1 second instead of every 60) corroborates this, showing brief spikes where the policed class's queue hits its maximum and experiences tail drops lasting under two seconds, precisely correlated with the timestamps of the application's reported failures; neither the standard 60-second interface counters nor a naive check of the interface's own drop counter would have surfaced this, since the drop is a deliberate policy action recorded in a QoS-specific counter.
Trade-offs & pitfalls
'The counters are clean' is often treated as proof there is no network-side loss, but standard interface counters have a real, specific blind spot for both short-duration events and policy-driven drops recorded elsewhere (in QoS policy-map counters, not the interface's own error/drop counters); before concluding the network is innocent, confirm you have looked at the highest time-resolution telemetry actually available on that hardware, and at any QoS policy applied to the affected traffic class, not just the generic interface counters. Distinguishing true drops (tail drop, policing) from bufferbloat-induced delay matters because the fixes are different: one calls for capacity, queue-management, or policy-rate changes, the other for buffer-sizing and queue-discipline tuning.
Explain the TCP three-way handshake and what you would look for in each packet to confirm that a client, load balancer, and backend are all progressing normally.
Sample Answer
Direct answer
The TCP three-way handshake is SYN, SYN-ACK, ACK: the client sends a SYN (synchronize) with its initial sequence number, the server replies with a SYN-ACK acknowledging that sequence number and providing its own, and the client completes the exchange with an ACK acknowledging the server's sequence number; only after all three arrive is the connection considered established, and confirming each leg in a capture tells you exactly how far a connection attempt actually got.
Structured elaboration
- SYN (client to server): look for the SYN flag set, an initial sequence number, and the client's advertised window size and options (like window scale and MSS, Maximum Segment Size); seeing this packet leave the client and never seeing a reply tells you the client sent correctly, so any failure is downstream of the client.
- SYN-ACK (server, or load balancer, to client): look for both SYN and ACK flags set, the acknowledgment number equal to the client's initial sequence number plus one, and the server's own initial sequence number; seeing this arrive confirms the SERVER (or whatever device generated it, which could be a load balancer terminating the connection on the server's behalf) received the SYN and is willing to proceed.
- ACK (client to server, completing the handshake): look for the ACK flag set (SYN now clear) with the acknowledgment number equal to the server's initial sequence number plus one; this is the client confirming receipt of the SYN-ACK, and its arrival at the server is what actually transitions the connection to ESTABLISHED on that side.
- How this confirms multiple hops are progressing correctly at once: capturing at the client, at a load balancer in the middle, and at the backend lets you see each leg of this three-part exchange independently at each point; if the client's SYN reaches the load balancer but the load balancer's own SYN to the backend never gets a SYN-ACK back, you've isolated the failure to specifically the load-balancer-to-backend leg, distinct from a client-to-load-balancer problem.
Worked example
A capture at the client shows a SYN sent with sequence number 1000. A capture at the load balancer shows that SYN arriving, and shows the load balancer generating its OWN new SYN toward the backend with a different sequence number (since a load balancer terminating TCP originates a fresh connection to the backend, rather than simply forwarding the same segments). If the backend-side capture never shows a SYN-ACK returning to the load balancer, the failure is isolated precisely to the load-balancer-to-backend leg, even though the client's own handshake with the load balancer completed successfully and the client may see no error at all yet, just an unusually long wait.
Trade-offs & pitfalls
A common mistake is treating "the client sees a successful connection" as proof the WHOLE path is healthy end to end; when a load balancer terminates TCP, it establishes an entirely separate handshake with the backend, and a failure on that separate leg is invisible to the client until the load balancer's own timeout or retry logic eventually surfaces an error. Always identify where TCP is actually being terminated versus just forwarded when reasoning about which "handshake" you're really looking at.
Given: RTT = 120 ms, MSS = 1460 bytes, TCP advertised window (rwnd) = 65,535 bytes, window scale negotiated = 4 (shift left 4). Calculate the theoretical maximum throughput in Mbps for a single TCP flow using the formula: throughput = (rwnd * 2^scale) / RTT. Explain how TCP window scaling, buffer sizes, and congestion control influence real achievable throughput and what tuning you'd apply for high-latency high-bandwidth WANs.
Sample Answer
Direct answer
Effective window = rwnd times 2 to the scale power = 65,535 times 16 = 1,048,560 bytes. Throughput = (effective window in bits) divided by RTT = (1,048,560 times 8) / 0.12 seconds = 69,904,000 bits per second, approximately 69.9 Mbps.
Structured elaboration
- Why window scaling matters here: TCP's original window field is only 16 bits, capping the advertised window at 65,535 bytes without scaling; a scale factor of 4 (a left-shift of 4 bits, meaning multiply by 2^4=16) extends the EFFECTIVE window far beyond that 16-bit limit, which is exactly why window scaling exists: modern high-bandwidth links would otherwise be capped at a small fraction of their real capacity by the original field's size alone.
- Why the calculation divides by RTT rather than some other time value: TCP's throughput ceiling, in the idealized case of no loss and full utilization of the window, is bounded by how much data can be "in flight" (unacknowledged) at once, which is the window; that data can be in flight for at most one round trip before an acknowledgment is needed to open the window further, so dividing the window's SIZE by the ROUND-TRIP TIME gives the maximum sustainable throughput under ideal conditions.
- What this number represents, and what it doesn't: 69.9 Mbps is a THEORETICAL CEILING based purely on window size and RTT; it assumes no packet loss, immediate acknowledgment, and the window fully utilized at all times, none of which real traffic guarantees; real achievable throughput on this same link, given actual congestion control behavior (which grows the window gradually and backs off on loss) and any real loss on the path, will typically be somewhat below this ceiling, sometimes substantially so if loss is nontrivial.
- What tuning would actually help on a real high-latency, high-bandwidth link: ensuring window scaling is negotiated at all (some older systems or restrictive middleboxes disable or strip it); increasing OS-level socket buffer sizes to allow the window to grow large enough to make use of scaling in the first place (a scaled window is only useful if the OS's own buffers are sized to support it); and, on paths with any real loss, tuning or selecting a congestion-control algorithm that recovers throughput faster after a loss event, since aggressive back-off on a high-bandwidth-delay-product link can otherwise leave substantial capacity unused for an extended recovery period.
Worked example
Verified by direct calculation: effective window = 65,535 × 2⁴ = 65,535 × 16 = 1,048,560 bytes. Converting to bits: 1,048,560 × 8 = 8,388,480 bits. Dividing by RTT in seconds (120 ms = 0.12 s): 8,388,480 / 0.12 = 69,904,000 bits per second = 69.904 Mbps. This was computed directly (not estimated) and matches the formula given exactly.
Trade-offs & pitfalls
This ceiling is a theoretical maximum assuming continuous, lossless utilization of the full scaled window; treating it as the throughput you should EXPECT to actually observe, rather than an upper bound, is a common misreading. On a real WAN path, actual achieved throughput depends heavily on whether the OS's socket buffers are actually sized to support this scaled window (a correctly negotiated scale factor is wasted if the buffer itself is too small to ever fill it) and on how much real loss the congestion-control algorithm has to react to.
Unlock Full Question Bank
Get access to all Network Troubleshooting and Diagnostics interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.