Networking Fundamentals and Protocols Questions
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
Walk through how TCP congestion control evolves during a long-lived connection: slow start, congestion avoidance, fast retransmit, and fast recovery. State which sender-side variable changes at each stage and what event triggers the transition to the next stage.
Sample Answer
Direct answer
Over the life of a connection, TCP's congestion window grows exponentially in slow start, switches to growing linearly in congestion avoidance once it approaches a known safe ceiling, and reacts to loss with fast retransmit and fast recovery rather than always restarting from scratch.
Structured elaboration
- Slow start: the connection begins with a small congestion window (historically 1 segment; modern stacks start higher, commonly around 10 segments per RFC 6928) and roughly DOUBLES the window every round trip, since each of the ACKs for the previous batch triggers sending two new segments. This continues until either loss occurs, or the window reaches a threshold called
ssthresh(slow start threshold), at which point the sender switches strategies. - Congestion avoidance: once at or above
ssthresh, growth switches from exponential to roughly linear (classically, additive increase of about one segment per round trip), a much more cautious probe for additional capacity. - Fast retransmit: if the sender sees three duplicate ACKs (the receiver repeatedly acknowledging the same byte, implying a specific segment is missing but LATER data did arrive), it retransmits the missing segment immediately, without waiting for the retransmission timer to expire, since three duplicate ACKs is strong, specific evidence of loss rather than simple reordering.
- Fast recovery: after a fast retransmit, rather than collapsing all the way back to slow start,
ssthreshis set to about half the current window, and the window itself is set near that halved value, so the sender doesn't have to re-earn all its previous progress from a window of one segment; it resumes near where it estimates the path can actually sustain.
Worked example
Picture a connection whose window has grown to 64 segments in flight when a single segment is lost and detected via three duplicate ACKs (not a full timeout). Fast retransmit resends the missing segment immediately. Fast recovery sets ssthresh to roughly 32 (half of 64) and the window to near that value, then resumes congestion avoidance's linear growth from there, rather than collapsing to slow start's small initial window and doubling all the way back up. Contrast this with a RETRANSMISSION TIMEOUT (no duplicate ACKs arrived at all, meaning the loss was severe enough that the whole flight of data went missing): that's a much stronger loss signal, and the sender resets ssthresh to half the current window but drops the actual window all the way back to slow start's minimum, since a timeout implies the path may be far more broken than a few duplicate ACKs would suggest.
Trade-offs & pitfalls
It's a common mistake to say TCP always halves its window on any loss and moves on; a full retransmission timeout is treated far more conservatively (full reset to slow start) than a fast-retransmit-detected loss (a much gentler recovery), because the ABSENCE of any duplicate ACKs at all is itself informative: it suggests either a much larger loss event or a badly congested/broken path, not just one unlucky dropped segment.
A packet capture of a failed connection attempt shows: the client sends SYN, the server responds SYN-ACK, the client retransmits SYN several times but never sends the final ACK, and the server's SYN-ACK is retransmitted once before the connection times out. List the plausible root causes for the client never completing the handshake (consider both network-level and host-level causes), and explain what evidence on the client versus the server would distinguish them.
Sample Answer
Direct answer
When a client keeps retransmitting its SYN and never sends the final ACK (while the server's SYN-ACK is only retransmitted once before the connection times out), the most likely causes are that the client's connect timeout hasn't fired yet, that something between the two hosts is dropping the ACK specifically (an asymmetric path or a stateful device confused about direction), or that the client-side application itself never actually attempted the ACK due to a bug. A different failure shape, the server sends SYN-ACK and the client immediately sends RST, points to a different family of causes entirely: the client rejecting the connection outright.
Structured elaboration
For the "client retransmits SYN, no final ACK" pattern, work through causes in order of likelihood:
- Asymmetric routing dropping only the return-to-forward-direction ACK path. If the SYN-ACK reaches the client (we know it does, since the client keeps retransmitting new SYNs rather than giving up, meaning it IS getting a response of some kind) but the client's ACK can't get back to the server along some other path, a stateful firewall or NAT device on that asymmetric path may be dropping the ACK because it doesn't recognize the connection's state in that direction.
- A middlebox is rewriting or dropping specific TCP options in the SYN-ACK that the client's stack doesn't handle gracefully, causing it to silently discard the SYN-ACK and retry instead of ACKing it.
- Client-side firewall or security policy is specifically blocking outbound ACKs to that destination while permitting the outbound SYNs, an unusual but real misconfiguration.
- MTU (Maximum Transmission Unit)-related silent packet loss on the SYN-ACK's return path if it happens to be an unusually large segment (rare for a SYN-ACK specifically, since it typically carries little payload, but worth ruling out if other symptoms point that way).
For the different shape (SYN-ACK followed by an IMMEDIATE client RST): this usually means the client-side application decided, upon establishing the connection, that it doesn't actually want it, for example an application-level timeout that already expired while the handshake was in flight, a client-side connection pool that raced two connection attempts and is aborting the loser, or a security tool on the client actively resetting connections that don't match an expected certificate or policy.
Worked example
To distinguish these hypotheses in practice, compare timestamps and evidence on BOTH ends: if the server's capture shows the SYN-ACK leaving on time but the client's capture never shows it arriving, the problem is in the path (asymmetric routing, a device eating it). If the client's capture shows the SYN-ACK arriving cleanly but no ACK is ever generated by the client's own stack, the bug is on the client host itself (application logic, local firewall) rather than the network path.
Trade-offs & pitfalls
A common mistake is assuming a stuck handshake is always a network problem; a client-side timeout race (the application gives up right as the handshake completes) produces an outwardly identical-looking symptom to a network drop and is only distinguishable by comparing what each side's own capture actually shows, not by reasoning about the network path alone.
Describe IPv4 fragmentation and reassembly: how the Identification, Flags, and Fragment Offset header fields work together, and how a receiver reassembles fragments back into the original packet. What failure modes (missing fragment, overlapping fragments, reassembly timeout) can occur, and why might you want to avoid fragmentation happening at all?
Sample Answer
Direct answer
IPv4 fragmentation splits a packet too large for a link into smaller pieces, each carrying enough header information (Identification, Flags, Fragment Offset) for the destination to reassemble the original packet, and it can fail in several distinct ways, a missing fragment, overlapping fragments, or a reassembly timeout, any of which prevents the original packet from ever being reconstructed.
Structured elaboration
Three IPv4 header fields work together to make fragmentation and reassembly possible:
- Identification: a 16-bit value the ORIGINAL packet is stamped with; every fragment of that same original packet carries the SAME Identification value, so the receiver knows which fragments belong together.
- Flags: includes the "Don't Fragment" (DF) bit, which tells routers along the path NOT to fragment this packet (instead, if it's too large for the next link, they must drop it and report the problem back via Internet Control Message Protocol, ICMP), and the "More Fragments" (MF) bit, set on every fragment except the LAST one, telling the receiver more pieces are still coming.
- Fragment Offset: a 13-bit field (in units of 8 bytes) indicating WHERE in the original packet's payload this particular fragment's data belongs, letting the receiver reassemble fragments that might arrive out of order.
The receiver buffers incoming fragments sharing the same Identification (and same source/destination address pair and protocol) until either all fragments have arrived (the last one identified by MF=0) and can be stitched back together in Fragment-Offset order, or a reassembly timer expires first.
Worked example
If a fragment carrying offset 0-1000 and another carrying offset 1000-2000 arrive, but the fragment that should have covered 2000-2500 (the actual final piece, with MF=0) never arrives, perhaps dropped somewhere along the path, the receiver has no way to know reassembly is even complete; it will hold the two partial fragments in its reassembly buffer until its reassembly timeout (commonly around 30 to 60 seconds depending on the OS) expires, then discard them and the entire original packet is lost, even though 2 of its 3 pieces successfully arrived. Overlapping fragments (two fragments both claiming to cover the SAME offset range, but with different content) are treated with suspicion by modern stacks specifically because that pattern was historically exploited in fragmentation-based attacks (evading firewall inspection by presenting different content to the firewall's re-assembly logic than to the actual destination host's).
Trade-offs & pitfalls
Fragmentation is something you generally want to AVOID rather than rely on: it multiplies the chance of a single original packet failing to arrive (since ALL its fragments must arrive for it to be usable at all, a single lost fragment dooms the whole packet even if every other fragment made it), it's more CPU-expensive for receivers to reassemble than to simply process one appropriately-sized packet, and many network devices (especially older or security-focused ones) handle fragments inconsistently or drop them outright, treating fragmentation itself as suspicious traffic. This is exactly why Path MTU (Maximum Transmission Unit) Discovery exists: to let a sender learn the right packet size UP FRONT and avoid needing fragmentation at all.
After a TCP connection closes, the socket that initiated the close sits in TIME_WAIT for a period before the port is reusable. Explain why TIME_WAIT exists, what a half-open connection is, and how you would detect an unusually large number of sockets stuck in TIME_WAIT on a busy server. What are the trade-offs of the common mitigations for socket exhaustion caused by this?
Sample Answer
Direct answer
TIME_WAIT is the state the side that sent the FINAL ACK of a connection close sits in for a fixed period (commonly twice the maximum expected segment lifetime, often around 60 seconds on Linux) before the connection's resources are fully released. It exists so a delayed, duplicate packet from an old connection can't be mistaken for part of a brand-new connection reusing the same address/port pair. A half-open connection is one where only one side still believes the connection is alive; the other side has already reset, crashed, or otherwise abandoned it without a clean FIN exchange.
Structured elaboration
TIME_WAIT exists to protect two things: (1) it guarantees the final ACK the closing side sent actually gets through, by giving time to retransmit it if the peer's FIN gets retransmitted (meaning the ACK was lost); (2) it prevents a stray, delayed packet from a previous incarnation of a connection (same 4-tuple: source IP, source port, destination IP, destination port) from being delivered into a brand new connection that happens to reuse the same 4-tuple before the old segments have had time to disappear from the network.
To detect a large number of sockets stuck in TIME_WAIT on Linux, ss -tan state time-wait | wc -l (or the older netstat -ant | grep TIME_WAIT | wc -l) gives a live count; watching this metric over time distinguishes a normal, self-draining backlog from a genuine problem.
Worked example
A server that closes millions of short-lived outbound connections per hour (for instance, a service making one HTTP call per request to an upstream) can exhaust its available ephemeral source ports if TIME_WAIT sockets accumulate faster than they expire, because each TIME_WAIT socket still holds its port reserved. The two standard mitigations are: raise the number of available client-side (ephemeral) ports and/or reuse connections via keepalive/connection pooling so fewer connections churn through TIME_WAIT in the first place; and, on the SERVER side specifically, enabling SO_REUSEADDR and, where safe, tcp_tw_reuse lets a new outgoing connection reuse a TIME_WAIT 4-tuple once TCP timestamps confirm it's safe to do so, rather than waiting out the full timer.
Trade-offs & pitfalls
Disabling or drastically shortening TIME_WAIT globally (rather than tuning port ranges or reuse settings) is the wrong fix: it reintroduces the exact correctness problem TIME_WAIT was designed to prevent, stray old packets landing in a new connection and corrupting it. The safe levers are reducing HOW MANY connections churn through the state (pooling, keepalive) and widening the ephemeral port range, not shrinking the safety window itself.
Explain the difference between TCP flow control and TCP congestion control: what each protects against, and one concrete mechanism each uses (the receive window versus the congestion window / slow start). Describe a real scenario where confusing the two would lead you to apply the wrong fix.
Sample Answer
Direct answer
Flow control protects the RECEIVER from being overwhelmed by data it can't process fast enough; congestion control protects the NETWORK from being overwhelmed by more traffic than the path between sender and receiver can actually carry. They use different signals and different windows, and confusing them leads to fixing the wrong end of the problem.
Structured elaboration
Flow control is implemented via the receive window (often written rwnd): the receiver advertises, in every ACK, how many more bytes of buffer space it currently has available. If the receiver's application is slow to read data out of its socket buffer, the advertised window shrinks, telling the sender to slow down regardless of how healthy the network path is. This is purely an endpoint-to-endpoint negotiation; the network in between plays no role.
Congestion control is implemented via the congestion window (cwnd), a value the SENDER maintains based on inferred network conditions, growing during slow start and congestion avoidance, and shrinking sharply when loss (or, with explicit congestion notification, an ECN mark) signals that the path is overloaded. Crucially, the sender is never allowed to have more unacknowledged data in flight than the SMALLER of rwnd and cwnd, whichever is more restrictive wins.
Worked example
Consider a video-conferencing server pushing data to a client on a fast, low-latency corporate network with no packet loss anywhere on the path. If the client's own CPU is pegged and its application isn't draining its TCP receive buffer fast enough, throughput will stall even though the network itself has plenty of headroom, that's a rwnd (flow control) problem, and the fix is on the client (free up CPU, drain the socket faster), not on the network. Conversely, if the same server saturates a congested shared WAN link, cwnd will shrink due to loss even though the receiver's buffer is empty and eager for more data, that's a congestion-control problem, and the fix is either reducing the amount of data in flight or addressing the actual network bottleneck, not the receiver.
Trade-offs & pitfalls
Confusing the two leads to genuinely wrong fixes: tuning net.ipv4.tcp_rmem (receive buffer size, which affects flow control's rwnd) will do nothing for a connection that's actually congestion-limited, and conversely, tuning the congestion control algorithm won't help a connection that's flow-control-limited by a slow receiving application. The fastest way to tell them apart in practice: watch which window value (rwnd vs cwnd) is the SMALLER, limiting one during the stall, most modern TCP instrumentation (ss -i on Linux) exposes both.
Unlock Full Question Bank
Get access to all 26 Networking Fundamentals and Protocols interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.