Networking Fundamentals and Protocols Questions
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
You are given a capture showing fragmented IPv4 packets and an ICMP Type 3 Code 4 (Fragmentation Needed) message reporting a next-hop MTU of 1400 bytes. Explain the role of the Don't Fragment (DF) bit and how Path MTU Discovery is supposed to behave here, then describe why PMTUD commonly fails in production (hint: something in the path is dropping the ICMP message) and what fixes actually resolve it for both TCP and UDP traffic.
Sample Answer
Direct answer
An ICMP Type 3 Code 4 message ("Fragmentation Needed and Don't Fragment was Set") is a router along the path telling the sender its packet was too large for the next hop's MTU (Maximum Transmission Unit, here 1400 bytes) and, because the Don't Fragment (DF) bit was set, the router dropped it rather than fragmenting it, expecting the sender to resend at a smaller size. Path MTU Discovery commonly fails in production because something along the path (often a firewall with an overly broad "block all ICMP" rule) discards that very ICMP message before it reaches the sender, so the sender never learns to shrink its packets and its large packets just keep silently disappearing.
Structured elaboration
The DF bit tells every router along the path "do not fragment this packet under any circumstances, if it doesn't fit, drop it and tell me why." Path MTU Discovery relies entirely on that "tell me why" part actually reaching the sender: the sender starts by assuming its LOCAL interface's MTU is usable end-to-end, sends with DF set, and if a router along the path can't forward it at that size, the router sends back exactly this ICMP message, reporting the smaller MTU it needs (1400 bytes here). The sender is then supposed to shrink its packet size to that reported value and retry.
The reason PMTUD commonly fails in production: many firewalls and security appliances, misconfigured to block ALL ICMP as a blanket "security" measure, silently discard the "Fragmentation Needed" message on its way back to the sender. The sender then never learns it needs to shrink its packets, keeps sending at the original (too-large) size with DF still set, and those packets keep getting silently dropped at the same router, forever, with no error ever surfacing to the sender, the classic "large transfers hang, small transfers succeed" symptom (small packets happen to fit under the constrained MTU and sail through fine, while anything larger vanishes without explanation).
Worked example
To reconstruct what happened from the capture: the fragmented IPv4 packets observed likely represent an EARLIER part of the same flow that happened to still get through (perhaps fragmented by an intermediate device before DF took full effect, or from a portion of traffic that didn't have DF set), while the ICMP message with next-hop MTU 1400 is the router's report on a LATER, DF-set packet it could not forward. To fix this for BOTH TCP and UDP traffic: for TCP, the most common resilient fix is MSS clamping on a network device at the edge (rewriting the TCP Maximum Segment Size option in SYN packets passing through, so TCP negotiates a small-enough segment size up front and the oversized-packet problem never occurs at all, PMTUD independent); for UDP, since there's no equivalent MSS negotiation, the application itself must either send appropriately small datagrams from the start or correctly handle PMTUD feedback (which requires NOT blocking the relevant ICMP messages on the path, the actual root-cause fix). In both cases, the truly correct long-term fix is ensuring ICMP "Fragmentation Needed"/"Packet Too Big" messages are explicitly PERMITTED through every firewall along the path, rather than working around their absence.
Trade-offs & pitfalls
MSS clamping is a pragmatic, widely-used workaround specifically because it doesn't depend on ICMP getting through at all, but it only helps TCP; it does nothing for UDP traffic hitting the exact same oversized-packet problem, which is why "block all ICMP" as a firewall policy is a genuinely bad default rather than a harmless-looking hardening step, it breaks a real, load-bearing part of how IP networking is supposed to self-correct.
Explain how IPv6 fragmentation differs from IPv4: why IPv6 moved fragmentation responsibility entirely to the source host (routers along the path no longer fragment), the role of the IPv6 fragmentation extension header, and what this design implies for Path MTU Discovery and for devices in the middle of the path.
Sample Answer
Direct answer
IPv6 removed the ability for ROUTERS along the path to fragment a packet; only the ORIGINATING host may fragment an IPv6 packet, using a dedicated fragmentation extension header, which pushes the entire responsibility for correctly sizing packets (and reacting to Path MTU, Maximum Transmission Unit, Discovery feedback) onto the sender rather than sharing it with the network.
Structured elaboration
In IPv4, any router along the path could fragment a packet too large for its outgoing link (unless the Don't Fragment bit was set), a convenience that came at real cost: routers doing fragmentation work on the fly, and receivers having to reassemble fragments that could have been created by ANY device along the path, not just the sender. IPv6's designers removed this entirely: routers now MUST drop an IPv6 packet that's too large for the next link and send back an ICMPv6 "Packet Too Big" message reporting the link's MTU, rather than fragmenting it themselves. If the ORIGINATING host still needs to send data larger than the smallest MTU along the path, it must fragment the packet itself, before sending, using the IPv6 Fragment extension header (which carries the same conceptual information as IPv4's Identification/Flags/Offset fields, just in a dedicated extension header rather than baked into the main IP header).
Worked example
This design change makes Path MTU Discovery load-bearing in IPv6 in a way it wasn't strictly required to be in IPv4: since no device in the middle of the path can rescue an oversized packet by fragmenting it, a sender that doesn't properly discover and respect the path's minimum MTU will simply have its packets dropped along the way (with an ICMPv6 message reporting why, PROVIDED that message isn't itself being filtered somewhere, the classic PMTUD-blackhole failure mode). This is also why a middlebox in the middle of an IPv6 path is a materially different kind of obstacle than in IPv4: it cannot silently paper over an oversized-packet problem by fragmenting on the sender's behalf; it can only drop and (hopefully) report back.
Trade-offs & pitfalls
The IPv6 design trades away routers' ability to "fix" an oversized packet on the fly in exchange for a simpler, more predictable job for routers (never having to fragment) and a stronger incentive for senders to get Path MTU Discovery right in the first place, rather than relying on the network to quietly compensate for an oversized packet. The practical implication: if ICMPv6 "Packet Too Big" messages are filtered anywhere along an IPv6 path (the same class of misconfiguration that breaks IPv4 PMTUD), there is NO fallback fragmentation happening in the middle of the path to mask the problem, the failure is more absolute than the equivalent IPv4 case.
Explain the seven layers of the OSI model. For each layer, state its primary responsibility, the name of its protocol data unit (PDU), and one or two protocols or technologies that commonly operate there. Then explain why identifying which layer a failure sits at is useful before you jump to a fix.
Sample Answer
Direct answer
The OSI model splits network communication into seven layers, each handing off a well-defined unit of work to the layer above and below it: Physical, Data Link, Network, Transport, Session, Presentation, and Application. Knowing which layer a protocol or symptom belongs to lets you reason about failures systematically instead of guessing.
Structured elaboration
| Layer | Responsibility | PDU (protocol data unit) | Common protocols/tech |
|---|---|---|---|
| 7. Application | Provides the interface applications use to talk over the network | Data | HTTP, DNS |
| 6. Presentation | Translates, encrypts, and compresses data into a form the application layer can use | Data | TLS, character encoding |
| 5. Session | Establishes, manages, and tears down a logical session between two hosts | Data | RPC session handling |
| 4. Transport | End-to-end delivery between processes: reliability, ordering, flow control | Segment (TCP) / Datagram (UDP) | TCP, UDP |
| 3. Network | Logical addressing and routing across networks | Packet | IP, ICMP |
| 2. Data Link | Framing and addressing on a single link (same broadcast domain) | Frame | Ethernet, ARP |
| 1. Physical | Raw bit transmission over a physical medium | Bit | Ethernet PHY, fiber, radio (Wi-Fi) |
The mnemonic that matters more than memorizing names is the DIRECTION of responsibility: each layer only needs to trust the layer directly below it to deliver its unit of data, and it only exposes a clean interface to the layer above. That's what lets a Transport-layer protocol like TCP work identically over Ethernet, Wi-Fi, or a VPN tunnel: it never needs to know which Layer 1/2 technology is underneath.
Worked example
Say a user reports "the site is down." Layer-by-layer reasoning turns that vague complaint into a specific hypothesis:
- Physical/Data Link symptom: the NIC shows no link light, or
ip linkreports the interface asDOWNa cable or switch-port problem. - Network layer symptom:
pingto the server's IP times out but the local gateway responds a routing problem somewhere between here and there. - Transport layer symptom:
pingsucceeds but a TCP connection to the port hangs or resets a firewall, a service that isn't listening, or a transport-level issue. - Application layer symptom: the TCP connection completes but the HTTP response is an error or garbage the service is up but misbehaving.
Each of these is a different team, a different fix, and a different urgency. That's the actual payoff of the model: it turns "the site is down" into "which layer, which evidence."
Trade-offs & pitfalls
The OSI model is a teaching and troubleshooting framework, not how real stacks are literally implemented. Session and Presentation are rarely separate pieces of code in modern systems; TLS, for instance, is commonly described as "sits between Transport and Application" rather than cleanly as Layer 6. Don't over-fit a real symptom to exactly one layer: a firewall dropping SYN packets looks like a Transport-layer symptom (connection never establishes) but the actual cause and fix live at a security-policy layer that OSI doesn't model at all.
After a TCP connection closes, the socket that initiated the close sits in TIME_WAIT for a period before the port is reusable. Explain why TIME_WAIT exists, what a half-open connection is, and how you would detect an unusually large number of sockets stuck in TIME_WAIT on a busy server. What are the trade-offs of the common mitigations for socket exhaustion caused by this?
Sample Answer
Direct answer
TIME_WAIT is the state the side that sent the FINAL ACK of a connection close sits in for a fixed period (commonly twice the maximum expected segment lifetime, often around 60 seconds on Linux) before the connection's resources are fully released. It exists so a delayed, duplicate packet from an old connection can't be mistaken for part of a brand-new connection reusing the same address/port pair. A half-open connection is one where only one side still believes the connection is alive; the other side has already reset, crashed, or otherwise abandoned it without a clean FIN exchange.
Structured elaboration
TIME_WAIT exists to protect two things: (1) it guarantees the final ACK the closing side sent actually gets through, by giving time to retransmit it if the peer's FIN gets retransmitted (meaning the ACK was lost); (2) it prevents a stray, delayed packet from a previous incarnation of a connection (same 4-tuple: source IP, source port, destination IP, destination port) from being delivered into a brand new connection that happens to reuse the same 4-tuple before the old segments have had time to disappear from the network.
To detect a large number of sockets stuck in TIME_WAIT on Linux, ss -tan state time-wait | wc -l (or the older netstat -ant | grep TIME_WAIT | wc -l) gives a live count; watching this metric over time distinguishes a normal, self-draining backlog from a genuine problem.
Worked example
A server that closes millions of short-lived outbound connections per hour (for instance, a service making one HTTP call per request to an upstream) can exhaust its available ephemeral source ports if TIME_WAIT sockets accumulate faster than they expire, because each TIME_WAIT socket still holds its port reserved. The two standard mitigations are: raise the number of available client-side (ephemeral) ports and/or reuse connections via keepalive/connection pooling so fewer connections churn through TIME_WAIT in the first place; and, on the SERVER side specifically, enabling SO_REUSEADDR and, where safe, tcp_tw_reuse lets a new outgoing connection reuse a TIME_WAIT 4-tuple once TCP timestamps confirm it's safe to do so, rather than waiting out the full timer.
Trade-offs & pitfalls
Disabling or drastically shortening TIME_WAIT globally (rather than tuning port ranges or reuse settings) is the wrong fix: it reintroduces the exact correctness problem TIME_WAIT was designed to prevent, stray old packets landing in a new connection and corrupting it. The safe levers are reducing HOW MANY connections churn through the state (pooling, keepalive) and widening the ephemeral port range, not shrinking the safety window itself.
Explain the differences between TCP and UDP in terms of connection model, reliability, ordering, and flow/congestion control. For each protocol, name two real-world services that should use it and explain why. Then describe a scenario where you would build a custom reliable protocol on top of UDP rather than simply using TCP.
Sample Answer
Direct answer
TCP is connection-oriented and guarantees reliable, in-order delivery with built-in flow and congestion control, at the cost of handshake setup latency and head-of-line blocking; UDP is connectionless, with no delivery guarantees, ordering, or congestion control, trading reliability for minimal overhead and lower latency. The right choice depends on whether the application can tolerate loss and reordering itself, or needs the transport layer to handle it.
Structured elaboration
| Property | TCP | UDP |
|---|---|---|
| Connection model | Connection-oriented (handshake required) | Connectionless (no setup) |
| Reliability | Guaranteed delivery via retransmission | Best-effort, no retransmission |
| Ordering | In-order delivery guaranteed | No ordering guarantee |
| Flow control | Yes (receive window) | None |
| Congestion control | Yes (built into the protocol) | None (must be built by the application, if needed at all) |
| Overhead | Higher (handshake, ACKs, header size) | Lower (no handshake, smaller header) |
Two examples per protocol: TCP is the right choice for a database connection or a file transfer, where losing or reordering even one byte silently would corrupt the result, and the application has no interest in reimplementing reliability itself. UDP is the right choice for live video/voice calls or DNS queries, where a single lost or late packet is better DISCARDED and moved past (an old, out-of-order audio frame is useless once its playback moment has passed) than retransmitted at the cost of added latency that would make the whole stream feel laggy.
Worked example
A scenario where building custom reliability ON TOP of UDP beats plain TCP: a real-time multiplayer game sending frequent position updates. TCP's strict in-order delivery means a single lost packet blocks EVERY later packet from being delivered to the application until the lost one is retransmitted and received (head-of-line blocking), even if those later packets contain fresher, more relevant position data. Building a thin reliability layer over UDP lets the application decide per-message whether it's worth retransmitting (a player's current position, three updates old, usually isn't worth retransmitting, a newer update has probably already superseded it) rather than being forced into strict in-order delivery for data where "newest wins" matters more than "nothing lost."
Trade-offs & pitfalls
A common mistake is treating "TCP is reliable, UDP isn't" as the END of the analysis; the real question is whether your application's OWN definition of correctness matches TCP's specific guarantees (strict ordering, full reliability) or would be better served by a custom scheme that's more permissive in exactly the ways TCP is rigid. Building your own reliability on UDP is real engineering work (implementing retransmission, sequencing, and congestion awareness yourself), not a shortcut, it's justified specifically when TCP's guarantees don't match what the application actually needs.
Unlock Full Question Bank
Get access to all 26 Networking Fundamentals and Protocols interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.