Networking Fundamentals and Protocols Questions
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
Explain the differences between TCP and UDP in terms of connection model, reliability, ordering, and flow/congestion control. For each protocol, name two real-world services that should use it and explain why. Then describe a scenario where you would build a custom reliable protocol on top of UDP rather than simply using TCP.
Sample Answer
Direct answer
TCP is connection-oriented and guarantees reliable, in-order delivery with built-in flow and congestion control, at the cost of handshake setup latency and head-of-line blocking; UDP is connectionless, with no delivery guarantees, ordering, or congestion control, trading reliability for minimal overhead and lower latency. The right choice depends on whether the application can tolerate loss and reordering itself, or needs the transport layer to handle it.
Structured elaboration
| Property | TCP | UDP |
|---|---|---|
| Connection model | Connection-oriented (handshake required) | Connectionless (no setup) |
| Reliability | Guaranteed delivery via retransmission | Best-effort, no retransmission |
| Ordering | In-order delivery guaranteed | No ordering guarantee |
| Flow control | Yes (receive window) | None |
| Congestion control | Yes (built into the protocol) | None (must be built by the application, if needed at all) |
| Overhead | Higher (handshake, ACKs, header size) | Lower (no handshake, smaller header) |
Two examples per protocol: TCP is the right choice for a database connection or a file transfer, where losing or reordering even one byte silently would corrupt the result, and the application has no interest in reimplementing reliability itself. UDP is the right choice for live video/voice calls or DNS queries, where a single lost or late packet is better DISCARDED and moved past (an old, out-of-order audio frame is useless once its playback moment has passed) than retransmitted at the cost of added latency that would make the whole stream feel laggy.
Worked example
A scenario where building custom reliability ON TOP of UDP beats plain TCP: a real-time multiplayer game sending frequent position updates. TCP's strict in-order delivery means a single lost packet blocks EVERY later packet from being delivered to the application until the lost one is retransmitted and received (head-of-line blocking), even if those later packets contain fresher, more relevant position data. Building a thin reliability layer over UDP lets the application decide per-message whether it's worth retransmitting (a player's current position, three updates old, usually isn't worth retransmitting, a newer update has probably already superseded it) rather than being forced into strict in-order delivery for data where "newest wins" matters more than "nothing lost."
Trade-offs & pitfalls
A common mistake is treating "TCP is reliable, UDP isn't" as the END of the analysis; the real question is whether your application's OWN definition of correctness matches TCP's specific guarantees (strict ordering, full reliability) or would be better served by a custom scheme that's more permissive in exactly the ways TCP is rigid. Building your own reliability on UDP is real engineering work (implementing retransmission, sequencing, and congestion awareness yourself), not a shortcut, it's justified specifically when TCP's guarantees don't match what the application actually needs.
Compare TCP congestion control algorithms Reno, NewReno, Cubic, and BBR at a conceptual level: how each reacts to packet loss or ECN, and their steady-state behavior on a high-bandwidth-delay-product cloud link versus the shared public internet. For a large file transfer across a satellite link (high RTT, low but non-zero loss), which would you prefer and why?
Sample Answer
Direct answer
Reno, NewReno, Cubic, and BBR represent an evolution in how TCP infers and reacts to congestion: the Reno family reacts to LOSS with a fixed halving of the window, Cubic grows more aggressively on high-bandwidth links using a cubic function of time since the last loss, and BBR abandons loss as the primary signal entirely, instead modeling the path's actual bandwidth and round-trip time directly.
Structured elaboration
- Reno: the classical algorithm. Slow start, congestion avoidance with linear (additive) growth, and on ANY loss, halves the window and re-enters a conservative recovery. Its big limitation on high-bandwidth-delay-product links is that halving the window after a single loss throws away a huge amount of earned capacity, and the subsequent linear regrowth takes a long time to recover it.
- NewReno: a refinement that fixes a specific weakness in Reno's fast recovery when MULTIPLE segments are lost within one window; Reno's original recovery logic could exit fast recovery prematurely and fall back to a slow, timeout-driven recovery for the second lost segment, while NewReno correctly stays in fast recovery until ALL the losses from that window are repaired.
- Cubic (the default on Linux for a long time): grows the window as a cubic function of the time elapsed since the last loss event, growing very slowly right after backing off, then accelerating, then leveling off as it approaches the window size where the last loss occurred, and probing gently past it. This makes Cubic much better at fully utilizing high-bandwidth, high-latency ("long fat") links than Reno's linear growth, since it isn't purely tied to round-trip-time-limited additive increase.
- BBR (Bottleneck Bandwidth and Round-trip propagation time): rather than reacting to loss at all, BBR periodically probes to directly estimate the bottleneck link's bandwidth and the path's minimum round-trip time, then paces its sending rate to match that estimate. This lets it largely ignore ordinary, non-congestive packet loss (which loss-based algorithms mistake for congestion), a real advantage on paths where a small amount of loss is normal and NOT actually a congestion signal (satellite links, some wireless links, or lossy long-haul fiber).
Worked example
For a large file transfer over a satellite link, characterized by very high round-trip time (often 500ms+) and some baseline non-congestive loss (a normal characteristic of the medium, not a sign of an overloaded path), BBR is generally the stronger choice: a loss-based algorithm like Cubic will repeatedly (and wrongly) interpret that baseline loss as congestion and needlessly shrink its window, capping throughput well below what the link can actually sustain, while BBR's bandwidth-and-RTT model isn't fooled by loss that isn't actually caused by queue buildup.
Trade-offs & pitfalls
BBR isn't a universal win: on a link SHARED with loss-based flows (Cubic, Reno), BBR's willingness to keep sending through non-congestive loss can let it grab a disproportionate share of a congested bottleneck's capacity from more conservative Reno/Cubic flows sharing that same link, an active area of real-world congestion-control fairness research, not a settled solved problem.
A public-facing TCP service is hit by a SYN flood: an attacker sends a rapid stream of SYNs (often spoofed) so the server allocates state for many half-open connections and runs out of resources for legitimate ones. Explain how SYN cookies let the server avoid this without changing the three-way handshake the legitimate client sees, and what the server gives up (in terms of TCP options) while cookies are active.
Sample Answer
Direct answer
A SYN flood works by sending a stream of SYN segments (often with spoofed source addresses) so the server allocates per-connection state for each one and replies with a SYN-ACK, but the final ACK never arrives, exhausting the server's backlog of half-open connections. SYN cookies let the server defer allocating any state at all until the final ACK actually shows up, so a flood of SYNs that never complete costs the server almost nothing.
Structured elaboration
Without SYN cookies, a normal server implementation stores an entry in its SYN backlog queue for every SYN it receives, holding onto that entry until the handshake completes or times out. A large flood of SYNs fills the backlog with half-open connections, so genuine clients' SYNs get silently dropped once the queue is full, that's the denial of service.
With SYN cookies enabled, once the backlog nears capacity, the server stops storing per-connection state up front. Instead, when it receives a SYN, it encodes the essential information it would normally have stored (effectively a hash of the source/destination IP, ports, and a secret, folded into the initial sequence number it sends back in the SYN-ACK) directly into its own sequence number field, and discards the SYN backlog entry immediately. If the final ACK ever arrives, the server checks that its acknowledgment number is the cookie value plus one, reconstructs the connection state on the spot, and only THEN commits real resources, at the exact moment a real, completing client shows up. If the SYN was part of the flood and no valid ACK ever comes, the server never allocated anything for it in the first place.
Worked example
From the legitimate client's point of view, the handshake looks completely normal: SYN, SYN-ACK, ACK, exactly as usual, it never knows cookies were involved. The difference is entirely server-side bookkeeping. The cost of the cookie scheme is that the server must give up storing some TCP options across the handshake (since there's no room left in a single 32-bit sequence number to also encode arbitrary options like the negotiated MSS (Maximum Segment Size) or window scale precisely), which is why some cookie implementations round MSS to one of a small handful of common values instead of preserving whatever exact value the client requested.
Trade-offs & pitfalls
SYN cookies alone do not stop a volumetric flood from consuming link bandwidth or CPU cycles processing the flood of SYNs in the first place, they only protect the connection-state table from being exhausted. A senior answer distinguishes "the handshake structure survives an attack that tries to exhaust connection state" (what SYN cookies solve) from "the network survives an attack that tries to exhaust bandwidth or CPU" (a separate problem requiring rate limiting, upstream filtering, or scrubbing, outside the scope of the handshake mechanism itself).
Explain bufferbloat: why excessive buffering in network devices increases latency and jitter under load even though it reduces packet loss, and how Active Queue Management algorithms such as fq_codel counteract it. Why does bufferbloat specifically interfere with TCP's own congestion signals?
Sample Answer
Direct answer
Bufferbloat happens when routers or switches along a path have excessively large buffers that queue packets during congestion instead of dropping them, which keeps loss low but lets queuing delay grow essentially unbounded, adding latency and jitter that TCP's own congestion control never gets a clear enough signal to react to. Active Queue Management algorithms like fq_codel counteract it by proactively dropping or marking packets BEFORE the queue grows large, restoring a timely loss signal.
Structured elaboration
A classic loss-based congestion-control algorithm relies on packet loss (or an explicit congestion mark) as its primary "back off" signal. If a device's buffer is very large, it can absorb a burst of excess traffic by queuing it rather than dropping it, so loss never actually happens, but every packet sitting in that oversized queue now experiences extra delay waiting its turn. Ironically, a device built to be MORE forgiving (bigger buffer, less loss) ends up making the user experience WORSE (much higher and more variable latency) precisely because it hides the congestion signal the sender needs to slow down.
fq_codel (Fair Queuing with Controlled Delay) attacks this from two angles: "fair queuing" gives each active flow its own small queue so one bulk-transfer flow can't monopolize the buffer and starve a latency-sensitive flow (like a video call) sharing the same link; "controlled delay" tracks how long packets are actually sitting in the queue, and once a packet has been queued longer than a target delay (commonly around 5ms), it starts dropping packets to force the responsible flow's congestion control to back off, well before the queue grows large enough to cause serious latency.
Worked example
On a home internet connection with a large, un-managed buffer at the router, a single large upload (like a cloud backup) can push queuing delay from a normal few milliseconds up to several hundred milliseconds or more, which is directly noticeable as a video call over the same connection becoming choppy or laggy, even though no packets for the video call are actually being dropped, they're just sitting in a queue behind the bulk upload's traffic for hundreds of milliseconds. Enabling fq_codel on that router's egress queue caps how long any packet can wait, keeping the video call's latency low even while the bulk upload continues in the background.
Trade-offs & pitfalls
Bufferbloat is easy to misdiagnose as "the link doesn't have enough bandwidth" when the real problem is excess, unmanaged queuing delay on a link that has PLENTY of bandwidth; the fix (AQM, not more bandwidth) is the opposite of what a naive read of "things feel slow" would suggest. A useful diagnostic: latency under load (with a saturating background transfer running) that's dramatically higher than idle-latency is the signature of bufferbloat, not a raw throughput problem.
Define bandwidth, throughput, and goodput, and explain at least four distinct reasons an application might observe lower throughput or goodput than the link's rated bandwidth. For a given slow transfer, what is the fastest way to tell which of your reasons is the actual cause?
Sample Answer
Direct answer
Bandwidth is the link's raw theoretical capacity; throughput is what's actually achieved end-to-end, which is always less than or equal to bandwidth; and goodput is throughput minus any overhead that isn't USEFUL application data (protocol headers, retransmissions of data that eventually succeeds anyway). An application commonly sees less than the rated bandwidth for several genuinely distinct reasons, and telling them apart is the fast path to the right fix.
Structured elaboration
Four common, distinct reasons throughput or goodput falls short of bandwidth:
- Retransmissions: any lost segment that has to be resent consumes bandwidth without contributing new useful data, so a lossy path shows lower goodput even at unchanged raw throughput.
- Protocol overhead: every layer's headers (Ethernet, IP, TCP) consume bytes that count toward raw throughput but never toward the application's actual USEFUL payload, goodput specifically excludes them.
- Congestion-window or flow-control limits: if the connection's congestion window or receive window is smaller than the path's bandwidth-delay product, the sender is idle waiting for ACKs rather than continuously filling the pipe, capping achieved throughput well below the link's rated bandwidth regardless of loss.
- Application-level inefficiency: an application that reads/writes in small chunks, adds its own serialization overhead, or simply isn't pipelining requests efficiently can bottleneck well below what the transport layer itself is capable of moving.
Worked example
The fastest way to localize which of these applies to a specific slow transfer: check retransmission counters first (ss -i's retrans field), if they're elevated, loss/retransmission is a real contributor. If retransmissions are low but the window (cwnd/rwnd from the same tool) is small relative to the path's bandwidth-delay product, the connection is window-limited rather than loss-limited, and no amount of "fixing loss" will help. If BOTH retransmissions are low and the window is comfortably larger than the bandwidth-delay product, yet throughput is still poor, suspect application-level inefficiency (measure CPU usage and how the application is actually issuing reads/writes) rather than anything at the transport layer at all.
Trade-offs & pitfalls
The most common diagnostic mistake is assuming a throughput shortfall is automatically a "network problem" and reaching straight for network-level tuning; a genuinely application-bound bottleneck (small, unbuffered writes; serialized rather than pipelined requests) is common and invisible to any amount of TCP-level tuning, wasting real effort until the actual bottleneck is correctly localized first.
Unlock Full Question Bank
Get access to all 27 Networking Fundamentals and Protocols interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.