Networking Fundamentals and Protocols Questions
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
An object-storage service needs to optimize TCP transfers for multi-gigabyte uploads and downloads between clients and storage nodes. What transport- and OS/NIC-level levers would you investigate to raise throughput, and how would you decide between using several parallel connections versus one well-tuned connection for a large transfer?
Sample Answer
Direct answer
For large object-storage transfers, the levers worth investigating, roughly in order of impact, are: making sure window scaling and socket buffers are large enough for the path's bandwidth-delay product, enabling NIC-level offloads so the CPU isn't the bottleneck at high throughput, and deciding whether to use several parallel connections or one well-tuned connection based on whether the limiting factor is per-connection window size or something else entirely (like a single flow being unfairly rate-limited by a middlebox).
Structured elaboration
First, confirm the connection's window (after scaling) and OS socket buffers are large enough to cover the path's bandwidth-delay product, undersized buffers here silently cap throughput regardless of how good everything else is (this is the same BDP-sizing exercise as tuning any other high-bandwidth, high-latency transfer). Second, check NIC-level segmentation offloads: TSO/GSO let the OS hand large chunks of data to the NIC and have the NIC itself split them into wire-sized frames, and LRO does the reverse on receive, coalescing many small incoming frames before handing them to the OS; without these, the CPU has to do that segmentation/coalescing work itself, and at multi-gigabit throughput that CPU cost can become the actual bottleneck well before the network link itself is saturated. Third, decide on Nagle's algorithm (which delays sending small writes to coalesce them into fewer, larger segments): for large sequential transfers, Nagle is rarely the bottleneck since writes are already large, but for anything issuing many small writes interleaved with reads, disabling it (TCP_NODELAY) avoids needless latency.
Worked example
The parallel-versus-single-connection decision comes down to WHAT is actually being limited. If a single connection's throughput is capped by its own maximum achievable window (even after correct BDP-based tuning, some paths or middleboxes limit an individual flow's effective window more aggressively than the path's own capacity), opening several parallel connections lets the AGGREGATE throughput exceed what one connection alone could reach, effectively working around a per-flow limit by using multiple flows. But parallel connections add real complexity: more connection-management overhead, more complexity assembling the transferred object back together correctly if it's split across connections, and, on a link SHARED with other traffic, an unfair grab of a disproportionate share of available bandwidth compared to a well-behaved single flow. If the true bottleneck is the raw link capacity itself (not a per-flow cap), splitting into parallel connections doesn't help, and one well-tuned single connection is simpler and just as fast.
Trade-offs & pitfalls
It's tempting to reach for "just add more parallel connections" as a default fix for slow transfers, but that only helps when the limiting factor is genuinely per-connection (a window/RTT/middlebox constraint on a single flow), and on a genuinely bandwidth-saturated shared link, more parallel connections from one client mostly just take a larger, less fair share of that link away from other traffic rather than achieving any real net throughput gain.
Describe the four layers of the TCP/IP (Internet) model, map each to its corresponding OSI layer(s), and give one example protocol per TCP/IP layer. Explain in one or two sentences why two competing layering models exist and are both still used in practice.
Sample Answer
Direct answer
The TCP/IP (Internet) model has four layers: Link, Internet, Transport, and Application. It maps onto the seven-layer OSI model by collapsing OSI's bottom two layers into "Link" and its top three into "Application."
Structured elaboration
| TCP/IP layer | Corresponding OSI layer(s) | Example protocol |
|---|---|---|
| Application | Application, Presentation, Session (7, 6, 5) | HTTP |
| Transport | Transport (4) | TCP |
| Internet | Network (3) | IP |
| Link | Data Link, Physical (2, 1) | Ethernet |
Both models describe the same reality; they just carve it up at different granularity. OSI was designed as a general, vendor-neutral reference standard, drawn up before the protocols that would actually win were finalized, so it separates concerns the real internet protocol suite never bothered to split (nothing on the modern internet implements a distinct Session or Presentation layer as its own protocol). TCP/IP is the model that grew directly out of the protocols that were actually built and deployed, so it only has as many layers as the real stack needs to describe.
Worked example
Take a single HTTPS request. In TCP/IP terms: Application layer is HTTPS/TLS+HTTP, Transport is TCP, Internet is IP, Link is Ethernet or Wi-Fi. In OSI terms, the same request spans Layer 7 (HTTP semantics), arguably Layer 6 (TLS's encryption/format role, though in practice TLS libraries sit logically between transport and application), Layer 4 (TCP), Layer 3 (IP), and Layers 2/1 (the local network technology). Neither model is "more correct"; TCP/IP is more useful for describing what's actually running, OSI is more useful for a shared vocabulary when discussing where a NEW piece of functionality should live.
Trade-offs & pitfalls
A common mistake is treating the OSI-to-TCP/IP mapping as one-to-one instead of many-to-one. Session and Presentation don't disappear in TCP/IP, they're just not modeled as distinct layers because no widely deployed protocol maps cleanly onto only one of them. When someone asks "what OSI layer does TLS operate at," the honest answer is "it doesn't map cleanly, and that's expected," not a forced single number.
TCP Fast Open (TFO) lets a client send data in the SYN packet, skipping a full round trip on repeat connections to the same server. Explain how a middlebox built to expect a standard three-way handshake can misinterpret or drop TFO traffic, and what fallback behavior a client implementation needs so a TFO attempt never makes a connection LESS reliable than a plain handshake would have been.
Sample Answer
Direct answer
TCP Fast Open (TFO) lets a client include application data directly in its SYN packet on a REPEAT connection to a server it has connected to before, skipping the usual "wait for the handshake to finish before sending anything" round trip; the risk is that some middleboxes, built assuming a SYN never carries a meaningful payload, may drop, strip, or misforward that data, or even the whole packet, since it doesn't look like a standard handshake to them.
Structured elaboration
TFO works by having the server issue a cryptographic cookie to the client the FIRST time they connect (during a normal handshake, no fast-open data yet). On any SUBSEQUENT connection attempt, the client includes that cookie plus its actual application data directly in the SYN packet. If the server validates the cookie, it can begin processing that data immediately, effectively saving a full round trip compared to the standard "handshake completes, THEN data starts flowing" sequence.
The middlebox risk comes from an assumption baked into a lot of older network equipment: that a SYN packet is control-plane-only and never carries a meaningful payload. Some middleboxes (older NAT devices, certain firewalls, some load balancers) built on that assumption may strip TCP options they don't recognize (breaking the cookie mechanism itself), drop SYN packets that carry payload as anomalous or suspicious, or in some documented cases, mishandle the connection entirely.
Worked example
The required fallback behavior: if a TFO attempt's SYN (carrying the cookie and data) gets no response, or gets a response that doesn't correctly acknowledge the fast-open data, a correct client implementation must FALL BACK to a standard handshake, retransmitting the SYN WITHOUT the fast-open data and cookie, and only sending the actual application data after the connection is fully, conventionally established. This fallback is what keeps TFO safe to enable broadly: in the worst case (a hostile middlebox on the path), the connection degrades to ordinary TCP behavior rather than failing outright, the client just loses the round-trip savings TFO was trying to provide, it never loses connectivity because of it.
Trade-offs & pitfalls
The whole design of TFO is built around the assumption that middlebox interference IS common enough to plan for, not a rare edge case, that's precisely why the fallback-to-standard-handshake behavior is a REQUIRED part of any compliant implementation, not an optional nicety. An implementation that assumes TFO will always succeed and doesn't implement the fallback path correctly risks connections silently failing specifically on paths that include one of these older devices, exactly the class of path where the round-trip savings would have mattered most (since those paths tend to be higher-latency, longer routes).
Explain bufferbloat: why excessive buffering in network devices increases latency and jitter under load even though it reduces packet loss, and how Active Queue Management algorithms such as fq_codel counteract it. Why does bufferbloat specifically interfere with TCP's own congestion signals?
Sample Answer
Direct answer
Bufferbloat happens when routers or switches along a path have excessively large buffers that queue packets during congestion instead of dropping them, which keeps loss low but lets queuing delay grow essentially unbounded, adding latency and jitter that TCP's own congestion control never gets a clear enough signal to react to. Active Queue Management algorithms like fq_codel counteract it by proactively dropping or marking packets BEFORE the queue grows large, restoring a timely loss signal.
Structured elaboration
A classic loss-based congestion-control algorithm relies on packet loss (or an explicit congestion mark) as its primary "back off" signal. If a device's buffer is very large, it can absorb a burst of excess traffic by queuing it rather than dropping it, so loss never actually happens, but every packet sitting in that oversized queue now experiences extra delay waiting its turn. Ironically, a device built to be MORE forgiving (bigger buffer, less loss) ends up making the user experience WORSE (much higher and more variable latency) precisely because it hides the congestion signal the sender needs to slow down.
fq_codel (Fair Queuing with Controlled Delay) attacks this from two angles: "fair queuing" gives each active flow its own small queue so one bulk-transfer flow can't monopolize the buffer and starve a latency-sensitive flow (like a video call) sharing the same link; "controlled delay" tracks how long packets are actually sitting in the queue, and once a packet has been queued longer than a target delay (commonly around 5ms), it starts dropping packets to force the responsible flow's congestion control to back off, well before the queue grows large enough to cause serious latency.
Worked example
On a home internet connection with a large, un-managed buffer at the router, a single large upload (like a cloud backup) can push queuing delay from a normal few milliseconds up to several hundred milliseconds or more, which is directly noticeable as a video call over the same connection becoming choppy or laggy, even though no packets for the video call are actually being dropped, they're just sitting in a queue behind the bulk upload's traffic for hundreds of milliseconds. Enabling fq_codel on that router's egress queue caps how long any packet can wait, keeping the video call's latency low even while the bulk upload continues in the background.
Trade-offs & pitfalls
Bufferbloat is easy to misdiagnose as "the link doesn't have enough bandwidth" when the real problem is excess, unmanaged queuing delay on a link that has PLENTY of bandwidth; the fix (AQM, not more bandwidth) is the opposite of what a naive read of "things feel slow" would suggest. A useful diagnostic: latency under load (with a saturating background transfer running) that's dramatically higher than idle-latency is the signature of bufferbloat, not a raw throughput problem.
Implement a user-space TCP handshake and retransmission simulator (Go or Python) that models the SYN / SYN-ACK / ACK exchange plus retransmission with exponential backoff on loss. The simulator should accept a configurable packet-loss rate and RTT distribution, and print a deterministic event timeline suitable for a unit test. Provide runnable code or complete pseudocode, and explain what your timeline shows about how backoff behaves as loss increases.
Sample Answer
Direct answer
A discrete-event simulator for the handshake models each of the three messages (SYN, SYN-ACK, ACK) as independently subject to loss, applies an exponential-backoff timer whenever the client doesn't hear back in time, and logs every event with a simulated timestamp, giving a deterministic, reproducible timeline for a given random seed.
Structured elaboration (approach)
The simulator advances a virtual clock rather than real wall-clock time: for each handshake attempt, it draws a one-way delay from the configured RTT distribution and independently decides (via a seeded random number generator, so results are reproducible) whether each message is delivered or lost, at the configured loss rate. If the full SYN/SYN-ACK/ACK sequence completes, the connection is marked ESTABLISHED. If anything is lost, the client's virtual timeout fires (starting at a base value and doubling on every subsequent retry, the exponential backoff), and it retransmits.
Worked example (code)
import random
class HandshakeSimulator:
def __init__(self, loss_rate, rtt_fn, base_timeout=1.0, max_retries=6, seed=0):
self.loss_rate = loss_rate
self.rtt_fn = rtt_fn # callable(rng) -> RTT in simulated ticks
self.base_timeout = base_timeout
self.max_retries = max_retries
self.rng = random.Random(seed) # seeded: reproducible timeline
self.events = []
self.time = 0.0
def _log(self, msg):
self.events.append((round(self.time, 3), msg))
def _delivered(self):
return self.rng.random() >= self.loss_rate
def run(self):
attempt, timeout = 0, self.base_timeout
while attempt < self.max_retries:
self._log(f"client sends SYN (attempt {attempt+1})")
one_way = self.rtt_fn(self.rng) / 2.0
if self._delivered():
self.time += one_way
self._log("server receives SYN, sends SYN-ACK")
one_way2 = self.rtt_fn(self.rng) / 2.0
if self._delivered():
self.time += one_way2
self._log("client receives SYN-ACK, sends ACK")
one_way3 = self.rtt_fn(self.rng) / 2.0
if self._delivered():
self.time += one_way3
self._log("server receives ACK, connection ESTABLISHED")
return True
self.time += timeout
self._log(f"client timeout after {timeout:.2f} ticks, retransmitting SYN")
timeout *= 2 # exponential backoff
attempt += 1
self._log("handshake failed after max retries")
return False
Executed with three configurations: (a) no loss, fixed 0.1-tick round trip completed instantly (t=0.15, ESTABLISHED). (b) 40% loss with jittery round trips exhausted all 6 retries before succeeding (in the actual run, the sequence never fully completed within 6 attempts, alternating SYN loss, SYN-ACK loss, and one run where the final ACK itself was lost after both earlier messages succeeded). (c) A deterministic reproducibility check confirmed identical event timelines across two runs with the same seed, and a dedicated 100%-loss run confirmed the backoff sequence doubles exactly as expected: 0.30, 0.60, 1.20, 2.40 ticks.
Trade-offs & pitfalls (edge cases and complexity)
Complexity is O(1) work per attempt, O(max_retries) total, trivial computationally; the real engineering value is in the event log's fidelity, not raw performance. A simplification worth stating honestly: this model retries the ENTIRE handshake from a fresh SYN on any failure, including a lost final ACK, whereas real TCP is more nuanced there (a lost final ACK is actually recovered by the SERVER retransmitting its SYN-ACK, not the client resending a fresh SYN); a higher-fidelity simulator would track which specific message needs retransmission rather than always restarting from SYN. Edge cases handled: loss of each of the three messages independently, exhausting max retries without ever establishing the connection, and backoff growing without bound (a production version would cap the backoff at some maximum rather than doubling forever).
Unlock Full Question Bank
Get access to all 34 Networking Fundamentals and Protocols interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.