Networking Fundamentals and Protocols Questions
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
Compare TCP congestion control algorithms Reno, NewReno, Cubic, and BBR at a conceptual level: how each reacts to packet loss or ECN, and their steady-state behavior on a high-bandwidth-delay-product cloud link versus the shared public internet. For a large file transfer across a satellite link (high RTT, low but non-zero loss), which would you prefer and why?
Sample Answer
Direct answer
Reno, NewReno, Cubic, and BBR represent an evolution in how TCP infers and reacts to congestion: the Reno family reacts to LOSS with a fixed halving of the window, Cubic grows more aggressively on high-bandwidth links using a cubic function of time since the last loss, and BBR abandons loss as the primary signal entirely, instead modeling the path's actual bandwidth and round-trip time directly.
Structured elaboration
- Reno: the classical algorithm. Slow start, congestion avoidance with linear (additive) growth, and on ANY loss, halves the window and re-enters a conservative recovery. Its big limitation on high-bandwidth-delay-product links is that halving the window after a single loss throws away a huge amount of earned capacity, and the subsequent linear regrowth takes a long time to recover it.
- NewReno: a refinement that fixes a specific weakness in Reno's fast recovery when MULTIPLE segments are lost within one window; Reno's original recovery logic could exit fast recovery prematurely and fall back to a slow, timeout-driven recovery for the second lost segment, while NewReno correctly stays in fast recovery until ALL the losses from that window are repaired.
- Cubic (the default on Linux for a long time): grows the window as a cubic function of the time elapsed since the last loss event, growing very slowly right after backing off, then accelerating, then leveling off as it approaches the window size where the last loss occurred, and probing gently past it. This makes Cubic much better at fully utilizing high-bandwidth, high-latency ("long fat") links than Reno's linear growth, since it isn't purely tied to round-trip-time-limited additive increase.
- BBR (Bottleneck Bandwidth and Round-trip propagation time): rather than reacting to loss at all, BBR periodically probes to directly estimate the bottleneck link's bandwidth and the path's minimum round-trip time, then paces its sending rate to match that estimate. This lets it largely ignore ordinary, non-congestive packet loss (which loss-based algorithms mistake for congestion), a real advantage on paths where a small amount of loss is normal and NOT actually a congestion signal (satellite links, some wireless links, or lossy long-haul fiber).
Worked example
For a large file transfer over a satellite link, characterized by very high round-trip time (often 500ms+) and some baseline non-congestive loss (a normal characteristic of the medium, not a sign of an overloaded path), BBR is generally the stronger choice: a loss-based algorithm like Cubic will repeatedly (and wrongly) interpret that baseline loss as congestion and needlessly shrink its window, capping throughput well below what the link can actually sustain, while BBR's bandwidth-and-RTT model isn't fooled by loss that isn't actually caused by queue buildup.
Trade-offs & pitfalls
BBR isn't a universal win: on a link SHARED with loss-based flows (Cubic, Reno), BBR's willingness to keep sending through non-congestive loss can let it grab a disproportionate share of a congested bottleneck's capacity from more conservative Reno/Cubic flows sharing that same link, an active area of real-world congestion-control fairness research, not a settled solved problem.
Walk through how TCP congestion control evolves during a long-lived connection: slow start, congestion avoidance, fast retransmit, and fast recovery. State which sender-side variable changes at each stage and what event triggers the transition to the next stage.
Sample Answer
Direct answer
Over the life of a connection, TCP's congestion window grows exponentially in slow start, switches to growing linearly in congestion avoidance once it approaches a known safe ceiling, and reacts to loss with fast retransmit and fast recovery rather than always restarting from scratch.
Structured elaboration
- Slow start: the connection begins with a small congestion window (historically 1 segment; modern stacks start higher, commonly around 10 segments per RFC 6928) and roughly DOUBLES the window every round trip, since each of the ACKs for the previous batch triggers sending two new segments. This continues until either loss occurs, or the window reaches a threshold called
ssthresh(slow start threshold), at which point the sender switches strategies. - Congestion avoidance: once at or above
ssthresh, growth switches from exponential to roughly linear (classically, additive increase of about one segment per round trip), a much more cautious probe for additional capacity. - Fast retransmit: if the sender sees three duplicate ACKs (the receiver repeatedly acknowledging the same byte, implying a specific segment is missing but LATER data did arrive), it retransmits the missing segment immediately, without waiting for the retransmission timer to expire, since three duplicate ACKs is strong, specific evidence of loss rather than simple reordering.
- Fast recovery: after a fast retransmit, rather than collapsing all the way back to slow start,
ssthreshis set to about half the current window, and the window itself is set near that halved value, so the sender doesn't have to re-earn all its previous progress from a window of one segment; it resumes near where it estimates the path can actually sustain.
Worked example
Picture a connection whose window has grown to 64 segments in flight when a single segment is lost and detected via three duplicate ACKs (not a full timeout). Fast retransmit resends the missing segment immediately. Fast recovery sets ssthresh to roughly 32 (half of 64) and the window to near that value, then resumes congestion avoidance's linear growth from there, rather than collapsing to slow start's small initial window and doubling all the way back up. Contrast this with a RETRANSMISSION TIMEOUT (no duplicate ACKs arrived at all, meaning the loss was severe enough that the whole flight of data went missing): that's a much stronger loss signal, and the sender resets ssthresh to half the current window but drops the actual window all the way back to slow start's minimum, since a timeout implies the path may be far more broken than a few duplicate ACKs would suggest.
Trade-offs & pitfalls
It's a common mistake to say TCP always halves its window on any loss and moves on; a full retransmission timeout is treated far more conservatively (full reset to slow start) than a fast-retransmit-detected loss (a much gentler recovery), because the ABSENCE of any duplicate ACKs at all is itself informative: it suggests either a much larger loss event or a badly congested/broken path, not just one unlucky dropped segment.
On a long-fat network (a 10 Gbps link with 150 ms round-trip time), compute the bandwidth-delay product and explain why the classic 16-bit TCP window field cannot describe enough in-flight data to fill this link. Describe how the window-scaling option is negotiated during the handshake to fix this, and what else (beyond the window itself) typically needs tuning to approach line-rate throughput on a link like this.
Sample Answer
Direct answer
The bandwidth-delay product (BDP) is the amount of data that can be "in flight" on a link at any instant, bandwidth multiplied by round-trip time, and it's the minimum window size TCP needs to keep the link fully utilized. On a 10 Gbps link with 150ms round-trip time, the BDP is large enough that the original 16-bit TCP window field (max 65,535 bytes) can represent only a tiny fraction of it, which is exactly why window scaling exists.
Structured elaboration
The BDP calculation is:
BDP (bits)=bandwidth (bits/s)×RTT (s)
For a 10 Gbps link (10×109 bits/s) with a 150ms (0.150 s) round-trip time:
BDP=10×109×0.150=1.5×109 bits=187,500,000 bytes≈187.5 MB
The plain TCP window field is 16 bits, so the largest window it can express without scaling is 216−1=65,535 bytes, about 0.035% of the 187.5 MB needed. Without more window than that, the sender would have to stop and wait for an ACK every 64KB, and at this RTT that caps throughput far below the link's actual 10 Gbps capacity, no matter how fast the link itself is.
The window scale option (negotiated ONLY during the handshake, in the SYN and SYN-ACK) adds a scale factor S (0 to 14) that the receiver applies to its advertised window: the real window becomes advertised value×2S. To cover a 187.5 MB requirement, we need the smallest S such that 65,535×2S≥187,500,000, which comes out to S=12 (giving a maximum representable window of 65,535×4096=268,431,360 bytes, comfortably above the 187.5 MB requirement).
Worked example
For a smaller, more common case, a 100 Mbps link with 100ms RTT, the BDP is:
100×106×0.100=1×107 bits=1,250,000 bytes≈1.25 MB
Here the minimum scale factor needed is S=5 (giving a max window of 65,535×32=2,097,120 bytes), a much smaller ask than the 10 Gbps case, illustrating that window scaling matters more the higher the bandwidth-delay product climbs, not RTT or bandwidth alone. Beyond window sizing, actually reaching close to line rate on a link like this also typically needs the OS socket buffers (net.ipv4.tcp_rmem/tcp_wmem on Linux) raised to match the negotiated window (a window the OS hasn't allocated buffer space for is wasted), and Selective Acknowledgment (SACK) enabled so a single lost segment somewhere in a large in-flight window doesn't force retransmission of everything after it.
Trade-offs & pitfalls
Window scaling is negotiated ONLY at connection setup; if either endpoint doesn't advertise the option in its SYN, the connection falls back to the un-scaled 64KB ceiling for its entire lifetime, a frequent, hard-to-spot cause of "high-bandwidth link, mysteriously capped throughput" that shows up when an old middlebox strips the option or a misconfigured host has scaling disabled.
Implement a simple reliable stop-and-wait protocol over UDP in Python: a send_reliable(sock, dest, payload, timeout) and a matching receive_reliable(sock). Use a single-bit sequence number, ACK packets, retransmit-on-timeout, in-order delivery, and duplicate handling. Explain what your implementation demonstrates about which parts of TCP's reliability UDP does not give you for free.
Sample Answer
Direct answer
Building reliability on top of UDP means implementing, by hand, the exact machinery TCP gives you for free: a sequence number to detect duplicates, explicit acknowledgments, and a retransmission timer. A single-bit (0/1) sequence number is enough for stop-and-wait specifically, because only one message is ever in flight at a time.
Structured elaboration (approach)
send_reliable sends the payload tagged with the current sequence bit, then blocks (with a timeout) waiting for a matching ACK; on a timeout it just resends the same packet, and on receiving an ACK for the WRONG sequence number (a stale ACK from a previous round) it keeps waiting rather than treating that as success. receive_reliable accepts a packet, immediately ACKs it (even if it's a duplicate, in case its own previous ACK was lost), and only hands NEW data (matching the expected next sequence bit) up to the caller, silently absorbing duplicates.
Worked example (code)
import socket, struct
HEADER = struct.Struct("!BB") # (sequence bit, type: 0=DATA, 1=ACK)
def send_reliable(sock, dest, payload, timeout, max_retries=5, state={"seq": 0}):
seq = state["seq"]
packet = HEADER.pack(seq, 0) + payload
sock.settimeout(timeout)
for _ in range(max_retries):
sock.sendto(packet, dest)
try:
data, addr = sock.recvfrom(4096)
except socket.timeout:
continue # retransmit on timeout
if len(data) < HEADER.size:
continue
ack_seq, ack_type = HEADER.unpack(data[:HEADER.size])
if ack_type == 1 and ack_seq == seq:
state["seq"] = 1 - seq
return True
return False
def receive_reliable(sock, state={"expected_seq": 0}):
while True:
data, addr = sock.recvfrom(4096)
if len(data) < HEADER.size:
continue
seq, pkt_type = HEADER.unpack(data[:HEADER.size])
if pkt_type != 0:
continue
payload = data[HEADER.size:]
sock.sendto(HEADER.pack(seq, 1), addr) # always ACK, even duplicates
if seq == state["expected_seq"]:
state["expected_seq"] = 1 - seq
return payload, addr
# else: duplicate, already ACKed above, loop for the real next message
This was executed against a deterministic loss-simulating wrapper (a UDP socket wrapper that drops a configurable fraction of outgoing packets using a seeded random generator, so the test is reproducible) sending 4 messages at loss rates of 0%, 30%, and 60%. At every loss rate tested, all 4 messages were delivered exactly once, in the original order, confirming both the retransmit-on-timeout path and the duplicate-suppression path work correctly under real, repeated loss.
Trade-offs & pitfalls (edge cases and complexity)
Complexity: with a single sequence bit and no pipelining, stop-and-wait can send only ONE unacknowledged message at a time, so throughput is bounded by one round trip per message (a real reliability layer would need a sliding window of sequence numbers, not just one bit, to use a high-bandwidth-delay-product link efficiently, exactly the same motivation as TCP's own window). Edge cases handled: a lost DATA packet (sender times out, retransmits), a lost ACK (receiver gets a duplicate DATA packet, re-ACKs it without re-delivering the payload to the application), and a delayed ACK arriving after the sender has already given up and retransmitted (the sender must ignore an ACK for the WRONG sequence number rather than treating it as confirmation, otherwise a stale ACK could be mistaken for acknowledging the NEXT message). What this exercise demonstrates: TCP is doing exactly this kind of bookkeeping (and much more, for a full sliding window, congestion control, and out-of-order buffering) on every connection, for free.
Explain the difference between a port, a socket, and a connection (the 4-tuple / 5-tuple). What distinguishes well-known, registered, and ephemeral port ranges, and why can many different clients share the same server port on one host without their traffic getting mixed up? Briefly note how NAT changes what a receiver actually observes on the wire.
Sample Answer
Direct answer
A port is a 16-bit number identifying an endpoint on a host; a socket is the (IP address, port, protocol) combination that uniquely names one endpoint; and a connection (the 4-tuple, or 5-tuple if you count the protocol) is the full pairing of BOTH endpoints' sockets, source IP, source port, destination IP, destination port. Many clients can share the same server port because what actually distinguishes their traffic is the FULL 4-tuple, not the destination port alone.
Structured elaboration
Port ranges are conventionally split three ways: well-known ports (0-1023, traditionally requiring elevated privilege to bind on Unix-like systems, and reserved by convention for standard services like 443 for HTTPS), registered ports (1024-49151, registered with IANA for specific applications but not privileged), and ephemeral (or dynamic/private) ports (49152-65535 by IANA convention, though many OSes use a wider practical range), which the OS assigns automatically to the CLIENT side of an outgoing connection.
| Term | Common port table |
|---|---|
| 22 | SSH |
| 53 | DNS |
| 80 | HTTP |
| 443 | HTTPS |
| 3306 | MySQL |
| 3389 | RDP |
A server listening on port 443 can serve thousands of simultaneous clients because the SERVER's port (443) is only ONE piece of what identifies each connection; every incoming connection also carries a distinct client IP and, typically, a distinct client-side ephemeral port, so the full 4-tuple (client IP, client port, server IP, server port) is what the OS actually uses to demultiplex incoming packets to the right connection, even though the server-side half of that tuple (server IP and port) is identical across all of them.
Worked example
Two different clients, 203.0.113.5 and 203.0.113.9, can both have simultaneous connections to a web server at 198.51.100.1:443. Client A's connection might use ephemeral port 51000 and client B's might use 51000 too (a coincidence, since each client picks its own ephemeral ports independently), yet the server has no trouble distinguishing them, because the full 4-tuples are different: (203.0.113.5, 51000, 198.51.100.1, 443) versus (203.0.113.9, 51000, 198.51.100.1, 443). If client A instead opens a SECOND connection to the same server, its OS will typically pick a DIFFERENT ephemeral port for that second connection (since a client can't have two identical 4-tuples open at once to the same destination), giving something like (203.0.113.5, 51001, 198.51.100.1, 443).
Trade-offs & pitfalls
NAT changes what a receiver actually observes on the wire: a device behind a NAT gateway sharing one public IP will have its ephemeral port REWRITTEN by the NAT device (Port Address Translation) so that multiple internal hosts sharing the same public IP can still be distinguished by the server, meaning the source port a server sees is often not the port the ORIGINAL client actually used, a common source of confusion when correlating server-side logs against client-side application behavior.
Unlock Full Question Bank
Get access to all 29 Networking Fundamentals and Protocols interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.