Network Design and Architecture Questions
Designing networks at the topology level: data-center fabrics (spine-leaf/Clos, ECMP path selection, underlay and overlay, VXLAN, GENEVE and EVPN-VXLAN, multi-pod and multi-site, lossless RDMA and AI fabrics, low-latency designs), campus and branch design (hierarchical access/distribution/core, small-office design), and WAN and backbone architecture (MPLS label switching, segment routing with SR-MPLS and SRv6, RSVP-TE versus SR-TE, provider POPs and private L3 VPN service, SD-WAN and MPLS migration, multihoming, internet edge, transit and peering cost, backbone and interconnect). Covers redundancy and resilience at the link, device and site level, QoS design, bandwidth and capacity growth planning, segmentation and tenant isolation at design level, the management network, inline service placement, physical constraints such as cabling, optics and power, and equipment and vendor selection with justification. The architecture layer above configuration. Excludes protocol configuration mechanics and BGP or OSPF policy, VLAN and STP operation, cloud VPC and hybrid-cloud design, firewall and zero-trust security design, load balancing, fault diagnosis, telemetry collection, device automation or SDN controller operation, and application-level failover and disaster recovery testing, which are covered elsewhere.
Segment a campus network for HR, Finance, Engineering and Guest. Which isolation controls do you apply at layer 2 and which at layer 3, and how does the design handle growth and compliance?
Sample Answer
Direct answer
Segment by what each group must reach, not by org chart. At Layer 2, give each group its own VLANs (virtual LANs, separate broadcast domains), assign users to them with 802.1X authentication (the switch keeps a port closed until the user or device logs in to a RADIUS server, a central server that checks credentials and replies with which VLAN to use), and harden the switch ports. At Layer 3, give each group its own VRF (a private routing table) and let traffic between groups cross only a stateful firewall (it tracks each connection, so replies to allowed traffic are accepted automatically) with a default-deny rule set (anything not explicitly permitted is dropped). Size each group's addresses with doubling headroom as one summarisable block, and keep the evidence auditors need.
Groups, VLANs and addresses
Assumed headcounts and devices per person: Engineering 800 users with 3 devices each, Finance 120 with 2, HR 60 with 2, Guest 400 concurrent clients at peak. Each block is sized for double today's devices, from the campus supernet (the large block that the group blocks are carved from) 10.20.0.0/16. Usable hosts are 2^(32 - prefix) - 2, minus the network and broadcast addresses: /19 gives 2^13 - 2 = 8,190; /23 gives 2^9 - 2 = 510; /24 gives 2^8 - 2 = 254; /22 gives 2^10 - 2 = 1,022.
| Group | VLANs | Block | Usable | Today | At 2x |
|---|---|---|---|---|---|
| Engineering | 301 to 316 (one per closet) | 10.20.0.0/19 | 8,190 | 2,400 (29%) | 4,800 (59%) |
| Finance | 120 | 10.20.32.0/23 | 510 | 240 (47%) | 480 (94%) |
| HR | 110 | 10.20.34.0/24 | 254 | 120 (47%) | 240 (94%) |
| Guest | 190 | 10.20.36.0/22 | 1,022 | 400 (39%) | 800 (78%) |
One block means one line. 10.20.0.0/19 covers 10.20.0.0 to 10.20.31.255, because its third octet runs from 0 to 31, so a single firewall rule or single route entry (permit source 10.20.0.0/19) names all of Engineering. If the closet subnets were scattered across the address plan, the rule would need one line per subnet and another for each new closet. Growing inside the block adds no line.
Engineering's /19 holds 32 subnets of /24; 16 are used today at 150 hosts each (59% full), and the other 16 (10.20.16.0/24 to 10.20.31.0/24) are the growth reserve, so doubling keeps every closet at 59%. Keeping each closet subnet small keeps broadcast domains small. 10.20.35.0/24 is left free so HR can grow to a /23 and stay one summary route. Finance has no free neighbour: at doubling its block is 94% full, so its growth path is a second block from the reserved 10.20.64.0/18 (64 /24s), at the cost of a second line in rules. Infrastructure uses 10.20.254.0/24 (128 /31 point-to-point links) and 10.20.255.0/24 for loopbacks.
Layer 2 controls
802.1X with VLAN assignment is what actually separates the groups at the port. The remaining rows harden the switch so the separation cannot be bypassed (rogue servers, forged ARP, VLAN hopping, which is a user tricking a switch into placing their traffic in another VLAN).
| Control | Purpose | How you prove it |
|---|---|---|
| 802.1X with RADIUS-assigned VLAN (the RADIUS reply carries three standard fields, Tunnel-Type=VLAN, Tunnel-Medium-Type=802 and Tunnel-Private-Group-ID set to the VLAN number, which together tell the switch "put this port in VLAN 120"; RFC 3580) | Port joins the VLAN of the user's group, not of the wall socket | Log in as a Finance and an HR user on the same port and confirm each lands in its own VLAN |
| MAC authentication bypass for printers and phones (the switch uses the device's MAC address as its credential) | Devices that cannot do 802.1X still land in a fixed VLAN | Plug in a printer, confirm its VLAN and that an unknown MAC gets none |
| DHCP snooping | Only trusted ports may send DHCP server replies; builds an IP-to-MAC binding table | Run a rogue DHCP server on a user port and confirm it is dropped |
| Dynamic ARP inspection | Checks ARP packets against the snooping bindings on untrusted ports | Send a forged ARP reply from a user port and confirm the drop is logged |
| Trunk hygiene: pruned allowed-VLAN lists, unused native VLAN (the VLAN whose frames cross a trunk without a tag), no dynamic trunk negotiation, unused ports shut | Stops VLAN hopping and VLAN sprawl | Compare each trunk's allowed list against the VLANs that closet needs |
| Guest wireless with client isolation (the access point refuses to forward traffic between wireless clients) | Guests cannot reach each other or internal networks | Two guest laptops fail to ping each other |
Layer 3 controls
- VRF per group at the distribution layer, with each group's SVIs (switch virtual interfaces, the gateway address for a VLAN) in its own VRF. Between VRFs there is no route unless the firewall provides it.
- Stateful firewall between VRFs: default deny, logged, rules by group summary block (one line per group, thanks to the summarised blocks: a rule for 10.20.0.0/19 covers every present and future Engineering subnet inside it).
- Guest VRF has only a default route to the internet zone, its own DNS and DHCP, and a rate limit.
- Policy matrix: with four groups there are 12 ordered pairs of source and destination group. One is allowed, 11 are denied.
| From | To | Rule | Pairs |
|---|---|---|---|
| HR | Finance | Allow only the payroll interface (one destination, TCP 443) | 1 allowed |
| HR | Engineering, Guest | Deny | 2 denied |
| Finance | HR, Engineering, Guest | Deny | 3 denied |
| Engineering | HR, Finance, Guest | Deny | 3 denied |
| Guest | HR, Finance, Engineering | Deny | 3 denied |
That is 1 allowed and 2 + 3 + 3 + 3 = 11 denied, 12 in all. HR, Finance and Engineering may reach shared services (DNS, DHCP, directory) through the services zone and the internet through the internet zone; both are outside these 12 pairs. Guest does not use the services zone: as the Guest VRF bullet says, it gets only its own DNS and DHCP and a default route to the internet zone, and it cannot reach the directory.
Growth and compliance
- Growth: new group = new VRF, a VLAN, and a /24 from the 10.20.64.0/18 reserve, plus one firewall rule block. Because blocks are summarised, rules do not multiply with VLAN count.
- Compliance: segmentation earns its audit value only when it is shown to work. If Finance handles payment card data, or a regulation limits who can see HR records, an auditor will ask for evidence. Keep the policy matrix under change control, firewall deny logs, 802.1X authentication logs, a quarterly review of the allow list with an owner, and the results of a periodic test that tries to connect from each group to each other group and records the denial.
- Failure mode to design against: a firewall outage cuts inter-group traffic, so run the pair as high availability, and test that failover keeps existing sessions.
You are sizing a new data center fabric for 100,000 servers, each with a 25 Gbps NIC, at a 3:1 oversubscription target at the leaf. How do you work out the number of ToR, leaf and spine switches and their port counts, and when does a single two-tier design stop being enough?
Sample Answer
Direct answer
Work from the server count down to ports, one tier at a time. This is a three-tier fabric, and I use one name per tier throughout: the ToR (top-of-rack switch, the rack-level switch servers plug into, which the question's 3:1 target applies to), the leaf (the middle tier that gathers a pod of ToRs) and the spine (the top tier). Oversubscription is the ratio of server-facing bandwidth to uplink bandwidth. With 48 x 25 Gbps ports per ToR and 4 x 100 Gbps uplinks, each ToR is exactly 3:1. That needs 2,084 ToRs for 100,000 servers. A two-tier design stops working when the spine port count caps the number of ToRs: with 64-port spines that is 64 ToRs, or 3,072 servers. At 100,000 servers you need a third stage: a pod design of 64 pods, 256 leaf switches and 124 spine switches above the 2,084 ToRs.
Step 1: ToR tier (computed)
A Clos network is a layered design where each stage connects to the next through many parallel links; the ToR, leaf and spine tiers here are three stages.
- Downlink 48 x 25 = 1,200 Gbps, uplink 4 x 100 = 400 Gbps, ratio 3:1 (the target).
- ToRs = ceil(100,000 / 48) = ceil(2,083.3) = 2,084. Server ports provided: 2,084 x 48 = 100,032. Uplinks: 2,084 x 4 = 8,336 x 100 Gbps.
- Assumption: one NIC per server on one ToR. Dual-homed servers (two ToRs per rack for resilience) need about 200,000 / 48 = 4,167 ToR ports-worth of switches (4,168 as whole pairs); the same method applies.
Step 2: how far a two-tier design goes
In a two-tier fabric each ToR connects to every spine, so a ToR with 4 uplinks needs 4 spines, and each spine has one port per ToR. The number of ports on a switch is its radix. A 64 x 100 Gbps spine (radix 64) therefore limits the fabric to 64 ToRs: 64 x 48 = 3,072 servers. Reaching 100,000 two-tier would need a spine radix of at least 2,084 ports, which no single fixed switch offers. That is the point where you add a stage.
Step 3: three-stage fabric with pods
Each pod is a small two-tier fabric: ToRs plus 4 leaf switches (one per ToR uplink, so ToR uplink k goes to leaf k). Leaves then uplink to spines, one spine plane per leaf position. A plane is a separate group of spines: leaf 1 of every pod connects only to the spines of plane 1, leaf 2 of every pod only to plane 2, and so on, so four leaf positions give four planes.
- Spine radix 64 means at most 64 pods (one port per pod in each spine). ToRs per pod = ceil(2,084 / 64) = 33, so 36 pods carry 33 ToRs and 28 carry 32 (36 x 33 + 28 x 32 = 2,084).
- Leaf: 64 ports. A pod has up to 33 ToRs and each ToR has one uplink to this leaf, so 33 ports face down. The other 64 - 33 = 31 ports go up, one to each spine of the leaf's plane. Leaf oversubscription 33:31 = 1.06:1. A 33-ToR pod offers 33 x 48 x 25 = 39.6 Tbps at the servers and 4 leaves x 31 x 100 = 12.4 Tbps up, about 3.2:1 end to end.
- Spines: each leaf has 31 uplinks and each goes to a different spine of its plane, so a plane needs 31 spines. Each spine has one port per pod (64 ports, 64 pods), connected to that pod's leaf of the plane's position. 4 planes x 31 = 124 spines, each with all 64 ports used by pods.
- Totals: 2,084 ToRs + 256 leaves (64 pods x 4) + 124 spines = 2,464 switches.
- The design has no headroom: spines are 64 of 64 ports used, so growth means a higher-radix spine (or a super-spine tier).
flowchart TB
PL1[Plane 1: 31 spines] --- P1L1[Pod 1 leaf 1]
PL1 --- P64L1[Pod 64 leaf 1]
PL2[Plane 2: 31 spines] --- P1L2[Pod 1 leaf 2]
PL2 --- P64L2[Pod 64 leaf 2]
PL3[Plane 3: 31 spines] --- P1L3[Pod 1 leaf 3]
PL3 --- P64L3[Pod 64 leaf 3]
PL4[Plane 4: 31 spines] --- P1L4[Pod 1 leaf 4]
PL4 --- P64L4[Pod 64 leaf 4]
P1L1 --- T1[Pod 1: 33 ToRs, each with one uplink to each of the 4 leaves]
P1L2 --- T1
P1L3 --- T1
P1L4 --- T1
P64L1 --- T64[Pod 64: 32 or 33 ToRs, same wiring]
P64L2 --- T64
P64L3 --- T64
P64L4 --- T64
Addressing plan and ARP/ND scale
- Route to the rack means each ToR is the Layer 3 boundary: the point where traffic is routed by IP address rather than switched by MAC address, and the ToR is its servers' gateway. Subnet arithmetic (a /n prefix holds 2^(32-n) addresses): a /26 is 2^6 = 64 addresses, 62 usable after the network and broadcast addresses, enough for 48 servers plus growth. 33 racks x 64 = 2,112 addresses fits a /20, which is 2^12 = 4,096. 64 pods x 4,096 = 262,144 = 2^18 addresses, which is one /14 such as 10.0.0.0/14. Pods then summarize to one /20 each toward the spine, so each spine holds 64 pod summaries instead of 2,084 rack prefixes (the routing table, the prefixes a switch holds, stays small).
- ARP (IPv4, "who has this IP address") and Neighbor Discovery (ND, the IPv6 equivalent) are confined to the rack: a ToR learns at most 62 neighbors. If Layer 2 is stretched with VXLAN/EVPN instead (VXLAN tunnels Layer 2 frames across the routed network and EVPN distributes the MAC and IP addresses over BGP), every tunnel endpoint (VTEP) may have to hold MAC and ARP entries for stretched segments, so limit how many VLANs stretch beyond a pod, and use ARP suppression (a leaf answers ARP itself from what it already knows), to keep tables bounded.
Trade-offs and pitfalls
- 3:1 is a statement about the ToR's server-facing side; a storage or machine-learning cluster that runs hot east-west needs 1:1 or 2:1, which means more uplinks per ToR and a larger fabric.
- Failure: losing a spine in a plane removes 1/31 of that plane's capacity, not a rack.
- Common mistake: sizing from ports alone and forgetting the spine radix ceiling and the cable count (8,336 ToR uplinks plus 256 x 31 = 7,936 leaf-to-spine links).
Design the topology for a small campus with about 200 users and four shared services (directory, DNS, file and print). Which physical topology would you choose, and what would make you change it as the company grows?
Sample Answer
Direct answer
Use a dual-star: two core switches in the server room (the central switches everything else plugs into), with every wiring closet (a small room on each floor holding the switches that user ports cable to) connected to both cores. Do not use a bus (obsolete), a ring (extra hops and a protocol needed to break the loop) or a full mesh (links grow quadratically). Change it when a user subnet nears its address limit, closet uplinks run busy, a second building appears, or the core runs out of ports.
The design
Assumptions: 200 users in four wiring closets (56, 56, 56 and 32 users), 10 wireless access points, 8 printers, 30% spare ports. The four services are directory (Active Directory, the Windows identity service), DNS (the service that turns names into IP addresses), file and print.
+----------+ 2x10G +----------+
| core-a |===========| core-b |
+----------+ +----------+
/ | | \ / | | \
(each closet stack has one 10G cable to core-a and one to core-b)
closet-A closet-B closet-C closet-D
(2 sw) (2 sw) (2 sw) (1 sw)
Wiring: the two switches in a closet are joined by stacking cables, so they behave as one logical switch. Each closet has one 10G uplink cable to core-a and one to core-b. In the three two-switch closets (A, B and C) the two cables leave from different member switches, so losing one switch still leaves a path for the users on the other. Closet D has a single switch, so both of its uplinks leave from that one switch: that survives the loss of a core or a cable but not the loss of that switch, which would take its 32 users off the network until it is replaced. I accept that at this size, keep a cold spare switch, and would add a second switch to closet D as it grows. Stacking cables stay inside the closet and are not counted as uplinks. The 2x10G between the cores is counted as one core-to-core link (a bundle of two cables).
Links: 4 closets x 2 uplinks (one cable to each core) + 1 core-to-core bundle = 9 links. A full mesh of the same 6 devices needs 6 x 5 / 2 = 15.
Access switches (48 ports each): ports needed per closet are users + access points + printers, times 1.3 for spare, rounded up to whole switches.
| Closet | Ports used (users + access points + printers) | With 30% spare | Switches | Ports installed |
|---|---|---|---|---|
| A | 56 + 3 + 2 = 61 | 80 | 2 | 96 |
| B | 61 | 80 | 2 | 96 |
| C | 61 | 80 | 2 | 96 |
| D | 32 + 1 + 2 = 35 | 46 | 1 | 48 |
Worked step for closet A: 56 + 3 + 2 = 61 ports; 61 x 1.3 = 79.3, rounded up to 80; 80 / 48 = 1.67, so 2 switches, which is 96 ports. Closet D: 32 + 1 + 2 = 35; 35 x 1.3 = 45.5, rounded up to 46; 46 fits in 1 switch of 48. That is 7 access switches (2 + 2 + 2 + 1), 7 x 48 = 336 ports. The 10 access points split as 3, 3, 3 and 1 across the closets, and the 8 printers as 2 per closet.
Addressing and VLANs. A VLAN is a virtual LAN: one switch network split into separate broadcast domains, each given its own subnet. How to read a prefix like /25: an IPv4 address has 32 bits and /25 fixes the first 25 as the network part, leaving 32 - 25 = 7 bits for hosts. 2^7 = 128 addresses, minus 2 (the all-zeros network address and the all-ones broadcast address, which no host can use) = 126 usable hosts. In the same way /26 leaves 6 bits: 64 - 2 = 62, and /27 leaves 5 bits: 32 - 2 = 30. The whole block 10.20.0.0/22 leaves 10 bits, so it spans 1,024 addresses, 10.20.0.0 to 10.20.3.255; every subnet below sits inside it and none overlap (a /25 covers 128 addresses, so the next /25 starts 128 higher).
| VLAN | Purpose | Subnet | Usable hosts | Planned use |
|---|---|---|---|---|
| 10 | Users, closet A | 10.20.0.0/25 | 126 | 56 |
| 20 | Users, closet B | 10.20.0.128/25 | 126 | 56 |
| 30 | Users, closet C | 10.20.1.0/25 | 126 | 56 |
| 40 | Users, closet D | 10.20.1.128/26 | 62 | 32 |
| 50 | Servers | 10.20.2.0/27 | 30 | 3 servers plus gateways |
| 60 | Switch management | 10.20.2.32/27 | 30 | 9 switches (7 access + 2 core) |
| 70 | Printers and access points | 10.20.2.64/27 | 30 | 18 (10 access points + 8 printers) |
Services: two directory servers that also run DNS (one per core so either core can fail), and one file and print server. The file and print server is a single point of failure that I accept at this size, and I would say so plainly.
Spanning tree (STP, the protocol that blocks redundant links to prevent loops): switches elect one root bridge, the reference switch every other switch calculates its loop-free path toward. Run Rapid STP (the faster-converging version) with core-a as the root bridge and core-b as the secondary, set by a lower priority number on core-a. With only STP, each closet uplink to the non-root core sits blocked until the other fails. Better still, pair the cores with multi-chassis link aggregation (two core switches present themselves to a closet as one switch, so both uplinks forward at once and neither is blocked).
Is the uplink big enough
Oversubscription is the ratio of the bandwidth users could demand to the bandwidth of the uplink that carries it. Worst case is a 96-port closet: 96 x 1 Gbps = 96 Gbps of port capacity over 2 x 10 Gbps = 20 Gbps of uplink, so 96 / 20 = 4.8 to 1 on paper. In practice, if about 61 hosts are busy at an assumed 20 Mbps each, that is 61 x 20 = 1,220 Mbps = 1.22 Gbps, which is 1.22 / 20 = 6% of the 20 Gbps and 1.22 / 10 = 12% of one 10G uplink if the other fails. Nothing here runs at 100%. Replace the 20 Mbps with your own measured 95th-percentile figure after the first month (sort the month's 5-minute readings and take the value that 95% of them fall below, which ignores brief spikes but shows the sustained busy hours).
What would make me change it
The thresholds in this table are illustrative starting values, not standards. Set yours from your own baseline.
| Trigger | Threshold | Change |
|---|---|---|
| A user subnet fills | Above 75% of a /25, which is 94 hosts (0.75 x 126 = 94.5, rounded down) | Split the VLAN or move to a larger block |
| Uplink load | 95th percentile above 5 Gbps on one uplink | Add uplinks or move to 25G |
| Spare ports | A closet under 10% free | Add a switch before it hits zero |
| Second building or site | Any | Routed link with its own subnet, not a stretched VLAN |
| Services need different trust levels | e.g. a finance share | Put them behind a firewall in their own VLAN |
| Core ports exhausted | Under 4 free 10G ports per core | Add a distribution layer so closets stop landing on the cores |
Past a few hundred more users the usual move is a three-tier design (access switches, a distribution layer that aggregates closets and does the routing between VLANs, and a core on top), which keeps each failure domain small. A failure domain is the set of users or services that one fault can take down together.
Pitfalls
- VLAN 70 spans all closets. That is acceptable here, but each extra stretched VLAN widens the spanning-tree domain.
- Two cores only give redundancy if each closet really has a cable to each. Test by pulling one uplink per closet during the install.
- Put the servers on VLAN 50 behind the cores (cabled to the core switches themselves), not in a closet, so a closet failure does not take out directory and DNS together.
What is the difference between traffic shaping and policing, and how do common queuing approaches decide which packets go first? Where in an enterprise would you apply markings?
Sample Answer
Direct answer
Both policing and shaping hold traffic to a configured rate using a token bucket (a counter that refills at the allowed rate and is spent as packets pass). They differ in what happens to a packet that arrives with no tokens: a policer drops it (or re-marks it), a shaper queues it and sends it later. Queuing is the separate decision of which waiting packet leaves the interface next. Markings (DSCP, Differentiated Services Code Point, a 6-bit value in the IP header) are written once at the edge you trust and read by every queue downstream. On the public internet they are often reset to zero, so they matter only inside networks you or your carrier control.
Policing versus shaping
Cisco's documentation puts it plainly: a policer typically drops traffic, though it can instead change a packet's marking, and a shaper typically delays excess traffic in a buffer. Shaping applies to traffic leaving an interface, while policing can be applied in either direction. The terms in the token bucket:
- CIR (committed information rate): the average rate allowed.
- Bc (committed burst): how many bytes can pass at once, which is the bucket depth.
- Tc (time interval): mean rate = burst size / time interval.
With numbers: CIR = 10 Mbps and Bc = 15,000 bytes (120,000 bits) give Tc = Bc / CIR = 120,000 / 10,000,000 = 0.012 s = 12 ms. The bucket refills completely every 12 ms, and 120,000 bits per 12 ms is the 10 Mbps mean rate. The 15,000 bytes are the same 10 packets of 1,500 bytes used in the example below.
| Policing | Shaping | |
|---|---|---|
| Excess traffic | Dropped or re-marked | Queued, sent later |
| Latency added | None | Up to the queue length |
| Effect on TCP | Loss, so the sender backs off with a saw-tooth (TCP's rate climbs until a packet is lost, then halves, so a plot of its rate looks like saw teeth) | Smoother, with longer round trips |
| Where | Edge, in or out (customer rate limit, control-plane protection) | Egress, to match a slower link or a contracted rate |
Use shaping when you must stay under a carrier's committed rate without losing packets, and policing when you must enforce a limit on traffic you do not want to buffer, such as untrusted or bulk traffic.
Worked example: 20 Mbps offered to a 10 Mbps limit
167 packets of 1,500 bytes arrive 0.6 ms apart (20 Mbps for about 100 ms). The policer has a 15,000-byte bucket (10 packets); the shaper has a 32-packet queue and drains at 10 Mbps (1.2 ms per packet).
PKT = 1500 # bytes per packet
BITS = PKT * 8
RATE = 10e6 # 10 Mbps policer and shaper rate
BURST = 10 * PKT # bucket depth: 15,000 bytes
arrivals = [i * BITS / 20e6 for i in range(167)] # 20 Mbps offered for about 100 ms
# Policer: tokens refill at RATE, a packet that finds too few tokens is dropped
tokens, last, passed, dropped = BURST, 0.0, 0, 0
first_drop = None
for i, t in enumerate(arrivals):
tokens = min(BURST, tokens + (t - last) * RATE / 8)
last = t
if tokens >= PKT:
tokens -= PKT
passed += 1
else:
dropped += 1
first_drop = i if first_drop is None else first_drop
print(f"policer: {passed} passed, {dropped} dropped (first drop at packet {first_drop}), no packet delayed")
# Shaper: packets wait in a 32-packet queue and leave at exactly RATE
QUEUE = 32
serial = BITS / RATE # 1.2 ms per packet
free_at, queued_until, sent, dropped, worst = 0.0, [], 0, 0, 0.0
for t in arrivals:
queued_until = [d for d in queued_until if d > t]
if len(queued_until) >= QUEUE:
dropped += 1
continue
depart = max(t, free_at) + serial
free_at = depart
queued_until.append(depart)
sent += 1
worst = max(worst, depart - t - serial)
print(f"shaper: {sent} sent, {dropped} dropped, worst queueing delay {worst * 1000:.1f} ms, last packet leaves at {free_at * 1000:.1f} ms")
Output:
policer: 92 passed, 75 dropped (first drop at packet 19), no packet delayed
shaper: 114 sent, 53 dropped, worst queueing delay 37.2 ms, last packet leaves at 136.8 ms
The policer passes the 10-packet burst plus what the refill allows (92 packets, about 11 Mbps over this window) and drops 75, with no added delay. The shaper delays instead: the queue builds until the worst packet waits behind 31 others, 37.2 ms of delay for the worst packet, and it still drops 53 once the 32-packet queue fills, because the offered load stays at twice the rate for the whole window. A shaper absorbs bursts shorter than its queue and drops when overload outlasts it.
Tracing the first packets by hand. Packets arrive every 0.6 ms (1,500 x 8 bits / 20 Mbps), and the policer's bucket refills 10,000,000 / 8 x 0.0006 = 750 bytes in each gap. Packet 0 finds 15,000 bytes and spends 1,500, leaving 13,500. Packet 1 gets 750 back and spends 1,500, leaving 12,750. Each packet nets minus 750, so the bucket empties after packet 18. Packet 19 finds only 750 bytes and is dropped; packet 20 finds 750 + 750 = 1,500 and passes; from there the policer alternates pass and drop, which is half of 20 Mbps, the 10 Mbps limit. The shaper instead holds packets: packet 0 leaves at 1.2 ms; packet 1 arrives at 0.6 ms, finds the line busy until 1.2 ms, leaves at 2.4 ms and so waited 0.6 ms; every later packet waits 0.6 ms longer than the one before, so packet 62 waits 37.2 ms, and packet 63 finds all 32 queue slots full and is dropped.
How queues decide which packet goes first
Each row below fixes the weakness of the row above it.
| Method | Rule | Weakness |
|---|---|---|
| FIFO (first in, first out) | Arrival order | A bulk transfer delays voice |
| Strict priority | Serve the high queue until empty | Can starve everything below if unlimited |
| Weighted fair queuing (WFQ, class-based as CBWFQ) | Each class gets a share by weight | No strict latency guarantee |
| Low-latency queuing (LLQ) | Strict priority queue that is capped, plus CBWFQ for the rest | Needs the priority class sized correctly |
Weights in practice: classes weighted 50, 30 and 20 on a congested 10 Mbps link get 5, 3 and 2 Mbps. If the 20 class goes idle, the other two share its capacity in proportion and get 6.25 and 3.75 Mbps. LLQ is the usual enterprise answer: voice in the capped priority queue, other classes by weight. A queue also needs a drop policy for when it fills; plain tail drop (a full queue discards the packet that just arrived) discards new arrivals, while weighted random early detection (WRED) starts dropping packets at random before the queue is full, and does so earlier for lower-priority traffic.
Where to apply markings
First, classification (deciding which class a packet belongs to, by port, address, application signature or existing mark) comes before marking (writing the value). Mark as close to the source as you trust the device, then let every later hop act on the mark.
| Place | What to do |
|---|---|
| Access switch port | The trust boundary (the first device whose arriving marks you stop believing and rewrite yourself). Trust the mark from known devices such as desk phones; re-mark everything from PCs and servers |
| Wireless | Map between the Wi-Fi priority and DSCP (Wi-Fi frames carry a User Priority from 0 to 7 that selects one of four access categories: voice, video, best effort and background; RFC 8325 maps EF to User Priority 6, the voice access category), and do not pass through marks from unauthenticated devices |
| Data center edge | Mark by source and destination, not by what the application sets |
| WAN edge | Re-mark to the carrier's agreed class menu; re-classify inbound traffic |
| Layer 2 and MPLS | A VLAN tag carries a 3-bit priority (CoS, class of service), and MPLS carries a 3-bit Traffic Class, so only 8 classes survive there; by default the MPLS value is the top 3 bits of the DSCP |
The same value appears in different notations: EF is 46 in decimal, 101110 in binary, and 0xb8 in the full type-of-service byte (46 x 4 = 184 = 0xb8), because the DSCP occupies the top 6 bits of that byte.
Why markings are often bleached or ignored on the public internet
The public internet has no end-to-end service contract. RFC 8100 notes that many networks re-mark unknown or unexpected DSCPs to zero when traffic enters, so a mark that is valid in your network can arrive as best effort. One reason is that a carrier cannot let customers who mark everything as high priority claim its premium classes. Only agreed interconnections preserve marks, and even then only for the agreed classes. The practical conclusion is that a mark guarantees nothing across the internet: your benefit comes from the queues at your own egress, where the mark is honored.
Pitfalls
- Shaping above the real link rate, so the queue forms in the carrier's device instead of yours.
- Marking from untrusted hosts, which lets any application take priority.
- An unbounded priority queue that starves other classes.
- Policing TCP bulk traffic too tightly: the resulting loss collapses throughput.
Explain the split between underlay and overlay in a modern data center network. What does each own, and what design problems appear when the boundary between them is drawn badly?
Sample Answer
Direct answer
The underlay is the physical, routed IP network that delivers packets between fabric switches. The overlay is the virtual network built on top: tenant Layer 2 and Layer 3 segments tunnelled across the underlay, typically VXLAN (Virtual Extensible LAN) with a BGP-EVPN control plane. The underlay owns reachability, ECMP (equal-cost multipath: the switches spread traffic across every path of equal cost instead of using just one), MTU (maximum transmission unit: the largest packet a link will carry) and fast failure detection. The overlay owns tenants, segments, MAC and IP reachability for workloads, and policy. Trouble appears when the underlay is asked to know about tenants or when the overlay assumes things the underlay does not provide.
What each owns
| Underlay | Overlay | |
|---|---|---|
| Carries | Outer IP packets between VTEP loopbacks (VTEP: VXLAN tunnel endpoint, the leaf or host that encapsulates) | Tenant frames and packets inside tunnels |
| Protocols | Point-to-point routed links (each link is its own small IP subnet between exactly two switches, with no shared VLAN), eBGP (BGP between different autonomous systems, the usual way to run routing between fabric switches) or an IGP (interior gateway protocol, a single-domain routing protocol such as OSPF or IS-IS), BFD | VXLAN data plane (the part that wraps and forwards packets), BGP-EVPN control plane (the part that advertises which MAC and IP addresses live behind which VTEP, carried in BGP), VNIs (VXLAN Network Identifiers) |
| Knows about | Loopback and link prefixes only: dozens to hundreds of routes | Tenant MACs, IPs, VRFs (virtual routing and forwarding instances: one private routing table per tenant): tens of thousands |
| Scale limit | Number of switches | Number of tenant endpoints and segments |
| Failure handling | Reroute around a link in ECMP | Withdraw routes, move endpoint |
| Identifiers | IP addresses | 24-bit VNI: 2^24 = 16,777,216 segments, versus 4,094 usable VLAN IDs |
The packet shows the boundary
A tenant frame is wrapped in a VXLAN header (8 bytes, carrying the 24-bit VNI), then UDP (destination port 4789 per RFC 7348, source port derived from a hash of the inner headers for ECMP entropy), then an outer IP header addressed to the remote VTEP's loopback. The underlay only ever sees the outer IP and UDP.
A picture of one packet, with illustrative addresses (leaf-1 loopback 10.0.0.1, leaf-2 loopback 10.0.0.2, tenant hosts 192.168.10.5 and 192.168.10.9):
[ outer Ethernet ][ outer IP 10.0.0.1 -> 10.0.0.2 ][ UDP dst 4789, src 51234 ][ VXLAN VNI 10010 ][ inner Ethernet ][ inner IP 192.168.10.5 -> 192.168.10.9 ][ payload ]
underlay forwards on these three outer headers |<---------- tenant frame, invisible to the underlay ---------->|
Entropy here means variety in the header fields that a switch hashes to pick a path. Two flows with different inner addresses get different UDP source ports (51234 here is illustrative), so the hash sees different values and spreads the flows across the equal-cost paths. The source port is picked from a hash of the inner headers, as RFC 7348 recommends.
Worked example: MTU
Encapsulation adds outer IPv4 (20) + UDP (8) + VXLAN (8) + inner Ethernet header (14) = 50 bytes. A host sending a 1,500-byte IP packet therefore produces an outer IP packet of 1,500 + 14 + 8 + 8 + 20 = 1,550 bytes. RFC 7348 says VTEPs must not fragment, and recommends raising the MTU across the physical network. Rule: set the underlay IP MTU at least 1,550 (IPv6 underlay: 1,570) and in practice a jumbo value such as 9,216 on all fabric links, and confirm it hop by hop with a don't-fragment ping sized 1,550 minus 28 = 1,522 bytes of payload. The 28 is the ping's own headers: 20 bytes of IPv4 header plus 8 bytes of ICMP echo header (type, code, checksum, identifier, sequence number), so 1,522 + 28 = 1,550 bytes on the wire; if that ping passes with the don't-fragment bit set, the link carries a 1,550-byte IP packet. If one link is left at 1,500, small packets (pings, TCP handshakes) work and large transfers stall: the classic overlay symptom.
Problems when the boundary is drawn badly
- Tenant routes leaked into the underlay. Every tenant prefix now sits in the fabric IGP, tables grow with tenants, and one tenant's flap triggers fabric-wide reconvergence. Fix: underlay carries loopbacks and links only; tenant routes live in EVPN and VRFs.
- MTU mismatch (above): the overlay assumes the room the underlay never gave.
- No entropy. If the VTEP sets a constant UDP source port, every tenant flow between the same two VTEPs has identical outer addresses, protocol and ports. The underlay's ECMP hash (it hashes outer source and destination IP, protocol and ports) sees one flow per VTEP pair and uses one path, leaving the others idle.
- Shared failure domain. One BGP session or one process carrying both underlay and overlay means a policy mistake in tenant routing drops VTEP reachability. Keep separate sessions or separate address families (the BGP mechanism that carries different kinds of routes, such as plain IPv4 prefixes versus EVPN routes, in distinct lists), and keep policy off the underlay.
- Blind troubleshooting. The overlay can look healthy while the underlay drops, and the reverse. Probe VTEP loopback to loopback in the underlay separately from tenant-to-tenant tests.
- Convergence mismatch. Underlay failure detection (BFD, bidirectional forwarding detection, a fast liveness check) should be faster than the overlay's holdtimers (how long a BGP session waits without hearing from its neighbour before declaring it dead), or the overlay keeps sending to a dead VTEP.
Trade-off
Host-based overlays (VXLAN or Geneve, a similar tunnel format with extensible headers, run on the hypervisor) keep the physical fabric even simpler but make the underlay team blind to tenant flows; network-based overlays on the leaf are more visible but put tenant scale into switch tables.
Unlock Full Question Bank
Get access to all 33 Network Design and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.