Network Design and Architecture Questions
Designing networks at the topology level: data-center fabrics (spine-leaf/Clos, ECMP path selection, underlay and overlay, VXLAN, GENEVE and EVPN-VXLAN, multi-pod and multi-site, lossless RDMA and AI fabrics, low-latency designs), campus and branch design (hierarchical access/distribution/core, small-office design), and WAN and backbone architecture (MPLS label switching, segment routing with SR-MPLS and SRv6, RSVP-TE versus SR-TE, provider POPs and private L3 VPN service, SD-WAN and MPLS migration, multihoming, internet edge, transit and peering cost, backbone and interconnect). Covers redundancy and resilience at the link, device and site level, QoS design, bandwidth and capacity growth planning, segmentation and tenant isolation at design level, the management network, inline service placement, physical constraints such as cabling, optics and power, and equipment and vendor selection with justification. The architecture layer above configuration. Excludes protocol configuration mechanics and BGP or OSPF policy, VLAN and STP operation, cloud VPC and hybrid-cloud design, firewall and zero-trust security design, load balancing, fault diagnosis, telemetry collection, device automation or SDN controller operation, and application-level failover and disaster recovery testing, which are covered elsewhere.
Design a resilient campus network for a building with 1,500 users. Sketch the topology, show where single points of failure are removed, and explain how you would add a second building later.
Sample Answer
Direct answer
For 1,500 users in one building, build a two-tier campus: access switches on each floor, each dual-homed to a pair of distribution switches that also act as the building's core (a collapsed core), joined by an MLAG (multi-chassis link aggregation: one logical port-channel, meaning several physical links bundled as one, spread across two physical switches) so both uplinks forward. Dual-homed means an access switch connects to both distribution switches, and a collapsed core means the distribution pair also does the job of a separate core layer. Routing starts at the distribution pair, which is also where each VLAN's default gateway lives. Every layer-2 and layer-3 element above the access layer has a twin, a firewall pair connects the building to the WAN, and a second building later attaches as its own distribution pair over routed links, with no layer 2 between buildings.
Sizing assumptions (stated, then computed)
6 floors with 250 users each (1,500 users). Each floor needs 250 user ports plus about 10 for access points (the 60 access points below, 10 per floor), 260 ports, and printers fit in the spare capacity; adding 25% spare gives 325, which is 7 switches of 48 ports (336 ports) per floor, 42 access switches in the building. Wireless: 2 devices per user is 3,000 devices, served by about 60 access points (about 50 devices each).
flowchart TB
fw[Firewall pair and WAN edge]
da[dist-a]
db[dist-b]
fw --- da
fw --- db
da ---|MLAG peer link| db
subgraph F["Per floor (6 floors, 7 access switches each)"]
a1[access-f1-1 ... access-f1-7]
ap[Wi-Fi APs on PoE ports]
end
da --- a1
db --- a1
a1 --- ap
da -.->|routed 100G links later| dz[dist-c and dist-d in building B]
db -.-> dz
Each access switch has one 10G uplink to each distribution switch (2 x 10G), so 48 x 1G ports over 20G is 48 / 20 = 2.4:1 oversubscription (if every port sent at full rate at once, there would be 2.4 times more traffic than uplink), normal for office traffic. If one distribution switch fails, each access switch is left with one 10G uplink, 4.8:1 in the worst case, still within what user traffic needs but worth alerting on (48 / 10 = 4.8).
Port budget. Each distribution switch (assume 48 x 10G plus 8 x 100G ports) uses 42 of the 10G ports for the 42 access switches (87.5%, 6 spare) and 6 of the 8 100G ports: 2 for the MLAG peer link, 2 to the firewalls, 2 reserved for building B (75%). A seventh floor would need 7 ports and only 6 are spare, so the growth path is an aggregation switch pair per group of floors, not a bigger distribution switch.
Where each single point of failure is removed
| Layer | Failure removed by | What remains |
|---|---|---|
| Access uplink | Two uplinks to two different distribution switches in an MLAG | The access switch itself: its 48 users lose service. Accepted, with a spare on site |
| Distribution device | Pair of switches, dual power supplies on separate feeds and UPS | Software bug affecting both (stagger upgrades) |
| Default gateway | HSRP or VRRP between the pair (below) | The takeover delay while the standby waits out the hold time (VRRP default about 3.6 s, HSRP recommended 10 s); clients keep the same virtual MAC, so their ARP entries stay valid and only the switches must relearn where that MAC lives |
| Firewall / WAN | Firewall pair and two WAN edge routers (an assumed pair, one per ISP circuit), two ISP circuits | Both circuits sharing one conduit or one upstream: ask each provider for its physical route |
| Wi-Fi controller, DHCP, DNS | Redundant instances | Config drift between instances |
| Power and cooling | Dual feeds per closet | Shared riser |
Segmentation and addressing
All inside 10.20.0.0/16 (the plan was checked for overlaps in Python; it uses 10,496 of 65,536 addresses, 16%):
| Network | VLAN scope | Prefix | Usable | Planned use |
|---|---|---|---|---|
| Wired data | one per floor (6) | /23 each, 10.20.0.0 to 10.20.10.0 | 510 | 250 users, 49% used |
| Voice | one per floor (6) | /24 each, 10.20.12.0 to 10.20.17.0 | 254 | 150 phones per floor assumed, 59% used |
| Staff Wi-Fi | two, floors 1 to 3 and 4 to 6 | /21 each, 10.20.24.0 and 10.20.32.0 | 2,046 | 1,500 devices each, 73% used |
| Guest Wi-Fi | one, isolated | /22, 10.20.40.0 | 1,022 | 400 peak guests, 39% used |
| Servers and printers | one | /24, 10.20.44.0 | 254 | |
| Management | one | /24, 10.20.45.0 | 254 | switches and access points |
| Point-to-point links | /24, 10.20.46.0 | 254 | /31 links for building B later |
How the usable counts arise: a /23 leaves 32 - 23 = 9 host bits, so 2^9 = 512 addresses, minus the network and broadcast addresses is 510 (250 / 510 = 49%). A /24 gives 2^8 - 2 = 254 (150 / 254 = 59%). A /22 gives 2^10 - 2 = 1,022 (400 / 1,022 = 39%). A /21 gives 2^11 - 2 = 2,046 (1,500 / 2,046 = 73%).
Voice gets its own VLAN so it can be prioritized (DSCP, Differentiated Services Code Point, value EF, Expedited Forwarding, is the Telephony class in RFC 4594) and kept apart from data. Guest traffic lives in its own VRF and reaches only the internet through the firewall. Staff Wi-Fi is split in two VLANs rather than one 3,000-address broadcast domain.
Gateways: SVI with HSRP or VRRP
Each VLAN has an SVI (switched virtual interface, the layer 3 interface for a VLAN) on both distribution switches. HSRP (Hot Standby Router Protocol, Cisco proprietary, RFC 2281) or VRRP (Virtual Router Redundancy Protocol, RFC 5798) gives clients one virtual gateway address. HSRP's recommended hello (the keepalive message the active switch sends) is 3 seconds and hold (how long the standby waits without one before taking over) 10 seconds. VRRP's default advertisement interval is 1 second and a backup takes over after 3 x interval + skew. RFC 5798 defines skew as (256 - priority) x interval / 256, a small offset that makes higher-priority backups time out slightly sooner. At the default priority of 100 with the 1 s interval, skew is 156 / 256 = 0.609 s, so takeover is 3 + 0.609 = 3.609 s. The spanning-tree root is the switch every other switch treats as the center of the loop-free tree, so traffic between access switches flows toward it. Tune both pairs so the gateway master is the same switch that is the spanning-tree root for those VLANs. Traced example, for any VLAN or switch still under spanning tree (for example a legacy closet switch with two plain uplinks, one of them blocked, instead of an MLAG bundle): if dist-a is the gateway master but dist-b is the root, that switch's forwarding uplink goes to dist-b, so a packet from a user travels access switch to dist-b, across the MLAG peer link to dist-a to be routed, and back across the peer link, using it twice for no benefit. With both roles on dist-a the forwarding uplink goes to dist-a and the packet is routed immediately. On MLAG-bundled access switches no uplink is blocked, so this penalty largely disappears there, but aligning the roles is still the default so the exceptions behave.
Spanning tree: MST versus PVST
The MLAG makes the access-to-distribution fabric loop-free, so spanning tree becomes a safety net against someone plugging a cable between two closet ports. This plan has 17 VLANs (6 wired data, 6 voice, 2 staff Wi-Fi, guest, servers and printers, management). PVST (per-VLAN spanning tree, one tree per VLAN) runs 17 trees and 17 root elections. MST (multiple spanning tree, IEEE 802.1s) maps many VLANs to a few instances and runs a handful of trees, which is lighter for the switch CPUs and easier to reason about. Choose MST with one instance for floors 1 to 3 and one for floors 4 to 6 if you want some load distribution, or one for everything when the MLAG already balances traffic. Edge ports (those facing users, not other switches) get portfast, which lets the port forward immediately instead of waiting through spanning-tree listening and learning states, and BPDU guard, which shuts the port if a switch-to-switch control frame (BPDU) arrives, so an unauthorized switch is shut down.
Adding the second building
Building B repeats the template with its own distribution pair and 10.21.0.0/16 (no overlap with 10.20.0.0/16). Connect each building B distribution switch to each building A distribution switch with a routed /31 link on the reserved 100G ports, four links in total, running OSPF (Open Shortest Path First) or eBGP (Border Gateway Protocol between autonomous systems). There is no VLAN trunk between the buildings, so a loop in one building cannot reach the other. When there are three or more buildings, insert a dedicated core pair instead of daisy-chaining distribution pairs.
Monitoring
Collect SNMP (Simple Network Management Protocol) and streaming telemetry for interface utilization and errors, syslog for MLAG, spanning-tree and HSRP/VRRP state changes, and flow records for top talkers. Useful alerts: either MLAG member down, spanning-tree topology change counts rising, gateway state flapping, uplink utilization above 70% for a sustained period, DHCP pool above 80%, PoE (Power over Ethernet, the switch powering phones and access points over the cable) budget on any access switch above 80%. The 70% and 80% lines are starting points to tune from your baseline.
Why have most data centers moved from a three-tier access, aggregation and core design to spine-leaf? Describe what each tier does and how east-west and north-south traffic cross the fabric.
Sample Answer
Direct answer
Three-tier (access, aggregation, core) was built for north-south traffic: clients at the edge talking to servers behind the core. Modern workloads talk server to server (east-west), and three-tier handles that badly: traffic climbs to an aggregation pair or the core and back, oversubscription (the ratio of server-facing bandwidth to uplink bandwidth; 6:1 means servers could send six times what the uplinks can carry) compounds at every tier, and Spanning Tree Protocol (STP, the loop-prevention protocol that blocks redundant Layer 2 links) leaves paid-for links idle. Spine-leaf is a two-stage Clos fabric (a network where every leaf connects to every spine): every leaf is the same number of hops from every other leaf, all links are routed and forward at once through equal-cost multipath (ECMP, spreading flows across equal-cost routes), and capacity grows by adding spines or leaves.
What each tier does
Three-tier (legacy)
| Tier | Job | Typical limit |
|---|---|---|
| Access | Server or user ports, the VLAN edge (where each server's VLAN, a Layer 2 broadcast domain, begins) | Few uplinks, each access switch is an oversubscription point |
| Aggregation (distribution) | Aggregates access switches, Layer 2/3 boundary (below it traffic is switched by MAC address, above it routed by IP), gateways, firewalls and load balancers attach here | Pair of boxes; Layer 2 below it means STP |
| Core | Fast Layer 3 transport between aggregation blocks and out to WAN/Internet | Few large chassis; a scaling ceiling |
Spine-leaf
| Tier | Job |
|---|---|
| Leaf (often the top-of-rack, ToR, switch) | Server ports plus uplinks to every spine; usually the Layer 3 gateway; VXLAN tunnel endpoint (VTEP) if an overlay is used (an overlay carries Layer 2 segments in tunnels over the routed underlay, the physical network of leaves and spines) |
| Spine | Pure transit between leaves: no servers attach, no policy, only fast routing |
| Border leaf | A leaf that connects the fabric to firewalls, WAN and Internet |
How traffic crosses the fabric
- East-west (server to server): leaf-1, then any one spine chosen by ECMP, then leaf-2. Always leaf, spine, leaf: three switches, two switch-to-switch hops, for every pair of racks. Inside one rack it is just the ToR.
- North-south (clients or the Internet to servers): enters at the border leaf, crosses a spine to the destination leaf, and the reverse for replies. Firewalls and the WAN edge hang off the border leaf pair so they do not sit in the middle of every east-west flow.
- In three-tier, rack-to-rack traffic stays under one aggregation pair when it can, but crosses the core when the racks sit in different aggregation blocks, so latency and available bandwidth differ by where a workload landed.
Worked example: compounding oversubscription
Assumptions (illustrative, computed): access switch with 48 x 10 Gbps server ports and 2 x 40 Gbps uplinks; eight access switches per aggregation pair; each aggregation switch has 2 x 100 Gbps to the core.
- Access (the switch servers plug into, called the ToR in spine-leaf): 480 / 80 = 6:1. Aggregation block: 8 x 80 = 640 Gbps down against 2 x 2 x 100 = 400 Gbps up = 1.6:1. End to end 6 x 1.6 = 9.6:1 if every link forwards. The ratios multiply because each tier squeezes what the tier below let through: if all servers transmit at once, the access uplinks pass 1/6 of the demand and the aggregation uplinks pass 1/1.6 of that, so a server gets 1 / 9.6 of its line rate, 10 / 9.6 = about 1.04 Gbps of its 10 Gbps.
- If STP blocks one of the two access uplinks (no multi-chassis link aggregation, a feature that lets two switches present themselves as one so a neighbour can use links to both), the access tier is 480 / 40 = 12:1, and the blocked links add cost but no capacity.
- Spine-leaf alternative: 48 x 25 Gbps = 1,200 Gbps down and 4 x 100 Gbps = 400 Gbps up gives 3:1 at the leaf (a worst-case 25 / 3 = 8.3 Gbps per server), spines are non-blocking, and all four uplinks carry traffic. Growth is additive: a fifth uplink and a fifth spine take the leaf to 2.4:1 without redesign.
Trade-offs and pitfalls
- Spine-leaf wants lots of cabling and optics (every leaf to every spine), and each leaf uplink goes to one spine, so a leaf with 4 uplinks reaches only 4 spines, and each spine has one port per leaf, so the spine port count caps the number of leaves, which is why very large sites add a third stage.
- Spine-leaf does not remove the need for Layer 2 adjacency for legacy apps: an overlay (VXLAN tunnels, with EVPN distributing host addresses over BGP) provides it over the routed underlay, rather than stretching VLANs and STP.
- Three-tier remains correct for campus networks, where traffic is mostly north-south. The move is driven by traffic pattern, not fashion.
- Mistake: putting servers on spines or hanging firewalls on a spine pair, which recreates the choke point the design removed.
A new SaaS API will serve 20,000 concurrent users, each making 0.1 requests per second with an average 32 KB request plus response. You want a 2x burst allowance and 20% protocol overhead. Estimate the minimum aggregate WAN capacity in Mbps and state your assumptions.
Sample Answer
Direct answer
About 1.23 Gbps (1,228.8 Mbps) of aggregate WAN capacity, which in practice means one 10 Gbps circuit (or two for redundancy), not a 1 Gbps one. The calculation is requests per second x bytes per request x 8 bits, then the burst and overhead multipliers.
The calculation
requests per secondaverage payloadwith 2x burstwith 20% overhead=20,000×0.1=2,000=2,000×32,000 B×8=512,000,000 bit/s=512 Mbps=512×2=1,024 Mbps=1,024×1.2=1,228.8 MbpsAssumptions I am stating: 1 KB = 1,000 bytes and 1 Mbps = 1,000,000 bit/s (decimal, as carriers sell circuits); the 32 KB covers the request and response together; the 2x burst multiplies the average, and the 20% overhead covers TLS, TCP/IP headers and retransmissions and is applied on top of the burst figure. If 32 KB means 32 KiB (32,768 bytes), the result is 1,258.3 Mbps, so the answer moves by about 2.4%, which does not change the circuit choice.
Direction matters
The aggregate figure adds both directions, but a link is sized per direction. A SaaS API sends far more bytes out (responses) than it receives. Assume 2 KB requests and 30 KB responses (an assumption, since the question gives only the 32 KB total): outbound is 2,000 x 30,000 x 8 x 2 x 1.2 = 1,152 Mbps and inbound is 2,000 x 2,000 x 8 x 2 x 1.2 = 76.8 Mbps. The two add up to 1,228.8 Mbps. The circuit that must carry 1,152 Mbps in the busy direction is the one that decides, and a symmetric 1 Gbps circuit is too small by itself (1,152 Mbps outbound exceeds 1,000 Mbps).
From estimate to circuit
| Question | Answer |
|---|---|
| Minimum aggregate | 1,228.8 Mbps (both directions combined), 1,152 Mbps in the busy direction |
| Standard circuit size | 10 Gbps; 1,152 Mbps uses 11.5% of it, leaving room for growth |
| Redundancy | Two circuits from different providers or paths; each alone can carry 100% of the load, so after a failure the one left runs at 11.5% in the busy direction (1,152 / 10,000). While both are in service and share traffic, each carries about 5.8% (576 Mbps), so the pair survives a loss |
| Growth | At 11.5%, the load can grow about 6.9x before one circuit reaches 80% (80 / 11.52 = 6.9) |
What the estimate leaves out
- Concurrency is a model, not a measurement: 0.1 requests per second per user is an average; real traffic is bursty and correlated (a marketing email, a retry storm). The 2x burst is a placeholder for the ratio of measured peak to average.
- Connections: new TLS handshakes (the opening exchange of certificates and keys that sets up an encrypted HTTPS connection) add certificates and round trips (heavier than the steady state per request) and are not in the 32 KB figure; a caching or CDN layer in front reduces both bytes and circuit size.
- Provisioned vs billed: with 95th-percentile billing, the provider records your usage every 5 minutes for the month, sorts the samples, throws away the busiest 5% (about 36 hours of a 30-day month) and bills the highest sample left, so short bursts are mostly free. With a committed rate you pay for a fixed number of Mbps all month whether you use it or not, so you pay for the burst capacity every day.
Pitfall
Do not round to a 1 Gbps circuit because 1.23 is close to 1: the question's own multipliers put you above it, and any real growth would saturate it immediately.
You are replacing VLANs with a VXLAN/EVPN overlay across a large data center that must support thousands of tenants and layer 2 adjacency between racks. How would you design it, including how tenant VLANs map to VNIs, how broadcast and unknown traffic is handled, what MTU the underlay needs, and how tenants stay isolated while still reaching shared or firewalled services?
Sample Answer
Direct answer
VXLAN is a tunnel that wraps a Layer 2 Ethernet frame inside a UDP/IP packet and carries it between leaf switches over a routed network, so two racks can share one Layer 2 segment without VLANs or Spanning Tree between them. The switch at each tunnel end is a VTEP (VXLAN tunnel endpoint). Build a BGP EVPN control plane (BGP carries the MAC and IP addresses of hosts between VTEPs) with VXLAN encapsulation. Every tenant gets its own IP-VRF (a routing table) identified by one Layer 3 VNI, and each of its VLANs becomes a Layer 2 VNI (VXLAN Network Identifier, the 24-bit segment ID in the VXLAN header; RFC 7348 allows up to 16 M segments). Use symmetric IRB (integrated routing and bridging: the leaf also acts as the router between subnets, and "symmetric" means both the sending and receiving leaf route through the tenant's L3 VNI) with an anycast gateway (the same gateway IP and MAC configured on every leaf, so a host's gateway never changes when it moves) on every leaf, ingress replication for broadcast, unknown-unicast and multicast (BUM) traffic, a routed underlay with an IP MTU (the largest IP packet a link carries, not counting the Ethernet header) of at least 1550 bytes (9,216 recommended on jumbo-capable gear), and a firewall pair on border or service leaves between tenant VRFs and shared services.
1. VLAN to VNI mapping with a structured scheme
Locally a leaf still uses a VLAN ID (12 bits, 4,094 usable) on server ports; that ID is only significant on that leaf and maps to a VNI that is fabric-wide. The scheme below follows these rules:
- Tenants are numbered 1 to 9,999.
- Layer 2 VNI = tenant x 100 + segment number (segments 1 to 99 per tenant). The highest value is 999,999.
- Layer 3 VNI = 5,000,000 + tenant. Range 5,000,001 to 5,009,999, so it cannot collide with any Layer 2 VNI, and it stays under the 24-bit limit of 16,777,215.
- The Route Target and Route Distinguisher are derived from the VNI, so a human can read a route and know the tenant.
| Tenant | Segment (local VLAN) | Layer 2 VNI | Layer 3 VNI |
|---|---|---|---|
| 17 | 10, 20, 30 | 1710, 1720, 1730 | 5000017 |
| 1204 | 10, 20 | 120410, 120420 | 5001204 |
| 2980 | 10, 20, 30, 40 | 298010, 298020, 298030, 298040 | 5002980 |
Reading a VNI: 298030 is tenant 2980, segment 30. Layer 2 stretch between racks exists only for segments whose Layer 2 VNI is configured on both leaves; the VLAN number can differ per leaf.
2. Which EVPN route types carry what
Symmetric IRB, traced (tenant 17, illustrative hosts): host A 10.17.10.5 in segment 10 on leaf-1 sends to host B 10.17.20.7 in segment 20 on leaf-9. Leaf-1 is A's anycast gateway and answers locally, routes the packet into tenant 17's VRF, finds B as a host route learned from a Type 2 route, and sends it in a VXLAN packet with the L3 VNI 5000017 to leaf-9's VTEP. Leaf-9 removes the tunnel header, routes the packet in the same VRF and bridges it to B in Layer 2 VNI 1720. The reply takes the mirror path. Leaf-1 never needed segment 20 configured (RFC 9135). In asymmetric IRB the ingress leaf would route straight into segment 20's VNI, so every leaf would need every segment of the tenant.
The EVPN route types involved:
- Type 3 (Inclusive Multicast Ethernet Tag): each leaf announces, per Layer 2 VNI, that it participates. Receivers build the flood list from these.
- Type 2 (MAC/IP advertisement): a host's MAC and IP learned at a leaf, with its Layer 2 VNI and, for symmetric IRB, the Layer 3 VNI plus the router MAC extended community (RFC 9135), the leaf's own MAC address that the sending leaf uses as the inner destination of the routed packet. Concrete record for host A (illustrative): MAC 00:50:56:aa:17:01, IP 10.17.10.5, Layer 2 VNI 1710, Layer 3 VNI 5000017, next hop leaf-1's VTEP address 10.0.0.1, tenant route target. A Type 2 route is how every other leaf learns "A is behind 10.0.0.1" without flooding.
- Type 5 (IP prefix): subnets and external prefixes into the tenant VRF; the label field carries the Layer 3 VNI (RFC 9136). Use it for the default route from the border leaf and for subnet summaries.
- Types 1 and 4 handle multihoming (servers dual-attached to two leaves) and are used when servers use EVPN multihoming instead of a vendor-specific pair of leaves: Type 1 (Ethernet Auto-Discovery) tells remote leaves which leaves share a server's links so they can send to either and drop a failed one in a single step, and Type 4 (Ethernet Segment) lets those leaves find each other and elect which one forwards flooded traffic to the server.
3. Broadcast, unknown unicast and multicast
Two choices:
- Ingress replication (head-end replication): the leaf copies each BUM frame to every other leaf that announced the VNI via Type 3. No multicast in the underlay.
- Underlay multicast: map VNIs to multicast groups (RFC 7348 style); the underlay needs PIM (Protocol Independent Multicast, the routing protocol that builds multicast trees) and state per group.
With 48 leaves a BUM packet in a VNI present on all of them is sent 47 times from the ingress leaf, against once with multicast. I choose ingress replication: the leaves have the replication capacity for 47 copies of low-rate traffic, and ARP suppression (ARP is the broadcast "who has this IP" question; the leaf answers it from its Type 2 table) removes most of the broadcast. Switch to underlay multicast when many VNIs span most leaves and carry high-rate multicast (for example a market-data feed), where 47 copies at the source would be too expensive.
4. MTU the underlay needs
VXLAN adds, per packet: outer Ethernet 14 + outer IP 20 + outer UDP 8 + VXLAN header 8. That is 14 + 20 + 8 + 8 = 50 bytes added to the inner Ethernet frame. A 1,500-byte inner IP packet becomes a 1,514-byte inner frame, so the outer IP packet is 1514 + 8 + 8 + 20 = 1,550 bytes (the IP MTU the underlay must carry) and the frame on the wire is 1,564 bytes before the frame check sequence (FCS, the 4-byte error check at the end of each Ethernet frame). If servers use 9,000-byte MTU, the underlay IP MTU must be at least 9,050. Set every fabric link, the SVIs (switch virtual interfaces, the routed interface of a VLAN) and loopback paths to 9,216 (check the platform maximum) so one value covers both. RFC 7348 says VTEPs must not fragment VXLAN packets and recommends sizing every MTU on the physical path for the encapsulation; an intermediate router may fragment, and the destination VTEP may silently discard fragments. So a single link left at 1,500 breaks full-size packets between racks (dropped, or fragmented and then discarded, depending on the DF bit and the platform): verify end-to-end with a ping at a size of 1,472 bytes of payload using don't-fragment inside the overlay, then 8,972 if hosts use jumbo frames.
5. Tenant isolation and shared or firewalled services
- Each tenant is a VRF with its own Layer 3 VNI; isolation is that no VRF imports another's Route Target.
- Shared services (DNS, patching) live in a shared-services VRF behind a firewall pair attached to a service leaf. The tenant VRF has a default route (Type 5) originated by the border leaf toward the firewall. The border leaf hands the tenant to the firewall by a VRF-lite sub-interface per tenant (a VLAN-tagged sub-interface on the firewall-facing link, placed in that tenant's VRF, with plain IP routing and no EVPN on it), so the only path from a tenant to shared services is through that firewall, which enforces policy and logs. I deliberately do not leak shared-service prefixes with Route Targets into tenants, because a leaked route is a direct path that never touches the firewall.
- The return path is symmetric: the shared-services VRF sends the tenant's prefix back through the same firewall, so state is kept.
- Capacity: one firewall sub-interface per tenant. As an example, 3,000 tenants need 3,000 sub-interfaces on one pair. One VLAN-tagged link has 4,094 usable VLAN IDs, so tenants beyond that need more than one tagged link or more than one firewall pair, because the numbering scheme itself allows 9,999 tenants. Verify the firewall's per-platform VLAN or sub-interface limit and split tenants across two pairs if needed.
6. Migration order
- Build the underlay and test MTU. 2. Bring up EVPN with one pilot tenant and no VLAN change. 3. Map the tenant's VLANs to Layer 2 VNIs on both leaves. 4. Move the gateway to the anycast gateway: configure the same gateway IP and MAC on every leaf in the segment, then retire the old gateway (for example on an aggregation switch), so hosts keep their default gateway address and only the device answering for it changes. 5. Retire the VLAN on the old aggregation trunks after the host MAC/IP appears as a Type 2 route.
Pitfalls
Mismatched VNI to VRF mapping on two leaves silently breaks symmetric IRB; BUM flooding when ARP suppression is off; leaving a link at the default MTU; sharing a Route Target across tenants.
Your single data center fabric is out of room. How do you scale to multiple pods and then multiple sites, where do you draw failure boundaries, and how do pods communicate with each other?
Sample Answer
Direct answer
Do not keep adding leaves to one fabric. Package the fabric as a repeatable pod (one leaf-spine unit with its own spine layer), join pods through a super-spine layer that is routed, not bridged, and when you outgrow one building or need disaster recovery, build a second site as an independent fabric and join the sites through border gateways over a routed data-center interconnect (DCI). Failure boundaries sit at the pod and at the site, pods talk to each other over equal-cost routed paths through the super-spine, and sites talk through the gateways.
Terms used below: a leaf is the top-of-rack switch servers plug into; a spine connects leaves; ECMP (equal-cost multipath) spreads traffic over every equal path; VXLAN (Virtual Extensible LAN, RFC 7348) tunnels layer 2 frames inside UDP; EVPN (Ethernet VPN, RFC 7432) is the BGP-based control plane that distributes MAC and IP reachability for those tunnels; a VTEP is the tunnel endpoint (normally each leaf); a VRF (virtual routing and forwarding instance) is one private routing table per tenant; blast radius is how much of the estate a single failure can take down; a pod and the layers above it are described next, and the sites come after that.
The pod design, with the numbers
Assumed building block (all figures computed, not measured):
| Item | Value | Result |
|---|---|---|
| Leaf | 48 x 25G server ports, 4 x 100G uplinks (one per spine) | 1,200G down vs 400G up = 3:1 oversubscription |
| Pod | 16 leaves, 4 spines | 768 server ports, 19,200G of server capacity |
| Spine | 32 x 100G ports: 16 to leaves, 4 to super-spines | 20 of 32 ports used (62.5%), 12 free |
| Spine to super-spine | 4 x 100G per spine | 1,600G down vs 400G up = 4:1 |
| Pod egress | 4 spines x 400G | 1,600G, so server capacity to pod egress is 12:1 |
| Super-spine | 4 planes x 4 devices = 16 devices, one link from each pod's spine in its plane | each device uses one port per pod: 4 server pods plus the border pod is 5 of its 32 ports, so 32-port devices allow 32 pods in all, border pod included |
| Scale | 4 server pods = 3,072 server ports; 32 pod slots with one taken by the border pod leaves 31 server pods = 23,808 |
Between two leaves in different pods there are 16 equal-cost paths (4 spines, then 4 super-spines in that spine's plane). Routing on the underlay is eBGP (BGP between different autonomous systems, where an autonomous system, AS, is a group of routers under one routing identity, a number). Adjacent tiers sit in different private ASes, so eBGP runs on every link between them. RFC 7938 documents this pattern (BGP as the only routing protocol in a large Clos data center, with ECMP across all equal paths), and its example scheme gives one AS to all top-tier devices, one AS to each set of second-tier devices in a cluster, and a unique AS to every top-of-rack switch, all from the private range 64512-65534.
flowchart TB
subgraph SS["Super-spine: 4 planes of 4 devices each"]
pl1["plane 1: ss-1a to ss-1d"]
pl2["plane 2: ss-2a to ss-2d"]
pl3["plane 3: ss-3a to ss-3d"]
pl4["plane 4: ss-4a to ss-4d"]
end
subgraph P1["Pod 1"]
a1[spine-1a] --- lf1["leaf-1 to leaf-16, each leaf has one link to every spine"]
b1[spine-1b] --- lf1
c1[spine-1c] --- lf1
d1[spine-1d] --- lf1
end
subgraph P2["Pod 2"]
a2[spine-2a] --- lf2["leaf-1 to leaf-16"]
b2[spine-2b] --- lf2
c2[spine-2c] --- lf2
d2[spine-2d] --- lf2
end
subgraph BP["Border pod"]
bsp[border spines] --- bl[border leaves and gateways, site A]
end
a1 --- pl1
b1 --- pl2
c1 --- pl3
d1 --- pl4
a2 --- pl1
b2 --- pl2
c2 --- pl3
d2 --- pl4
bsp --- SS
bl ---|routed WAN or DCI| bg2[border gateways, site B]
Each pod spine connects to all 4 devices of its own plane, so a leaf-to-leaf flow between pods picks one of 4 spines and then one of 4 super-spines in that plane: 4 x 4 = 16 paths. Border gateways sit in their own border pod and join the fabric the same way any pod does.
Where this design runs hot, and the growth path. The 12:1 pod egress is the two oversubscription stages multiplied: 3:1 at the leaf times 4:1 at the spine uplink, 3 x 4 = 12 (equivalently 19,200G down to 1,600G up). It only works if most traffic stays inside the pod. If inter-pod traffic exceeds one twelfth of the pod's server capacity (1,600G of 19,200G, about 8.3%), the spine uplinks run at 100% and queue. The growth path is already in the port budget: each spine has 12 free ports, so adding 4 more uplinks per spine (8 in total, 24 of 32 ports) together with 4 more super-spines per plane (8 per plane, which makes 4 x 8 = 32 equal-cost paths between pods) doubles pod egress to 3,200G and halves the ratio to 6:1. Ports are a shared budget on each 32-port spine: with 4 uplinks it can face up to 28 leaves, and after the growth step above (8 uplinks) it can face up to 24, so adding uplinks and adding leaves compete for the same ports.
Failure boundaries
| Failure | What is lost | Blast radius (computed) |
|---|---|---|
| Server link or leaf | that rack, 48 server ports | 1.6% of a 4-pod build (48 of 3,072) |
| One spine in a pod | 1 of 4 uplinks per leaf | leaf uplink falls 400G to 300G, ratio 3:1 becomes 4:1, no outage |
| One super-spine | 1 of the 4 uplinks on the one spine in that plane, in every pod | each pod's egress 1,600G to 1,500G (15 of 16 uplinks left); that one spine's uplinks 400G to 300G; the other three spines are unchanged |
| A whole pod (power, bad config push) | 768 server ports | 25% of a 4-pod build, other pods keep forwarding |
| A site | everything there | handled by the second site, not by the fabric |
Rules that keep those numbers true: the underlay carries only point-to-point /31 links and loopbacks (80 links per pod at 2 addresses each is 160 addresses, so one /24 per pod covers the underlay links); changes are pushed to one pod at a time; and layer 2 flooding domains are not shared across pods unless a specific workload needs it, because a flooded storm in one pod would otherwise be replicated into every pod carrying that VNI (VXLAN network identifier).
How pods communicate
- Routed (default). Tenants live in VRFs. A flow from a rack in pod 1 to a rack in pod 2 is routed at the ingress leaf, crosses the underlay over one of the 16 ECMP paths inside a VXLAN tunnel, and is delivered at the egress leaf. Both leaves do the routing step in the tenant VRF and the tunnel carries the packet in between, a pattern called symmetric integrated routing and bridging (symmetric IRB), so each leaf needs only its own local endpoints plus the tenant routes, not every remote segment. No layer 2 state is shared beyond the endpoint's own EVPN routes.
- Bridged, by exception. A VLAN that must exist in two pods gets the same VNI in both. The tunnel crosses the super-spine like any other IP traffic, so the underlay stays loop-free, but the broadcast domain now spans pods.
- Scaling the control plane. Leaves do not form a full mesh of overlay sessions: in-pod leaves peer with a pair of spines, and the pods exchange EVPN routes through the super-spine layer or a pair of dedicated speakers, so the number of BGP sessions grows with devices, not with devices squared. How those sessions are built follows the AS scheme: with a different AS per tier, as above, the overlay sessions are eBGP and the spines relay the EVPN routes, while an iBGP overlay uses route reflectors (BGP speakers that re-advertise routes to all their clients, so leaves need not form a full mesh of sessions) and needs one shared overlay AS. Pick one model and use it in every pod.
Going multi-site
Everything above is one site: pods joined by a super-spine. A second site repeats the same structure in another building, and the only link between the two is the gateway pair described here.
When one site is out of power, space, or tolerance for a shared failure, build the second site as its own fabric with its own pods. Each site gets a pair of border gateways that terminate the local overlay, re-originate routes toward the other site, and carry traffic between sites over routed links (the gateway-and-interconnect model that RFC 9014 describes for EVPN overlays, where a gateway detecting a failure withdraws its routes under normal EVPN procedures, so the other site stops sending to it). The point of the gateway is containment: the other site sees only the gateways, not individual leaves, so a leaf flap or MAC move storm in one site does not become a routing event in the other. Between sites use routed tenant VRFs by default, summarize site address blocks at the gateways (summarization means advertising one covering prefix, for example one /16 per site, instead of every individual subnet, so inside changes are invisible to the other site, at the cost of hiding which specific subnet failed), and give every stretched network an explicit owner and justification.
Pitfalls
- Treating the super-spine as a place to put services. It is a routing layer; border functions and firewalls belong on dedicated border pods.
- Uneven pods. If pod 3 has 20 leaves and others 16, ECMP is still equal per path but the spine uplink ratio is not, and capacity planning becomes per-pod.
- Underlay MTU. VXLAN adds 50 bytes to the tenant's Ethernet frame (outer Ethernet 14, IP 20, UDP 8, VXLAN 8), and VTEPs must not fragment (RFC 7348). A 1,500-byte tenant packet plus its 14-byte inner Ethernet header is 1,514 bytes; adding the 8-byte VXLAN, 8-byte UDP and 20-byte IP headers makes a 1,550-byte outer IP packet, so the underlay needs at least 1,550 bytes of IP MTU (the outer Ethernet header sits outside the IP MTU).
- Stretching everything to avoid re-addressing. It trades a one-time renumbering for a permanent shared failure domain.
Unlock Full Question Bank
Get access to all 26 Network Design and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.