Network Design and Architecture Questions
Designing networks at the topology level: data-center fabrics (spine-leaf/Clos, ECMP path selection, underlay and overlay, VXLAN, GENEVE and EVPN-VXLAN, multi-pod and multi-site, lossless RDMA and AI fabrics, low-latency designs), campus and branch design (hierarchical access/distribution/core, small-office design), and WAN and backbone architecture (MPLS label switching, segment routing with SR-MPLS and SRv6, RSVP-TE versus SR-TE, provider POPs and private L3 VPN service, SD-WAN and MPLS migration, multihoming, internet edge, transit and peering cost, backbone and interconnect). Covers redundancy and resilience at the link, device and site level, QoS design, bandwidth and capacity growth planning, segmentation and tenant isolation at design level, the management network, inline service placement, physical constraints such as cabling, optics and power, and equipment and vendor selection with justification. The architecture layer above configuration. Excludes protocol configuration mechanics and BGP or OSPF policy, VLAN and STP operation, cloud VPC and hybrid-cloud design, firewall and zero-trust security design, load balancing, fault diagnosis, telemetry collection, device automation or SDN controller operation, and application-level failover and disaster recovery testing, which are covered elsewhere.
An enterprise has two upstream ISPs, wants one preferred for outbound traffic and some influence over inbound traffic. Design the internet edge, and tell me what you would verify about the providers themselves before trusting it to survive a failure.
Sample Answer
Direct answer
Terminate each ISP (internet service provider) on its own edge router (edge-a and edge-b), run eBGP (external Border Gateway Protocol, the protocol that exchanges routes between different organizations) to both with your own AS (autonomous system: a network under one routing policy, identified by a number) number and your own address block, and run iBGP (the same protocol between your own routers) between the two edge routers. Outbound you choose the exit with LOCAL_PREF (a value only your AS sees, higher wins). Inbound you can only influence what other networks decide: advertise your block to both providers, make ISP-A the attractive path with AS-path prepending (adding extra copies of your AS number to the route's AS path, the list of networks a route has crossed, so it looks longer) toward ISP-B and, where the provider supports it, provider communities. Before trusting it to survive a failure, verify that the two providers really fail independently (physical path, upstream transit, filtering practice), because the BGP configuration cannot create diversity that is not there. The core of the design is three decisions: LOCAL_PREF for outbound, announcements plus prepending for inbound, and checking provider independence; the rest below supports those.
Outbound: you decide
RFC 4271 evaluates a route's degree of preference (LOCAL_PREF) first and only then compares AS_PATH length, so a higher LOCAL_PREF on routes learned from ISP-A makes ISP-A the exit for every destination regardless of path length. Set it on the inbound policy for ISP-A's session and share it over iBGP so both edge routers agree. If ISP-A fails, its routes disappear and ISP-B's lower-preference routes take over without any other change.
What to receive. A default route (a catch-all entry meaning send anything you have no better route for here) plus a few specifics from each provider is cheap and sufficient if you only need failover. Full tables (a route to every network on the internet, about 1.08 million IPv4 entries alone as of October 2026 according to bgp.potaroo.net, plus the IPv6 table) let you choose exits per destination, at the cost of router memory and operational care. Pick default-plus-partial unless you have a performance reason to steer individual destinations. With a default-only design the BGP session can stay up while the provider's upstream is broken (the default route is still announced), so add active probing of an external target through each link and lower the preference or withdraw announcements when the probe fails.
Inbound: you influence, others decide
Illustrative setup: AS 64500 (a number from the 64496 to 64511 documentation range, RFC 5398), the block 203.0.112.0/23 (illustrative), providers AS 64501 (ISP-A, preferred) and AS 64502 (ISP-B).
| Announcement | To ISP-A | To ISP-B | Why |
|---|---|---|---|
| 203.0.112.0/23 aggregate | plain | plain | Covers everything if one /24 gets filtered |
| 203.0.112.0/24 and 203.0.113.0/24 | plain | prepended 3 extra times | Remote networks see AS paths 3 hops longer through ISP-B, so they prefer ISP-A when other things are equal |
The /24s are the longest prefixes that will propagate: RFC 7454 notes IPv4 prefixes longer than /24 (and IPv6 longer than /48) are generally neither announced nor accepted, so splitting below /24 does not work. A remote network that hears both announcements compares AS paths. Through ISP-A it sees 64501 64500 (2 hops). Through ISP-B it sees 64502 64500 64500 64500 64500 (5 hops, because 64500 was added 3 extra times). With everything else equal, the shorter path through ISP-A wins. Keep the aggregate announced to both so that if ISP-A fails, traffic for the whole block still arrives through ISP-B.
To spread inbound load instead of preferring one provider, prepend the first /24 toward ISP-B and the second /24 toward ISP-A. Each provider then attracts roughly half of the networks that hear both.
Why prepending is weak. Remote ISPs apply their own LOCAL_PREF, and a route from a customer (a network that pays them) usually beats one from a peer (a network they exchange traffic with for free), so a network that is a direct customer of ISP-B will choose ISP-B whatever your prepending says. If you need more control, ask each provider whether it offers communities (BGP tags, RFC 1997) that lower the preference of your route inside its network or restrict where it is announced. The community meanings are provider-defined, so read the provider's published list, and test each one. MED (multi-exit discriminator) only helps choose between several links to the same provider, so it does not help here.
Do not become a transit network (a network that carries other networks' traffic between providers). Announce only your own prefixes outbound to each provider (filter anything learned from ISP-A before it goes to ISP-B), or a failure elsewhere can send ISP traffic through your edge.
Capacity and failure behavior
If peak traffic is 700 Mbps on two 1 Gbps circuits, one circuit failing leaves the other at 70%, which survives. At a 1.2 Gbps peak the one left would need to carry 120% of its capacity and drop traffic, so size each circuit for the full peak (or have a QoS policy that sheds bulk traffic first). Run BFD (bidirectional forwarding detection, a fast liveness check) on the sessions if the providers support it, because BGP alone detects a silent failure only when the hold timer (the time without a keepalive message after which a peer is declared dead) expires.
What to verify about the providers
- Physical diversity. Ask each for its fiber route into your building, separate building entrances, different conduit, different carrier hotel. Two providers in one trench is one failure.
- Upstream diversity. Look at how each provider reaches the rest of the internet (a looking glass, a public web page or server where an operator lets you query its routers, or a public BGP data service, shows the AS paths). If both buy transit from the same upstream, an incident there takes both down.
- Routing hygiene. Confirm both will accept your prefixes. If your block is provider-assigned (lent to you from ISP-A's address space, not owned by you), ISP-B will not announce it without a letter of authorization (LOA) from ISP-A saying you may use it. Separately, RPKI (resource public key infrastructure) lets the address holder publish signed records saying which AS may announce a prefix. Register a ROA (route origin authorization, an RPKI object stating which AS may originate a prefix) for the /23 with maximum length 24. Under RFC 6811, a /24 covered by a ROA that does not match it, for example a ROA with maximum length 23, is Invalid and a provider that drops RPKI-invalid routes will drop your /24s.
- Limits and filters. Ask what they accept (prefix lengths, max-prefix per session; RFC 7454 recommends a limit on routes accepted from a peer) and whether they honor TE communities.
- Capacity and DDoS handling. Whether the upstream capacity is oversubscribed and whether they offer scrubbing.
- Operations. Maintenance notices, escalation contacts, BFD support, SLA credits.
- Test it. Shut each session during a maintenance window and measure recovery time for traffic in both directions.
Pitfalls
- Announcing only /24s without the aggregate, then losing both when one provider filters.
- Assuming prepend 3 means inbound traffic will follow. Measure with flow data after the change.
- Both circuits entering through the same wall.
- A default-only design with no probe.
Your design needs more than 4,096 isolated segments. What are your options, and how do they compare on operational complexity and vendor support?
Sample Answer
Direct answer
A VLAN ID is a 12-bit field, so one Layer 2 domain tops out at 4,094 usable segments (IDs 0 and 4095 are reserved). Past that you have five real options, and in practice the choice is mostly the first one: a VXLAN (Virtual Extensible LAN, a tunnel that carries Ethernet frames inside UDP) overlay with an EVPN control plane (Ethernet VPN, BGP-based MAC and IP distribution); Geneve (Generic Network Virtualization Encapsulation), usually terminated on the hypervisor; QinQ double tagging (IEEE 802.1ad); Shortest Path Bridging (SPB, IEEE 802.1aq); and routed segmentation (one VRF, a virtual routing table, per tenant) with MPLS or plain IP between sites. Geneve is chosen mainly when the hypervisor already owns the tunnels, while QinQ and SPB are narrow cases (small provider-style edges, or a few vendors' campus gear). For a typical enterprise or cloud data center I would pick EVPN-VXLAN, because it gives 16,777,216 segment IDs, an open standards-based control plane, and the widest choice of switch hardware. If the segments never need to share a Layer 2 domain, per-tenant VRFs are cheaper to run than any overlay.
Terms used in the table
- VTEP (VXLAN tunnel endpoint): the switch (or host) port that wraps frames into VXLAN on the way in and unwraps them on the way out. Usually one per leaf.
- EVPN route types (RFC 7432): the kinds of BGP message EVPN uses. Type 2 advertises "this MAC (and IP) lives behind that VTEP"; Type 3 tells other VTEPs "send me broadcast and unknown-destination traffic for this VNI"; Type 1 and Type 4 describe a server wired to two leaves (multihoming).
- Multihoming mode: how a server cabled to two leaves is shared. All-active means both leaves forward its traffic; single-active means only one does at a time. Whether a model supports each mode is a vendor check.
- Anycast gateway: every leaf answers for the same gateway IP and MAC, so a workload that moves to another leaf keeps its default gateway unchanged.
- Flood list: the list of remote VTEPs a leaf copies broadcast and unknown-destination frames to for one VNI, so it grows with the number of leaves that carry that VNI.
How each option gets past 4,094
| Option | Segment space | Operational complexity | Vendor support to check |
|---|---|---|---|
| EVPN-VXLAN | 24-bit VNI (VXLAN Network Identifier), 16,777,216 (RFC 7348) | Medium: BGP underlay plus overlay, VTEP (VXLAN tunnel endpoint) design, 50-byte MTU change | Per model and release: VNI scale, EVPN route types, multihoming mode, anycast gateway. Cross-vendor EVPN interoperability must be proven in a lab |
| Geneve | 24-bit VNI, same 16,777,216 (RFC 8926) | Low on the network (it is just UDP 6081 traffic), high in the virtualization stack that owns the tunnels | Hypervisor or DPU (data processing unit, a programmable NIC) support. Whether a given switch can parse Geneve options is a per-model check, so keep the switches as a plain IP underlay |
| QinQ (802.1ad) | Two 12-bit tags: 4,094 x 4,094 = 16,760,836 usable combinations (4,096 x 4,096 = 16,777,216 raw, before the reserved IDs are removed) | Looks easy, scales badly: still Layer 2, provider switches must learn every customer MAC, and how customer Spanning Tree (STP) frames cross the provider network is an explicit design item (tunnel them or filter them) that is configured differently per platform, so check it before relying on QinQ for loop control | Broad, but the standards body also added Provider Backbone Bridges (802.1ah), introduced to address the customer-MAC learning load that plain 802.1ad leaves on provider switches |
| SPB (802.1aq) | 24-bit I-SID (service identifier), about 16 million services | Medium, but a different control plane to learn | Niche: concentrated in a few vendors, so a later multi-vendor or cloud need is hard |
| VRF per tenant, MPLS or IP between sites | Bounded by VRF and route scale, not by a header field | Low to medium: routing and ACLs only, no Layer 2 stretch | Universal |
Worked example
A platform needs 6,000 isolated tenant segments. A single VLAN space gives 4,094, so the design is 1,906 segments short. In VXLAN, 6,000 VNIs use 0.036% of 16,777,216.
The 16.7 million figure is a header field, not what your leaf supports. Each switch has its own limit on VNIs, MAC entries, ARP (address resolution) and ND (Neighbor Discovery) entries and flood lists. A leaf that serves 200 of the 6,000 tenants only needs 200 of them programmed, so size each leaf for its local tenants and keep the fabric-wide VNI count a planning number. Why the 4,094 limit stops being fabric-wide: on each leaf you map a local VLAN ID to a VNI (for example VLAN 100 on leaf 1 and VLAN 300 on leaf 2 can both map to VNI 10100). The VLAN ID then only labels the segment inside one switch, so the 4,094 limit applies per switch, not to the fabric. Not every platform allows the same VLAN ID to map to different VNIs on different ports or switches, so confirm that in the datasheet.
One leaf in numbers (illustrative values, not any vendor's datasheet). Suppose each of the 6,000 tenants has about 150 endpoints across the fabric, and this leaf serves 200 tenants with about 50 endpoints each:
| Resource on this leaf | Needed | Illustrative limit | Used |
|---|---|---|---|
| VNIs | 200 | 4,000 | 5% |
| MAC entries (200 x 150 endpoints in its VNIs, local plus remote) | 30,000 | 100,000 | 30% |
| ARP and ND entries | 30,000 | 64,000 | about 47% |
The leaf only installs tables for VNIs it carries, so the fabric-wide 6,000 is a planning number and the leaf's own numbers decide the purchase.
MTU: VXLAN adds outer Ethernet 14 + IPv4 20 + UDP 8 + VXLAN 8 = 50 bytes (70 with an IPv6 underlay, whose header is 40). A 1,500-byte inner payload needs a 1,550-byte underlay IP MTU, and 9,000-byte jumbo frames need 9,050. RFC 7348 says VTEPs must not fragment, so set the whole underlay MTU before the first tunnel comes up.
Migration path off shared VLANs, including cloud
The same fabric also has to move existing workloads from shared VLANs into isolated overlays that span on-premises and cloud. I would run it in six steps:
- Measure who talks to whom with flow records (IPFIX or sFlow) for at least a month, so each segment is defined by observed dependencies, not by org charts.
- Build the EVPN-VXLAN fabric beside the old VLAN network and bridge them: each legacy VLAN gets a VNI on the leaf, so a workload keeps its IP and gateway while the Layer 2 path changes underneath it.
- Move the gateway of a VLAN into a tenant VRF on the fabric (anycast gateway), then split the VLAN into per-application segments, one at a time, with a rollback that is simply re-trunking the old VLAN.
- Put default-deny policy between segments at a firewall or on leaf ACLs, and open only the flows seen in step 1.
- Extend to cloud by routing, not by stretching Layer 2: give each segment a non-overlapping CIDR, connect the tenant VRF to the cloud virtual network over a VPN or dedicated link, and apply the cloud's own security groups. Stretched Layer 2 across a WAN carries broadcast, loops and MAC flaps into the cloud.
- Retire the old VLAN only after its flow records show zero traffic for a full business cycle.
Trade-offs and pitfalls
- Recommendation: EVPN-VXLAN for the network, plus host-based overlay where the hypervisor already ships one. What flips it: a small provider-style edge with a few thousand customers and no EVPN skill might accept QinQ; a design where nothing needs Layer 2 adjacency should use per-tenant VRFs and skip the overlay.
- A segment ID is not isolation. Isolation comes from the policy and routing you attach to it; VXLAN itself has no encryption, so use MACsec (link-layer encryption) or IPsec (IP-layer encryption) between sites that you do not control.
- The commonest wrong turn is sizing on the 16.7 million figure instead of each platform's per-switch table limits.
- Mixed vendors: prove route types, multihoming and mobility interoperate in a lab before the purchase order.
Design an EVPN-VXLAN fabric that has to grow to 100,000 endpoints across several sites. Where do the scaling limits show up first, and what design choices keep you inside them?
Sample Answer
Direct answer
I would build four self-contained EVPN-VXLAN fabrics of about 25,000 endpoints each (EVPN is Ethernet VPN, the BGP-based control plane that distributes MAC and IP reachability) and join them through border gateways, instead of one flat 100,000-endpoint domain. The limits appear in this order: the per-leaf hardware tables (MAC, ARP and ND address-resolution entries, host routes), the number of EVPN route paths held by route reflectors and border gateways, the flood replication list for broadcast traffic, and convergence after a gateway or mobility event. The choices that keep you inside them are symmetric IRB, anycast gateways, ARP and ND suppression, routed (not stretched) inter-site traffic, and a hard limit on how many VLANs are allowed to span sites.
Terms in plain words. A leaf is the switch servers plug into; a spine connects leaves to each other. A route reflector (RR) is a BGP router that relays routes between many routers so that each leaf needs sessions to only two spines, not to every other leaf. IRB (integrated routing and bridging) means a leaf also acts as the router between subnets. A border gateway is a leaf-class switch that connects one fabric to other fabrics. ECMP spreads traffic over equal-cost paths.
Which choices matter most. Two decisions change the numbers: splitting into per-site fabrics joined by gateways (items 4 and 6 below), and symmetric IRB (item 1). The remaining items are standard hygiene that keeps those two working.
Reference design and sizing (my assumptions, stated)
- 4 sites, 100,000 / 4 = 25,000 endpoints (VMs, containers, hosts) each.
- 500 endpoints per leaf, so 25,000 / 500 = 50 server leaves per site, 200 server leaves across the four sites.
- Each site: 4 spines with 64 ports, plus 2 border gateway leaves. Every leaf (server or border) connects to every spine, so each spine needs 50 + 2 = 52 ports per site, which is 52 of 64 = 81.25%. The 12 spare ports are the growth path inside a site: growth beyond 64 leaves per spine needs a super-spine tier (a fifth stage: leaf, spine, super-spine, spine, leaf), not more leaves.
- Underlay: eBGP per RFC 7938, one private ASN per leaf, a shared ASN per spine tier. The 16-bit private range 64512 to 65534 holds 1,023 ASNs (RFC 6996), so 200 leaf ASNs fit, and the 4-byte range 4200000000 to 4294967294 gives more room. ECMP and multipath-relax are needed: BGP normally load-shares only over paths with identical AS paths, and because every leaf has its own ASN the paths through different spines differ, so multipath-relax tells BGP to share over them anyway.
Where the limits show up first (and the numbers)
| Rank | Limit | Flat 100,000 domain | Four sites, gateway-based |
|---|---|---|---|
| 1 | Leaf host and MAC tables | If a leaf must hold every host route and MAC of a stretched tenant, 100,000 entries. A hypothetical leaf rated 128,000 entries would sit at 78.1%, with only 28% growth left | A leaf holds its site's endpoints: at most 25,000, which is 19.5% of the same hypothetical 128,000 |
| 2 | EVPN paths on route reflectors (spines) | Up to 100,000 endpoints x 2 advertisers when servers are dual-homed = 200,000 paths | 25,000 x 2 = 50,000 paths per site |
| 3 | Flood replication | With ingress replication a leaf sends one copy per remote VTEP (VXLAN tunnel endpoint) in the VNI per BUM frame (broadcast, unknown unicast, multicast): up to 199 copies | Up to 51 copies inside a site (49 other server leaves plus the 2 border gateways, which sit on the flood list and make the copies to other sites) |
| 4 | Gateways | Not needed | The border pair now carries the whole interconnect, so the limit moves there: if it re-advertises every remote endpoint it holds 100,000 routes, so filter and summarize |
| 5 | Convergence | A flap in one big domain withdraws and re-advertises routes network-wide | A gateway failure is contained to inter-site traffic |
| These are my assumptions, not platform limits. Replace the 128,000 with the datasheet number of the leaf you actually buy, at the release you run. |
Design choices that keep you inside the limits
- Symmetric IRB (a VNI, VXLAN Network Identifier, names each segment and each tenant routing instance): inter-subnet traffic uses the VRF's L3 VNI, and per RFC 9135 each leaf keeps ARP entries only for its local hosts and bridge tables only for locally configured subnets, whereas asymmetric IRB requires every leaf to hold ARP entries and IRB interfaces for all subnets in the VRF. Traced example (illustrative addresses, one tenant VRF): host A 10.1.10.5 in VLAN 10 on leaf-1 sends to host B 10.1.20.7 in VLAN 20 on leaf-9. Symmetric: leaf-1 routes the packet into the VRF, sends it in the L3 VNI to leaf-9 using B's /32 host route, and leaf-9 routes it out to VLAN 20. Leaf-1 never needs VLAN 20 or B's ARP entry. Asymmetric: leaf-1 routes straight into VLAN 20's VNI, so leaf-1 must be configured for VLAN 20 and hold B's ARP entry, and by the same logic leaf-9 must hold VLAN 10 for the reply. Multiply that by every subnet in the VRF and every leaf, and at 100,000 endpoints it is the difference between a scalable and an unscalable leaf.
- Anycast gateway: the same gateway IP and MAC configured on every leaf (each leaf answers its own hosts locally), so a VM keeps its gateway when it moves.
- ARP and ND suppression (ARP asks "who has this IP" and ND is the IPv6 equivalent): RFC 7432 lets a leaf answer an ARP request itself when it already has the binding, which keeps ARP broadcasts off the fabric.
- Gateway-based interconnect (RFC 9014, with the flood replication list, the set of remote VTEPs a leaf copies each broadcast frame to, ending at the gateway instead of at every remote leaf): each fabric is its own failure and routing domain, joined by gateways, with the option of an unknown-MAC route so a leaf can send unknown unicast to the gateway instead of storing every remote MAC. Cost: the gateways become scale and failure points, so they are deployed in pairs and sized first.
- Route reflectors: the spines act as route reflectors inside a site (relaying EVPN routes so leaves need only two sessions); keep the number of EVPN sessions per leaf small (two) and set maximum-prefix limits on every EVPN session.
- Inter-site traffic is routed by default: advertise subnet routes (EVPN type 5, one route per subnet rather than one per host) between sites, and stretch Layer 2 only for the few VLANs that need it (for example a legacy cluster), with an explicit cap on stretched VNIs.
- Mobility controls: MAC mobility sequence numbers (RFC 7432) move a MAC between leaves, so set a duplicate-move threshold to catch loops and VMs that flap.
Worked check
Per site: 50 server leaves + 2 border leaves = 52 leaves, each with one link to each of 4 spines, so each spine uses 52 of its 64 ports (the 200 in the sizing list is the server-leaf total over all four sites, not a per-site count). At 25,000 endpoints per site and 2 advertisers each, a route reflector holds 50,000 paths per site. In a flat design the same route reflector would hold 200,000. If a design assumed a leaf could hold all 100,000 endpoints and your hardware rating is 128,000, you would be at 100,000 / 128,000 = 78.1%. One 25% growth event takes it to 125,000, which is 97.7% of the rating and still fits, with only 3,000 entries spare; the next growth event exceeds it. Headroom is 128,000 / 100,000 = 28% in total, which is too thin to plan on.
Trade-offs and pitfalls
- Recommendation: per-site fabrics with gateway interconnect and symmetric IRB. What flips it: if the endpoint count stays under what one leaf table holds with headroom and the sites are close, one fabric is simpler.
- Do not discover scale limits in production. Fill the tables in a lab at 130% before the design is signed off.
- Silent hosts and mobility generate routes you did not size for; count them.
- The commonest wrong turn is stretching every VLAN everywhere, which makes the 100,000 a single failure domain.
A fast-growing company is paying more every month for internet transit. How would you scale internet egress and peering cost-effectively, and how would you forecast when each investment makes sense?
Sample Answer
Direct answer
Measure what you are billed on first (usually the 95th percentile of 5-minute traffic samples), forecast it with an explicit growth rate, and add a capacity option when its fixed monthly cost falls below the transit spend it removes. Transit is what you pay an upstream ISP to carry your traffic to the whole internet, billed per Mbps with a commit (a minimum monthly volume you pay for even if you use less). Peering is swapping traffic directly with another network instead. In the model below the order is an internet exchange (IXP, a shared switch in a data center where many networks connect) port first, then a private interconnect (PNI, a dedicated direct link to one network) with the largest content source, and extra transit renegotiation throughout. Order each item about three months before its break-even month, because cross-connects and transport take time to deliver.
Options and what each removes
| Option | What it costs | What it removes | Risk |
|---|---|---|---|
| Renegotiate or add a transit commit | Cheaper per Mbps at higher commit | Price per Mbps | Over-commit if growth slows |
| IXP port plus transport | Fixed monthly: port, transport (the circuit from your site to the exchange), colocation (rented rack space in the exchange's data center) | Traffic to networks present at the exchange | Needs traffic to exist there |
| PNI to a large content network | Fixed monthly: ports and cross-connect (the physical cable between your equipment and theirs inside the same data center) | The largest single flows | One partner, you need a backup path |
| CDN or cache nodes at your edge | Hardware and hosting | Repeat content | Only helps cacheable traffic |
Method, with code you can run
The billing figure is usually the 95th percentile of samples: sort a month of 5-minute samples and discard the top 5% (36 hours of a 30-day month, which are free to burst). Check whether your contract measures per link or in aggregate. This script generates 30 days of samples with a pinned seed, computes the figure and finds the month each option pays for itself. The traffic is synthetic: each day is 288 slots of 5 minutes, t is the fraction of the day, and the base rate is 15,000 Mbps overnight (until t = 0.25) then rises along a sine-squared hump to about 24,000 Mbps in mid-afternoon (t is about 0.63) and eases off toward midnight, with 5% random noise per sample. first_month(fixed_cost, offload_share) walks forward one month at a time, growing the billed rate by 8% each month, and returns the first month in which the transit bill it would remove (rate x share of traffic moved x price) is at least the option's fixed monthly cost. Month 0 is today. The inputs (price per Mbps, 8% monthly growth, offload shares and fixed costs) are assumptions to replace with your own data.
import math, random
random.seed(40)
samples = [] # 30 days of 5-minute samples, Mbps
for day in range(30):
for slot in range(288):
t = slot / 288
base = 15000 + 9000 * math.sin(math.pi * max(0, t - 0.25) * 1.3) ** 2
samples.append(base * (1 + random.gauss(0, 0.05)))
ordered = sorted(samples)
p95 = ordered[int(len(ordered) * 0.95) - 1] # drop the top 5% of samples
PRICE = 0.50 # assumed $/Mbps/month, committed transit
GROWTH = 1.08 # assumed 8% traffic growth per month
print(f"samples={len(samples)} p95={p95:,.0f} Mbps max={ordered[-1]:,.0f} Mbps")
print(f"transit bill today = ${p95 * PRICE:,.0f} / month")
def first_month(fixed_cost, offload_share):
for month in range(60):
rate = p95 * GROWTH ** month
if rate * offload_share * PRICE >= fixed_cost:
return month, rate
for name, fixed, share in (("IXP port + transport", 4200, 0.20),
("PNI to a large content network", 9000, 0.35)):
month, rate = first_month(fixed, share)
print(f"{name}: pays for itself in month {month} at {rate:,.0f} Mbps "
f"(saves ${rate * share * PRICE:,.0f} vs cost ${fixed:,})")
Output:
samples=8640 p95=24,424 Mbps max=27,799 Mbps
transit bill today = $12,212 / month
IXP port + transport: pays for itself in month 8 at 45,207 Mbps (saves $4,521 vs cost $4,200)
PNI to a large content network: pays for itself in month 10 at 52,730 Mbps (saves $9,228 vs cost $9,000)
Reading the result
At 8% growth traffic doubles every 9.0 months (ln 2 / ln 1.08). The IXP pays for itself in month 8 and the PNI in month 10. With a three-month lead time (assumed), order the IXP port in month 5 and the PNI in month 7.
The forecast is only as good as its growth rate. Re-running the same model:
| Monthly growth | IXP break-even month | PNI break-even month |
|---|---|---|
| 4% | 14 | 19 |
| 8% | 8 | 10 |
| 12% | 5 | 7 |
So the decision date moves by 9 to 12 months across plausible growth rates. Track actual monthly 95th-percentile growth and re-forecast each quarter.
Operating the edge as it scales
- Spread load across transit links so the 95th percentile on each link, not just in aggregate, is controlled if you are billed per link.
- Keep enough transit headroom that losing a peering or the PNI does not saturate the remaining paths. A PNI that carries 35% of traffic must have a fallback path through transit.
- Offload shares (20% and 35% here; offload means traffic moved off paid transit onto the cheaper path) are assumptions. Measure them from flow data: sum the bytes to the destination networks that are present at the IXP, or that belong to the content network's ASNs (autonomous system numbers, the identifiers each network uses on the internet), before buying anything.
Pitfalls
- Buying a 100G port for a 10G need because the unit price is lower. Include the fixed monthly cost in the model.
- Forecasting from the mean. You pay on the 95th percentile, which was 24,424 Mbps here against a mean of 18,471.
- Treating transit price as constant. Commit tiers change it, so re-run the model after each renegotiation.
How would you design the management network for thousands of devices so that management traffic cannot congest the data plane and is not itself a single point of failure?
Sample Answer
Direct answer
Build a dedicated out-of-band (OOB) management network: a physically separate set of cheap switches, links and routers that carries only management traffic: SSH (remote login), SNMP (the older polling protocol for reading device counters), streaming telemetry (the device pushes its counters on a schedule), syslog (device log messages), NTP (clock synchronisation), config pushes and console (serial access). Because it shares no links or queues with production, management traffic cannot congest the data plane. Because every layer above the rack switch is paired (two uplinks from each rack switch, two aggregation switches per group, two cores), no single link or switch above the rack can cut management to more than one rack. The rack switch itself is a shared failure domain for its 40 devices, which the dual-port rule and the in-band fallback below cover. In-band management (over production links) stays configured only as a tested fallback.
In one line each: separate physical network; two paths above every rack switch; summarised addressing; locked-down access; polling limits on devices; console access as a last resort.
Sizing the example: 3,000 devices
Assumptions (stated, not measured): one management port per device, 40 devices per rack-level management switch, an average of 200 kbit/s of telemetry, logs and polling per device, and 1 Gbit/s management ports.
| Layer | Count and role | Result |
|---|---|---|
| Rack management switch (48 x 1G) | 40 devices + 2 uplinks + 6 spare ports | 3,000 / 40 = 75 switches (48 ports = 40 + 2 + 6) |
| Aggregation group | 20 rack switches per pair of aggregation switches (the middle layer that gathers rack switches; agg-a, agg-b); 3 groups of 20 plus 1 of 15 | 4 groups, 8 aggregation switches |
| OOB core | pair of routers/firewalls (oob-core-a, oob-core-b), each fed by every aggregation switch | 2 |
| Out-of-band to the world | console servers (boxes that connect to each device's serial console port and make it reachable over the network) with a second path (cellular or a separate ISP circuit) at each site | reaches devices when the whole WAN is down |
Load check: 40 devices x 200 kbit/s = 8 Mbit/s per rack switch uplink (0.8% of 1G). A 20-switch group carries 800 devices x 200 kbit/s = 160 Mbit/s, which is 1.6% of one 10G uplink from an aggregation switch to the core. The spare capacity is there for failure and bursts (a firmware download to every device at once), not for steady state.
Design rules
- Separate planes, not just separate VLANs. Every device's dedicated management port goes into a VRF (a separate routing table) on the device, so management routing never mixes with production. A VLAN on shared links would still share fate with the data plane.
- Address plan. One /26 (64 addresses, 62 usable) per rack management switch, carved from 10.250.0.0/19. Counting: a /19 has 32 - 19 = 13 host bits and a /26 has 6, so the /19 holds 2^(26-19) = 128 blocks of /26. Handing out the /26s one after another would not summarise: a group of 20 consecutive /26s is 1,280 addresses, which is not a power of two, so it needs two to four separate prefixes (checked with Python's ipaddress module), not one. So give each aggregation group its own aligned /21 (2,048 addresses = 32 blocks of /26; four /21s fill the /19 exactly): group 1 is 10.250.0.0/21, group 2 is 10.250.8.0/21, group 3 is 10.250.16.0/21 and group 4 is 10.250.24.0/21. Inside a /21 the /26s start at .0, .64, .128 and .192 of each third-octet value. Groups 1 to 3 use 20 of their 32 blocks (group 1 runs from 10.250.0.0/26 to 10.250.4.192/26) and group 4 uses 15 (10.250.24.0/26 to 10.250.27.128/26). That is 75 used and 128 - 75 = 53 spare, 12 per full group and 17 in group 4, so every group can grow to 32 racks without renumbering. The OOB core then carries 4 group routes (the four /21s), not 75.
- Redundancy at every layer. Each rack switch has 2 uplinks to agg-a and agg-b of its group (different line cards, different power feeds). In the sizing assumption each device has one management port, so a rack switch failure cuts management to its 40 devices; for the devices whose loss would hurt most (core routers, spines, firewalls) cable a second management port to a rack switch in a different rack (the 6 spare ports per switch have room), and rely on the tested in-band fallback and the console servers for the rest. Aggregation pairs have two links to each OOB core router. The core pair runs a first-hop redundancy protocol (FHRP, such as VRRP: two routers share one virtual gateway address so hosts keep working if one dies) or ECMP (equal-cost multipath: traffic is spread over several equal-cost paths, and if one dies the rest carry it) with BFD (bidirectional forwarding detection, a tiny fast hello between neighbours that detects a dead link in well under a second). Two kinds of availability arithmetic apply. Devices in series, where every one must work, multiply their availabilities and get worse: two 99.9% devices in a chain give 0.999 x 0.999 = 99.8%. Redundant paths in parallel are the reverse: the pair is down only when both are down, so you multiply the failure probabilities. If each path is 99.9% available (0.001 down) and failures are independent, both are down 0.001 x 0.001 = 0.000001 of the time, so the pair is 1 - 0.000001 = 99.9999%. Independence is the assumption that fails in practice (shared conduit, shared PDU), so audit physical diversity.
- Access control. Management access is reachable only through jump hosts or a bastion (one hardened server that engineers must log in to first before reaching any device) on the OOB core, with AAA (authentication, authorisation and accounting, provided by a central server speaking TACACS+ or RADIUS so each engineer logs in as themselves), per-user logging and a deny-by-default ACL on the core. The management network never routes to the internet or into production.
- Protect production from management load. On devices with a shared CPU, apply control-plane policing (a rate limit on traffic addressed to the device's own CPU) to SSH/SNMP so a polling storm cannot starve routing protocols. Rate-limit and stagger pollers and prefer streaming telemetry (device pushes on a schedule) over aggressive SNMP walks.
- Break-glass. Console servers on the OOB network give serial access when a device's management IP is unreachable. Keep a local emergency account and config backups held off-network.
Failure modes and pitfalls
- A rack switch failure cuts out-of-band management to its 40 single-port devices until it is replaced; keep a cold spare per site and use in-band access and the console servers meanwhile.
- A single OOB core pair is itself a shared fate for all 3,000 devices; mitigate with a second site's core and the cellular console path.
- Management VRF leaks (a route accidentally imported to production) silently re-couple the planes. Verify with a traceroute from a production host to a management address; it must fail.
- Same power and same upstream feed for the OOB switches and the production switches defeats the purpose. Put the OOB gear on separate UPS (battery backup) or at least separate PDU (power distribution unit, the rack power strip) circuits.
- Cost: 75 rack switches, 8 aggregation switches and two cores add real capex, but it is a rounding error next to the cost of a data-plane outage you cannot diagnose because the management path died with it.
Unlock Full Question Bank
Get access to all Network Design and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.