Network Automation and Software-Defined Networking Questions
Programmatic and automated network operation: network automation tooling (Ansible network modules, Netmiko, NAPALM, Nornir, Jinja2 templating, ZTP), choosing between CLI-scraping libraries and model-driven programmability (NETCONF, RESTCONF, gNMI, YANG) across multi-vendor fleets, idempotent and declarative-versus-imperative design, a network source of truth (NetBox) and intended-versus-running configuration drift reconciliation, Git-based versioning and review of device configuration and automation code, pre- and post-change validation (for example Batfish) including lab or emulated test environments, CI/CD and staged, canary or rollback-safe rollout of configuration to large device fleets with concurrency control and credential and secret handling for automation, automated site and leaf-switch provisioning, and event-driven remediation from streaming telemetry. Also software-defined networking: control and data-plane separation, controller architecture and state consistency, controller-managed overlays and tenant network virtualization, SD-WAN controller orchestration and programmable data planes (P4). Boundary: general scripting and Terraform practice, shell craft, IAM, incident response, protocol design, data-centre fabric and WAN topology design (including choosing EVPN-VXLAN or SD-WAN), and network fault diagnosis are covered elsewhere.
Design the orchestration architecture for an SD-WAN spanning several carriers and regions that needs active-active links and application-aware path steering. How does the controller distribute policy, validate carrier performance, and handle keys and failure of the controller?
Sample Answer
Direct answer
Use a layered, controller-based design. An overlay is the virtual network of tunnels that edge routers build between sites, and the underlay is the carrier network (MPLS, internet) it runs over. The design has a management plane (policy authoring, onboarding, monitoring), a redundant control plane (controllers that exchange routes, keys and policy with every edge device over authenticated control connections), and a data plane of edge routers (the router at each site that forwards user traffic) that keep forwarding on their own if the controllers vanish. Run each region's edges against at least two controllers, make carrier links active-active with probe-driven, per-application path selection, and size key lifetimes and policy retention to outlast the longest controller outage you accept. I use Cisco Catalyst SD-WAN vocabulary as the worked reference because its architecture is documented; the pattern is vendor-neutral, and what matters in an interview is the four roles (manager, controller, validator, edge), not the product names.
Architecture and policy distribution
- Roles (Cisco names): SD-WAN Manager (vManage, management plane), SD-WAN Controller (vSmart, the central control plane brain), SD-WAN Validator (vBond, authenticates new devices and helps them through NAT), and edge routers at each site.
- Overlay protocol: OMP (Overlay Management Protocol) runs inside DTLS or TLS control connections (TLS is the encryption used by HTTPS; DTLS is the same idea over UDP) and carries routes, next hops, keys and policy information. Centralized policy (path preference, application steering) is configured once on the control plane and influences how prefixes are advertised to edges, instead of being typed on each router; edges keep local intelligence for site-local decisions.
- Distribution: the manager pushes policy to the controllers, the controllers distribute it to the edges that match the policy's site lists (named groups of sites, for example all sites in one region), so a change is scoped by site list and applied region by region (canary region first), not to every edge at once.
Active-active links and application-aware steering
- Each edge has one transport per carrier (for example MPLS plus internet, or two internet carriers). Tunnels are built over each. Default behavior is load sharing across both (equal-cost style, meaning paths of equal preference are used at the same time) for best-effort traffic.
- Application-aware steering: classify traffic into classes (voice, video, business-critical, bulk), give each class a service-level agreement profile (maximum loss, latency, jitter) and a preferred and fallback path. The edge measures every tunnel and moves a class to a path that meets its profile. Example profile values for illustration only, to be set from your own application requirements: voice at most 1% loss and 150 ms latency, bulk unconstrained. A concrete policy in plain words (illustrative, not vendor syntax): for the 12 sites in the EU region, class
voicehas SLA profileloss <= 1%, latency <= 150 ms, preferred path MPLS, fallback internet; if MPLS breaches the profile for 3 consecutive probe windows, move voice to the internet tunnel and move it back after 10 clean windows; classbulkuses both paths equally. - Avoid flapping: require several consecutive probe windows over threshold before moving a class, and a longer recovery window before moving it back.
Validating carrier performance
The edge devices measure loss, latency and jitter (variation in delay) per tunnel using BFD (Bidirectional Forwarding Detection) probes, small hello packets exchanged through each tunnel, so a missing reply means the path is down or degraded, which also give fast failover. Do not trust only the carrier's own report: keep the edge-measured history per carrier and region, compare it with the contracted SLA, and alert on sustained breaches. Add probes toward the SaaS destinations that matter, because a clean tunnel between two sites does not prove the internet path to an application is clean.
Scale sizing example (computed): each site has two transports (two carrier links) and a tunnel can run between any transport at one site and any transport at the other, so one pair of sites needs 2 x 2 = 4 tunnels. A full mesh is 200 x 199 / 2 x 4 = 79,600 tunnels, and 199 x 4 = 796 BFD sessions per edge. With 8 regional hubs and 192 spokes connecting only to hubs, spoke-to-hub tunnels are 192 x 8 x 4 = 6,144 and a spoke holds 8 x 4 = 32 sessions. The 6,144 counts spoke-to-hub tunnels only; if the 8 hubs also mesh with each other, that adds 8 x 7 / 2 x 4 = 112 tunnels, for 6,256 in total. Each hub then terminates 192 x 4 + 7 x 4 = 796 sessions, the same as one edge in the full mesh, which is why hubs are sized for it. Hub-and-spoke with spoke-to-spoke built on demand (where the platform supports it) trades a hub detour for a probing load the edges can sustain.
Keys
Edges encrypt data-plane traffic with IPsec (the standard suite for encrypting IP packets between two routers). A symmetric key is one secret used for both encrypting and decrypting. Each edge generates a symmetric key per transport location (its link to one carrier) and sends it to a controller; the controller reflects it, meaning passes it on in its reachability advertisements, to the other edges. So if edge A wants to send to edge B over the MPLS link, A uses the key B advertised for that link: B holds the key, the controller only relays it, and the data itself never passes through the controller. Cisco documents that routers regenerate these keys every 24 hours by default. Consequence: because the new keys travel through the controller, key rollover needs a reachable controller, so set the rekey interval much longer than the controller outage you plan for, and verify the vendor's rekey behavior when controllers are unreachable in a lab before relying on it.
Failure of the controller
- Redundancy: if one controller becomes unavailable, the other controllers keep the overlay functioning (documented for this architecture), so give each edge connections to at least two controllers, in different regions and failure domains.
- If all controllers are lost, edges keep forwarding with the routes, policy and keys they already hold (stale routes are routes kept after the controller that taught them is gone). What you lose is change (new sites, policy updates, rekeying, onboarding through the validator). The retention period for stale routes after losing the controller is a configurable setting on these platforms: look up your platform's value and test it by cutting the controller links in the lab.
- Operational: back up controller configuration, practice a rebuild from the manager's backup, and monitor control-connection state as a first-class alarm.
Trade-offs and pitfalls
- Active-active across unequal carriers works only if the policy is class-aware: spreading voice evenly over a congested internet path and a clean MPLS path creates the choppy calls you were trying to avoid.
- Application-aware routing shifts paths on measurements, so measure at an interval and with a window that match the traffic (too aggressive causes flapping).
- Controller redundancy in one region protects against an instance failure, not a region failure.
Your company runs legacy network hardware and wants to try programmable forwarding hardware where the data plane itself can be reprogrammed. How would you judge whether the team is ready, design a low-risk pilot, define success criteria, and make sure production stays safe?
Sample Answer
Direct answer
Do not start with the hardware: start with a specific problem that fixed-function equipment cannot solve or solves poorly, then check readiness in skills, tooling and operations, and only then run a pilot that is off the critical path, can be removed by one routing change, and has numeric exit criteria agreed in advance. If the team cannot name that problem, or cannot run software-style change control on a data plane, the answer is "not yet", and I would say so.
Is the team ready? (a scored checklist)
The data plane is the part of a device that forwards each packet (look at the headers, pick an output port); the control plane is the software that decides the rules the data plane uses, such as routing protocols and controllers. On a fixed-function device (one whose packet-handling logic is built into the chip by the vendor) you can only configure the data plane; a programmable one lets you change what it does. Programmable data plane means P4 (a language for programming the data plane of network devices) running on a target (hardware or software that can execute a P4 program) with a defined architecture (the programmable blocks and interfaces the target offers). An analogy: P4 is the program, the target is the computer it runs on, and the architecture is that computer's instruction set. In practice a P4 program declares tables, each saying what header fields to match and which action to run on a match, for example match on destination address and act by setting the output port. P4 describes packet processing, not how table entries are filled: that is the control plane's job, commonly through P4Runtime, a control plane API for the data plane elements defined by a P4 program. For a first pass you need those four terms (P4, target, architecture, P4Runtime); the details below come in when you reach the pilot.
Score each, yes or no, and require yes on the first three. An illustrative filled-in scorecard for a team that wants custom telemetry: use case yes (it cannot get per-flow latency from its current switches), people no (one engineer knows P4), pipeline yes, software target yes, vendor support no, on-call no. The result is not yet: hire or train a second P4 engineer first.
- A named use case with a measurable gain (custom telemetry, a load balancer, a custom encapsulation, a scrubbing function), and a statement of why existing options (fixed-function features, meaning what the vendor's chip already does; streaming telemetry; a SmartNIC (a network card with its own processor) or a server) are insufficient.
- People: at least two engineers able to read and write P4 and the control software, so the pilot does not depend on one person.
- Pipeline: the team already uses version control, code review and CI (continuous integration, automated tests on every change) for network changes. A reprogrammable data plane is software, and a bug in the program can drop traffic for every flow on that device.
- A software target for test (BMv2, the P4 reference software switch, which its authors state is for developing, testing and debugging and not for production or performance) and a lab device of the real target.
- A vendor support path for the target hardware and its compiler.
- An on-call team trained for a failure mode they have not had before.
If 1 to 3 are not all yes, the readiness result is to fix those first and postpone the pilot.
Low-risk pilot design
- Placement: start with something that cannot break production: replay captured traffic in the lab, then a tap or mirror feed (a copy of real traffic, taken from a tap or a switch mirror port, that the device observes without forwarding it), then one device carrying a small, low-priority slice of real traffic, with ordinary fixed-function paths still available.
- Single writer: several controllers may connect to one device, so P4Runtime needs a way to pick the one allowed to write. Each controller sends a number, the election ID: the client with the highest election ID becomes primary for a role (a named slice of the device's tables), so only one controller writes. Think of it as a talking stick that goes to the highest number. Pin one pilot controller and treat a second writer as an incident.
- Program changes: installing a new pipeline goes through SetForwardingPipelineConfig (the P4Runtime request that loads a compiled program onto the device), with P4Info (a file the compiler produces that lists the program's tables, match fields and actions, so the controller knows what it can write) describing the program's tables and actions. Treat each as a release: compile, test on the software target, test on the lab hardware, then roll to the pilot device.
- Failure behavior: table entries are not required to be deleted when the controller disconnects, so check, on your target, exactly what stays programmed and for how long, and test it by killing the controller in the lab.
- Rollback by routing, not by reprogramming: raise the routing cost or withdraw the route so traffic returns to the legacy path in the time of a routing reconvergence (the routers recomputing paths after a change), and rehearse it before the pilot starts.
A concrete pilot (illustrative numbers)
Take the telemetry case. The legacy fabric cannot report per-flow queueing delay. The pilot is one programmable switch on a mirror feed of a single leaf's uplink, carrying no production forwarding. The mirror sends copies of about 1 percent of the flows to the pilot; the P4 program stamps and records delay for them. The first exit check is the 10 million packet replay below, then 2 weeks on the mirror feed, then, only if every criterion holds, a low-priority slice of real traffic through the device with the routing withdrawal ready.
Success criteria (decided before the pilot)
- Functional: replayed traffic gives identical forwarding decisions to the reference implementation, 0 mismatches over an agreed capture (for example 10 million packets, a number you set from your traffic mix).
- Performance: the target sustains the required rate at your packet-size mix without loss, measured with the lab traffic generator.
- Operations: a new program version goes from commit to pilot device within a stated time using the pipeline; incidents attributed to the pilot stay at zero Sev-1 (the most severe incident class, a customer-visible outage) over the trial period; on-call needed no vendor escalation for routine tasks.
- Value: the named use case delivers its measured gain (for example the telemetry detail you could not collect before).
If the value criterion fails, the pilot ends even if everything else passed.
Keeping production safe
The pilot device must not be a single point of failure; its path has a redundant legacy alternative. Change windows and a freeze rule apply to the pilot as to production. A kill switch (routing withdrawal) is owned by a named person. Every program is reviewed by someone other than the author.
Trade-offs and pitfalls
- Programmability buys flexibility, and the price is that you now own data plane software: the failure surface moves from vendor firmware to your own code.
- Choosing a pilot that is too important makes failure unacceptable and the evidence ambiguous; choosing one too trivial proves nothing: pick a real but non-critical use.
- What flips the answer to "go straight to a fixed-function or SmartNIC solution": the use case is satisfied by an existing feature, or the team has no P4 or CI maturity.
A cloud platform must give each of about 10,000 tenants its own isolated L3 network on shared Linux hosts. How would you build the host networking, handle addressing and tenant routing, isolate east-west traffic, and keep performance and monitoring manageable?
Sample Answer
Direct answer
Give each tenant its own VRF (virtual routing and forwarding instance: a separate routing table) on every host that runs one of its workloads, and carry tenant traffic between hosts inside VXLAN (Virtual Extensible LAN) tunnels with one VNI (VXLAN Network Identifier) per tenant. Use BGP (Border Gateway Protocol) EVPN (Ethernet VPN, a BGP address family that advertises tenant routes and MAC addresses so hosts do not have to flood and learn: the older switch method of sending unknown traffic everywhere and remembering who answers) as the control plane, run by FRR (an open-source routing daemon) on each host. Plain VLANs cannot do the job: a VLAN ID is a 12-bit number, which allows 4094 usable values (0 and 4095 are reserved), and RFC 7348 says that limit is inadequate for multi-tenant environments, while a VNI is a 24-bit value that allows up to 16 million segments. Ten thousand tenants is already 2.4 times the whole VLAN space.
The core of the design is the two lines of control: a VRF per tenant for routing isolation and a VNI per tenant for overlay isolation. The firewall, MTU, monitoring and offload sections below make that design safe and operable at 10,000 tenants.
The path a packet takes through the host
flowchart LR
W["Workload: network namespace or VM"] --> V["veth into the tenant VRF"]
V --> F["nftables forward hook: default drop"]
F --> B["Bridge + VXLAN device, VNI per tenant"]
B --> U["Underlay IP fabric, MTU 1550 or more"]
U --> R["Remote host: VNI, VRF, workload"]
C["FRR BGP EVPN"] -.-> B
| Layer | Linux object | Job |
|---|---|---|
| Workload | Network namespace (netns) or VM tap (a virtual network port for a VM), connected to the host by a veth (a pair of virtual Ethernet ports, like a cable: what goes in one end comes out the other) | A netns gives the workload its own interfaces, routes, firewall and sysctls. A VRF only gives one shared network stack a second routing table, so the workload gets the netns and the host side gets the VRF. |
| Tenant router | VRF device (ip link add vrf-A type vrf table 1001) | The VRF is the tenant's router on this host. Interfaces are attached with ip link set dev NAME master vrf-A, and their connected routes move into the VRF's table (kernel VRF documentation). |
| Tenant tunnel | Bridge enslaved to the VRF plus a VXLAN device with nolearning | Carries the tenant's L3 VNI. Learning is off because EVPN supplies the MAC and route entries instead of flood-and-learn. |
| Address of the host in the underlay | A VTEP (VXLAN tunnel endpoint) IP such as a loopback | Never visible to tenants and never in a tenant VRF. |
One VXLAN device per tenant would mean up to 10,000 devices on a host. FRR's EVPN documentation describes a single VXLAN device mode (ip link add vxlan0 type vxlan dstport 4789 local <VTEP IP> nolearning external vnifilter, then bridge vni add dev vxlan0 vni 100, plus a bridge VLAN id mapped to each VNI with tunnel_info id, plus a VLAN interface on the bridge enslaved to the VRF). In plain words: external makes one device handle many VNIs; vnifilter plus bridge vni add lists which VNIs this device accepts; and tunnel_info id maps a local bridge VLAN id to a global VNI, so the bridge can translate between the two. Each tenant on the host therefore borrows one local VLAN id, and a VLAN id is a 12-bit number with 4094 usable values. That VLAN id only has local meaning (it never leaves the host; the VNI is what travels), so one host can hold at most 4094 tenants this way, which is far above how many tenants have a workload on one host. Hosts create a tenant's VRF and VNI mapping when its first workload lands and remove it with the last one, so the object count per host follows resident tenants and not the 10,000 total.
Addressing and tenant routing
- Tenant address space. Each tenant picks its own CIDR (for example a /24 per subnet carved from the tenant's block). Overlap between tenants is legal because every VRF is a separate table; the platform's IP address management (IPAM) only enforces uniqueness inside one tenant. The demo below gives tenants A and B the same 10.0.1.0/24 and 10.0.2.0/24.
- Underlay. The fabric's own addresses (VTEP IPs, BGP peering) live in the default VRF from a range tenants never see.
- Routing between hosts. Use symmetric routing: the source host routes the packet inside the tenant VRF, encapsulates it with that VRF's L3 VNI, and the destination host decapsulates and routes again. Symmetric routing means both the sending and receiving host do a routing lookup in the tenant's VRF, with the tunnel in between carrying a single per-tenant VNI. Tenant prefixes travel as EVPN type-5 (IP prefix) routes (a route type that says "this address block is reachable behind this host"). In FRR the VRF-to-VNI mapping is a
vrf vrf1stanza containingvni 100;advertise-all-vnigoes insideaddress-family l2vpn evpnof the defaultrouter bgp <ASN>instance (ASN is the autonomous system number); andadvertise ipv4 unicastgoes insideaddress-family l2vpn evpnof the per-tenantrouter bgp <ASN> vrf vrf1instance. FRR documents that prefixes are not exported as type-5 routes until thatadvertiseline exists. - Route targets. FRR derives route targets automatically: import is the wildcard
*:VNIand export is(AS & 0xFFFF):VNI. A route exported for VNI 5001 is therefore imported only by the VRF that owns VNI 5001, which is the control-plane half of tenant isolation. A route target is a label attached to a BGP route that says which VRFs may import it. Worked example (illustrative AS 65001, VNI 5001):65001 & 0xFFFFis 65001, because 65001 is below 65536 and fits in two bytes, so the export label is65001:5001. Tenant A's VRF, which owns VNI 5001, imports anything labelled*:5001(any AS, VNI 5001), so it accepts this route. Tenant B's VRF owns VNI 5002 and imports*:5002, so it ignores the route. The*wildcard on the AS side is what lets every host import routes from any other host's AS, while the VNI half keeps tenants apart. - Leaving the cloud. Overlapping tenants cannot share one flat internet or on-premises route table, so each tenant's egress goes through a per-tenant NAT or gateway attachment, and any shared service (DNS, metadata, a managed database) is reached by an explicit, audited route import into that one VRF, never by leaking whole tables.
Isolating east-west traffic
Isolation comes from stacking independent layers so one mistake does not expose a tenant:
- Routing. A VRF has no route to another tenant's prefixes. The demo shows tenant A's workload receiving no reply for 10.9.9.2, a prefix that exists only in tenant B.
- Overlay. Frames carry a VNI, and a host only accepts VNIs it has configured.
- Host firewall. An nftables table (nftables is the Linux kernel's packet-filtering framework) hooked on
forwardwithpolicy drop, onect state established,related acceptline, and explicit per-tenant allow rules. Only routed traffic crosses this hook, so two workloads of one tenant that share a bridge need their own bridge-level filtering, and a test should confirm which path same-host traffic takes.
table inet tenant_fw {
chain forward {
type filter hook forward priority 0; policy drop;
ct state established,related accept
iifname "vrf-A" ip daddr 10.0.2.0/24 icmp type echo-request accept
counter comment "dropped by default"
}
}
Pitfall found while testing this on a Linux 7.0 kernel: a first version matched iifname "wv1A" (the host side of the workload veth) and silently dropped everything, including the allowed ping. For VRF-routed traffic the forward hook reported the VRF device (vrf-A) as the input interface. With the rule matching vrf-A, tenant A's ping passed and tenant B's identical ping was dropped. Check the interface name your own kernel presents before trusting a rule.
Performance
- MTU (maximum transmission unit). VXLAN adds about 50 bytes of outer Ethernet, IP, UDP and VXLAN headers (RFC 7348), so tenants that get a 1500-byte MTU need at least 1500 + 50 = 1550 on every underlay link, or a jumbo underlay. The demo sets 1550 on the underlay link, the kernel set the VXLAN device to 1500, and a 1500-byte packet with the do-not-fragment bit passed.
- ECMP spreading. RFC 7348 recommends deriving the outer UDP source port from a hash of the inner packet, which lets the fabric's equal-cost multipath (ECMP) spread tenant flows over all uplinks.
- Offload. Offload means letting the network card (NIC) do work the CPU would otherwise do. The kernel segmentation documentation lists UDP-tunnel segmentation types (such as SKB_GSO_UDP_TUNNEL; GSO is generic segmentation offload, splitting large packets into wire-size ones late or in hardware), so tunnel traffic can be segmented and checksummed in the NIC. Confirm each candidate NIC with
ethtool -kand compare iperf3 throughput with and without the tunnel before committing to hardware. - Control plane. Per-host cost follows resident tenants, as above. Watch the BGP table size with
show bgp l2vpn evpn summaryduring the scale test.
Keeping monitoring manageable
| Signal | Where it comes from | Why it matters |
|---|---|---|
| Per-tenant traffic | ip -s link show vrf-A (receive side only) plus the host-side veth counters of the tenant's workloads, summed per tenant by the collector | Bytes and packets per tenant in both directions, one series per tenant |
| Control plane | show bgp l2vpn evpn summary, show vrf vni, show evpn mac vni (FRR) | Session state, VNI-to-VRF mapping, learned MACs |
| Policy drops | nftables counters | Distinguishes a blocked flow from a broken path |
| Isolation | A scheduled negative probe: tenant A must NOT reach a canary only tenant B owns | Catches a leak that no throughput graph shows |
Four counters (receive and transmit bytes and packets) per tenant is 4 x 10,000 = 40,000 time series for the whole fleet. One measured caveat shapes where those counters come from: in the demo run the VRF device's TX counters stayed at 0 for forwarded traffic and only its RX counters moved, so the VRF device alone gives one direction. The other direction comes from the host-side veth of each workload, where RX is what the workload sent and TX is what was delivered to it. The VXLAN device is not a per-tenant source when one device carries every VNI. The collector reads the per-workload counters locally and exports only the per-tenant sum; per-workload series are exported for one tenant only while debugging.
Worked example, executed
The script builds two hosts as namespaces joined by an underlay link, gives tenants A (VNI 5001) and B (VNI 5002) overlapping addresses, and installs by hand the neighbor, forwarding-database and route entries that BGP EVPN would install. Save the script below as tenants.sh in the current directory and run it as root on Linux with iproute2 and iputils-ping, for example docker run --rm --privileged -v "$PWD":/w debian:stable-slim bash -c "apt-get update -qq && apt-get install -y -qq iproute2 iputils-ping && bash /w/tenants.sh".
#!/usr/bin/env bash
# Run as root in a Linux environment with iproute2 and iputils-ping (for example a privileged container).
# Two hypervisor "hosts" (h1, h2) as namespaces joined by an underlay link.
# Each tenant gets: a VRF, a bridge + VXLAN device carrying its L3 VNI, and one workload namespace per host.
set -euo pipefail
ip netns add h1; ip netns add h2
ip link add u1 type veth peer name u2
ip link set u1 netns h1; ip link set u2 netns h2
ip -n h1 addr add 172.16.0.1/30 dev u1; ip -n h2 addr add 172.16.0.2/30 dev u2
ip -n h1 link set u1 mtu 1550 up; ip -n h2 link set u2 mtu 1550 up # 1500 inner + 50 VXLAN overhead
# tenant_setup <name> <vni> <vrf-table> <router-mac-h1> <router-mac-h2> <transit-net>
tenant_setup() {
local t=$1 vni=$2 tbl=$3 m1=$4 m2=$5 n=$6
for h in 1 2; do
# r is the other host (h=1 gives 2, h=2 gives 1); ${!mac} reads the variable whose NAME is stored in mac (m1 or m2)
local r=$((3 - h)); local mac=m$h rmac=m$r
ip -n h$h link add vrf-$t type vrf table $tbl; ip -n h$h link set vrf-$t up
ip -n h$h link add br-$t type bridge
ip -n h$h link set br-$t address ${!mac} master vrf-$t up
ip -n h$h link add vx-$t type vxlan id $vni dstport 4789 local 172.16.0.$h nolearning
ip -n h$h link set vx-$t master br-$t up
ip -n h$h addr add 10.255.$n.$h/32 dev br-$t
# STAND-IN FOR EVPN (two lines): a type-2/type-5 route would tell this host the remote router's MAC and VTEP.
# neigh add ... nud permanent: a fixed ARP entry (remote router IP -> remote router MAC) that never ages out.
# bridge fdb add ... dst: forwarding entry saying frames for that MAC go inside VXLAN to the remote VTEP IP.
ip -n h$h neigh add 10.255.$n.$r lladdr ${!rmac} dev br-$t nud permanent
bridge -n h$h fdb add ${!rmac} dev vx-$t dst 172.16.0.$r self static
# tenant workload: veth into the VRF, namespace w<h>-<t>
ip netns add w$h-$t
ip link add wv$h$t type veth peer name wp$h$t
ip link set wp$h$t netns w$h-$t; ip link set wv$h$t netns h$h
ip -n h$h link set wv$h$t master vrf-$t up
ip -n h$h addr add 10.0.$h.1/24 dev wv$h$t
ip -n w$h-$t addr add 10.0.$h.2/24 dev wp$h$t; ip -n w$h-$t link set wp$h$t up; ip -n w$h-$t link set lo up
ip -n w$h-$t route add default via 10.0.$h.1
# STAND-IN FOR EVPN type-5: the other host's subnet, reachable via its router IP, installed in this tenant's VRF only.
# onlink: accept the next hop as directly reachable on br-<tenant> even though no address is configured on that subnet.
ip -n h$h route add 10.0.$r.0/24 vrf vrf-$t via 10.255.$n.$r dev br-$t onlink
done
}
tenant_setup A 5001 1001 02:00:00:00:0a:01 02:00:00:00:0a:02 1
tenant_setup B 5002 1002 02:00:00:00:0b:01 02:00:00:00:0b:02 2
for h in h1 h2; do ip netns exec $h bash -c 'echo 1 > /proc/sys/net/ipv4/ip_forward'; done
# tenant B alone owns a second subnet on h2
ip -n h2 addr add 10.9.9.1/24 dev wv2B
ip -n w2-B addr add 10.9.9.2/24 dev wp2B
ip -n h1 route add 10.9.9.0/24 vrf vrf-B via 10.255.2.2 dev br-B onlink
echo "--- A w1 -> A w2 (same addresses exist in tenant B)"
ip netns exec w1-A ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> B w2"
ip netns exec w1-B ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)"
ip netns exec w1-B ping -c1 -W1 10.9.9.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)"
ip netns exec w1-A ping -c1 -W1 10.9.9.2 2>&1 | grep -o '[0-9]* packets transmitted, [0-9]* received' || true
echo "--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)"
ip netns exec w1-A ping -c1 -W1 -M do -s 1472 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- MTUs on h1"
for dev in u1 vx-A; do ip -n h1 -o link show $dev | awk '{sub(/@.*/, "", $2); sub(/:$/, "", $2); print $2, $4, $5}'; done
echo "--- tenant VRF route tables on h1"
echo "vrf-A:"; ip -n h1 route show vrf vrf-A
echo "vrf-B:"; ip -n h1 route show vrf vrf-B
Output:
--- A w1 -> A w2 (same addresses exist in tenant B)
2 packets transmitted, 2 received
--- B w1 -> B w2
2 packets transmitted, 2 received
--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)
1 packets transmitted, 1 received
--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)
1 packets transmitted, 0 received
--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)
1 packets transmitted, 1 received
--- MTUs on h1
u1 mtu 1550
vx-A mtu 1500
--- tenant VRF route tables on h1
vrf-A:
10.0.1.0/24 dev wv1A proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.1.2 dev br-A onlink
vrf-B:
10.0.1.0/24 dev wv1B proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.2.2 dev br-B onlink
10.9.9.0/24 via 10.255.2.2 dev br-B onlink
How to read the script: tenant_setup runs once per tenant and builds, on each of the two hosts, a VRF, a bridge and VXLAN device carrying that tenant's VNI, and a workload namespace connected by a veth. The shell trick ${!mac} means "the value of the variable whose name is stored in mac", so on host 1 it reads m1, the argument holding host 1's router MAC; r=$((3 - h)) is arithmetic that gives the other host's number. The only steps that stand in for BGP EVPN are the neigh add ... nud permanent line (a fixed ARP entry for the remote router, never expiring), the bridge fdb add line (which tunnel destination to use for that MAC) and the final route add ... onlink line (the remote subnet, inside this tenant's VRF only; onlink accepts the next hop as directly reachable). In a real deployment EVPN advertises those three facts by itself. Everything else is the same.
Reading it: both tenants reach their own 10.0.2.2 through the same address plan, tenant A has no route to the prefix owned by tenant B, and each VRF holds only its own routes. The static entries stand in for EVPN, and the single-VXLAN-device and FRR lines above come from the FRR documentation and were not run here.
Trade-offs and pitfalls
- EVPN on the host versus a central controller. A controller-driven overlay such as OVN (Open Virtual Network, an Open vSwitch based controller that uses Geneve tunnels, a tunnel format similar in purpose to VXLAN) gives distributed firewalling and L2 features with one control point. EVPN on the host keeps the fabric on standard BGP that a network team already operates and avoids a controller as a single dependency. Recommend EVPN when the team is network-led and the model is routed (L3) per tenant; switch to the controller model if tenants need stretched L2 subnets and live migration with identical addresses.
- VRF is not a security boundary by itself. It separates routing only. A host compromise, a kernel bug or a mistaken route import defeats it, which is why the firewall and the negative probe exist.
- MTU black holes. If ICMP "fragmentation needed" is blocked in the underlay, small pings work while large transfers stall.
- Overlapping prefixes at every shared boundary (NAT, peering, shared services) need per-tenant handling, or the first two tenants who overlap collide there.
You are designing the controller for a software-defined network that must program forwarding state for millions of endpoints. How do you handle state consistency, partitioning, high availability and failure of the controller, and what happens to traffic if the controller is lost?
Sample Answer
Direct answer
Treat the controller as a distributed system whose failure must degrade the network, not stop it. Split state into a small, strongly consistent intent store (replicated by a consensus protocol, a method for several servers to agree on one ordered history of changes, with a majority quorum, meaning a change counts only once more than half the servers have it) and a large, eventually consistent device-state layer that is continuously reconciled. Partition the work so every switch has exactly one active controller owner. Keep the data plane (the switches that actually forward packets) independent of controller liveness, so losing every controller freezes the network at its last programmed state instead of blackholing it.
State consistency
- Intent store (desired state): tenants, segments, endpoint attachments, policy. It is small (thousands to millions of rows, not millions of flow entries) and needs linearizable writes (every read sees the latest committed write, as if there were one copy of the data), so run it on Raft (a consensus protocol where a leader replicates a log and an entry commits once a majority of servers has it). Raft needs a majority to make progress, and 2f+1 servers tolerate f failures: 3 nodes survive 1 failure, 5 nodes survive 2. The reason is that two majorities of the same group always share at least one server, so two halves of a split group can never both commit different histories. With 3 servers a majority is 2: if one fails, 2 remain and commit; if a partition leaves 2 on one side and 1 on the other, only the side of 2 can commit and the lone server stops accepting writes. With 2 servers a majority is also 2, so one failure stops the group, which is why even sizes add cost without adding tolerance.
- Device state (actual state): what each switch really holds. Never assume it equals intent. Each intent change carries a monotonically increasing version; a reconciler (a loop that repeatedly compares actual state with desired state and repairs differences) compares the version and content the switch reports against intent. It is level-triggered: it acts on the current difference, not on a one-time change notification, so a missed message is healed on the next pass rather than lost forever.
- Ordering across switches: a path across several switches is updated make-before-break: install the downstream hops first, flip the ingress last, remove old state after. This avoids a window where ingress sends traffic to a hop that has no entry yet.
Partitioning
- Shard by switch (or by pod/region), not by endpoint. Each switch is owned by one controller instance at a time, held through a lease (ownership that expires unless renewed, so a dead owner cannot hold a switch forever). Assign switches to instances with consistent hashing (a scheme where adding or removing an instance moves only a small share of the assignments) or a static pod map so a failure moves only that instance's switches.
- Use a fencing token on every ownership change (a number that rises with each new owner, which the target uses to reject any writer holding an older number) so a stale owner cannot overwrite a newer one. OpenFlow gives you this directly: a controller asks a switch to make it master (full access, and the switch demotes the previous master to slave, which is read-only), and role requests carry a 64-bit
generation_id; the switch discards a request whosegeneration_idis smaller than the largest it has seen and replies with a Stale error. A trace: controller X becomes master with generation_id 7. The network partitions, controller Y becomes master with generation_id 8, and the switch demotes X to slave. X comes back still believing it is master and sends a role request carrying 7; the switch has seen 8, 7 is smaller, so it replies Stale and X is refused. - Do not program per-endpoint state everywhere. With 2,000,000 endpoints on 2,000 leaf switches, each leaf holds only its local endpoints: 2,000,000 / 2,000 = 1,000 entries per leaf, instead of 2,000,000 on every switch. Remote endpoints are reached through an overlay (a virtual network built from tunnels on top of the physical network, called the underlay; traffic is encapsulated at the edge, with a mapping service that resolves endpoint-to-location) or aggregated routes.
High availability and controller failure
- Run 5 controller instances for 5 shards of 400 switches each (2,000 / 5 = 400), with the intent store on a 3 or 5 node Raft group. A failed instance's 400 switches are re-leased to survivors; a minority partition of the intent store becomes read-only (it stops accepting intent changes) while the data plane keeps forwarding. Split-brain is the failure where two controllers both believe they own a switch; the lease and fencing token described above prevent it.
- Make the control connection redundant: each switch lists more than one controller address and a standby holds the slave role.
What happens to traffic if the controller is lost
A flow entry is one forwarding rule in a switch's table (match these packets, send them out this port), and it may carry a timeout after which the switch deletes it. The switch keeps forwarding with the entries it has when the controller disappears, which is why the network freezes rather than stops. What happens next depends on the switch's failure mode, which you must choose deliberately. In OpenFlow, a switch that loses all controllers enters either fail secure mode (only packets and messages destined to the controller are dropped, and installed flow entries keep forwarding and still expire according to their timeouts) or fail standalone mode (the switch acts as a legacy Ethernet switch or router, usually only on hybrid switches). Consequences for the design:
- Set entry timeouts so critical forwarding state outlives your longest tolerated controller outage. A hard timeout of 5 minutes on an entry turns a 10 minute outage into a blackhole. Timeline: at minute 0 the controllers are lost and every entry still forwards; at minute 5 an entry installed at minute 0 expires and its traffic is dropped (blackholed); from minute 5 to minute 10 that traffic stays dropped because nothing can reinstall the entry.
- Prefer proactive programming (entries pushed before traffic arrives) over reactive programming (the first packet of a flow is sent to the controller as a packet-in message, and the controller then installs the entry). Reactive mode makes every new flow depend on a live controller.
- Run the underlay with a distributed routing protocol so link failures heal without the controller. New endpoint attachments and policy changes stop until it returns.
- On recovery, avoid a thundering herd (everything reconnecting and asking for work at the same moment): when the channel is re-established the switch keeps its entries, and the controller can read them with a flow-stats request to resynchronize. Reconcile in waves: 400 switches at 20 at a time is 400 / 20 = 20 waves, each switch holding about 1,000 local entries, so one shard reads up to 400 x 1,000 = 400,000 entries, paced instead of all at once.
Trade-offs and pitfalls
- A single global strongly consistent store for flow entries does not scale: use it for intent only.
- Strong consistency everywhere reduces availability during a partition; the commit here is to be consistent on intent and eventually consistent on device state, with reconcile as the safety net.
- Common wrong answer: "controller HA means a hot standby" with no answer for split-brain (two controllers both believing they own a switch). A lease plus fencing token (or OpenFlow
generation_id) is the answer. - What flips the design: a small network (under a few hundred switches) can use one Raft-backed controller cluster with no sharding.
That is every published Network Automation and Software-Defined Networking question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.