Network Automation and Software-Defined Networking Questions
Programmatic and automated network operation: network automation tooling (Ansible network modules, Netmiko, NAPALM, Nornir, Jinja2 templating, ZTP), choosing between CLI-scraping libraries and model-driven programmability (NETCONF, RESTCONF, gNMI, YANG) across multi-vendor fleets, idempotent and declarative-versus-imperative design, a network source of truth (NetBox) and intended-versus-running configuration drift reconciliation, Git-based versioning and review of device configuration and automation code, pre- and post-change validation (for example Batfish) including lab or emulated test environments, CI/CD and staged, canary or rollback-safe rollout of configuration to large device fleets with concurrency control and credential and secret handling for automation, automated site and leaf-switch provisioning, and event-driven remediation from streaming telemetry. Also software-defined networking: control and data-plane separation, controller architecture and state consistency, controller-managed overlays and tenant network virtualization, SD-WAN controller orchestration and programmable data planes (P4). Boundary: general scripting and Terraform practice, shell craft, IAM, incident response, protocol design, data-centre fabric and WAN topology design (including choosing EVPN-VXLAN or SD-WAN), and network fault diagnosis are covered elsewhere.
Tell me about a time an automated network change caused an unexpected problem in production, or nearly did. What was your role, how did you contain the impact, and what did you change in the automation or rollout process afterward?
Sample Answer
Direct answer
A strong answer is one real incident told in order: the change, the signal that something was wrong, what you did in the first minutes to limit the damage, how you restored service, and the specific changes to the automation and rollout that now prevent a repeat. The example below is an illustrative story with arithmetic you can check; use your own facts.
Story skeleton
Situation. I owned a playbook that kept the VLAN list on 40 access switches in line with our source of truth (the repository that holds the intended state). A cleanup moved VLAN data between two files and I ran the playbook with ios_vlans in state: overridden, which replaces all VLANs on the device with the configuration supplied. The moved file was missing one site's voice VLAN (the separate VLAN, a virtual LAN that splits one physical switch into isolated networks, that carries IP phone traffic; without it on a switch, the phones plugged into it cannot register with the phone system).
What happened. The first batch was a single canary switch (one device changed first, as an early warning, before the rest). The task itself finished green, but a check I had added after an earlier near-miss, a comparison of the VLAN count before and after the change, flagged that the canary had lost a VLAN. Because I had rolled out in batches (1, 2, 3, ...), the other 39 switches were untouched. serial is the Ansible play keyword that sets how many hosts are changed at a time; with it set to a list of increasing batch sizes, a failure scopes to the batch, not the whole host list (Ansible documentation).
Containment (my role).
- Stopped the run, and told the on-call and the site owner what had changed and on which switch.
- Restored the missing VLAN on the one affected switch from the pre-change configuration backup, then confirmed the phones registered again.
- Ran a read-only
state: gatheredpass (theios_vlansstate that reads the device's current VLANs back as data and changes nothing) on all 40 switches and compared each switch's VLANs with the source of truth, to confirm no other switch was missing a VLAN or held one the data did not list.
Impact arithmetic. 1 affected switch out of 40 is 1/40 = 2.5 percent of the fleet; 39 of 40 (97.5 percent) were never touched because the run stopped after the first batch. If I had run all 40 at once, every switch with the same data would have been hit (the story's switch counts, not measured figures).
What I changed afterwards.
- Guard: keep
serialin growing batches and add a failure threshold (max_fail_percentage, the percentage of hosts in a batch allowed to fail before Ansible aborts the play; at 0, one failure aborts it), so the first bad batch ends the run by itself instead of depending on me to notice and stop it (Tested on six hosts withserial: [1, 2, 3]: a failing host in the one-host first batch ended the play with or without this setting, because every host in that batch failed. The setting earns its place in the later, larger batches: when one of the two hosts in the second batch failed, the default let the third and later batches run, whilemax_fail_percentage: 0stopped the play before them.) - Pre-flight: a read-only check that runs before anything is changed; it reads each switch's current VLANs with
state: gathered, compares them with the intended data, and fails the pipeline if the run would delete any VLAN not named in the pull request (therenderedstate cannot do this, because it only turns your data into commands and never looks at the device). - Scope: prefer
mergedfor routine runs; useoverriddenonly through a reviewed change with a diff. - Review: the data files are validated for required entries in CI (continuous integration) before a merge.
- Blameless review (written without assigning fault to a person), shared with the team.
What I would do differently. Build the pre-flight diff before the first production use of overridden, not after this near-miss.
What interviewers look for
- A clear statement of your role and decisions, with timing of detection (how you found out) and of containment.
- Containment first, root cause second.
- Process fixes aimed at the pipeline (canary, batching, diff gate, scope), not a promise to be more careful.
Pitfalls
- Do not blame the tool or a colleague; describe the gap in the process.
- Do not invent precise outage durations you did not measure; use the ones you recorded.
- Include the contributing factor that made the blast radius larger or smaller (batching, backups, monitoring).
Your lab only approximates production, and some failures appear only under real traffic or vendor quirks. Before approving a network change, how would you validate it, and how does your approach differ for a syntax mistake, a design flaw and a platform bug?
Sample Answer
Direct answer
I approve a change when three questions have each been answered by the kind of evidence that can answer them: will the device accept it (syntax), will the network behave as intended (design), and will this hardware and software do what the documentation says (platform). A lab and static analysis can answer the first two well; the third can only be tested on production-like hardware, which in practice means a limited production canary with a rollback ready. Say clearly what remains unproven after each check, and size the rollout to that.
Three failure classes and the evidence for each
| Syntax mistake | Design flaw | Platform bug | |
|---|---|---|---|
| What it is | A line the device rejects or misreads: a typo, a reference to a prefix-list that does not exist | The configuration is valid and does what it says, but what it says is wrong for the network (a policy that causes a loop, a filter that blocks a needed path) | The device does not behave as documented under your feature, scale or software version |
| Where it is caught | Before the device: parsing and reference checks. On the device: a dry run that is not committed | In a model of the whole network, and by someone checking intent | Only by running on the same hardware and software |
| Evidence | Template renders, linter, Batfish parse and undefined-reference checks, device-side candidate validation | Batfish reachability and route-policy analysis, before and after comparison of expected routes, design review of failure scenarios | Release notes and known-bug search for the exact version, a canary on the same platform, post-change state comparison |
| What it cannot tell you | Whether the valid config is right | How the device implements it under load | Anything about devices you did not test |
Syntax mistakes: cheap and mechanical
- Render the configuration from templates and fail the build on a missing variable.
- Batfish (an open-source analyzer that reads configuration files without device access) answers named questions about a set of configs.
fileParseStatusreports, for each file, whether it parsed fully, partially or not at all, andparseWarninglists the warnings, such as lines it did not recognise.undefinedReferenceslists every place a config uses a named object that is never defined; each result row gives the file, the kind of object (for example a route-map), the missing name, and the line numbers where it is used. An empty result means every reference resolves. - Ask the device itself without committing. With NAPALM (a Python library giving one API across vendors),
load_merge_candidate()stages the candidate andcompare_config()returns the diff;discard_config()throws it away. Junos hascommit check, which validates a candidate (the proposed configuration, staged but not yet active) without activating it. NETCONF devices that advertise the:validatecapability (urn:ietf:params:netconf:capability:validate:1.1) can validate a candidate datastore. - Residual risk is low and the fix is cheap.
Design flaws: model, don't just lint
A syntactically perfect change can still be wrong. I check intent against behaviour:
- Run network-wide analysis: Batfish can compare two snapshots (the current configs and the proposed configs) with
differentialReachability, which returns example flows that succeed in one snapshot but not the other, so gained or lost connectivity shows up before anything is pushed (it returns examples rather than every flow, so rerun it after each fix).testRoutePoliciesfeeds sample routes through a route policy (the rules a router uses to accept, reject or modify routes) and shows what comes out. - Write down expected outcomes as testable statements ("branch A can still reach the data center on both uplinks; prefix X is not advertised to the internet") and check each before and after.
- Walk the failure scenarios: what if the link or peer the change relies on is down?
- Residual risk: the model uses its own idea of how a device behaves, and a lab topology that is smaller than production hides scale effects (route table size, number of ACL entries, convergence time).
Platform bugs: you cannot prove their absence in a lab
Lab and snapshot cannot catch problems that depend on production traffic, the vendor's real implementation or scale, which is exactly the limit stated in the question. So:
- Search the vendor's release notes and bug database for the exact software train and feature before approval.
- Run the change on the lowest-impact production device of the same model and version (a canary: a first, deliberately small deployment whose job is to reveal problems cheaply), not only in a lab with different hardware.
- Compare device state before and after (NAPALM getters such as
get_bgp_neighbors()andget_interfaces()return structured data) so unexpected changes show up even if no alarm fires. - Arm a rollback with a timer (commit-confirm: the device applies the change but reverts it by itself unless you confirm in time): Junos
commit confirmedrolls back automatically unless confirmed (default 10 minutes); IOS XE'sconfigure replace ... timewithconfigure confirmdoes the same job. - Widen in stages, and treat one clean canary as weak evidence: it excludes frequent faults, not rare ones.
Worked example
Change: raise BGP local preference for routes learned from a new transit provider so it becomes the preferred exit. BGP is the routing protocol that exchanges routes between networks. Local preference is a number a router attaches to a route inside your own network; when the same destination is reachable several ways, the route with the higher local preference wins (RFC 4271). A prefix-list is a named list of address ranges, and a route-map is a named set of match-and-set rules that here says "for routes matching this prefix-list, set local preference to a higher value".
- Syntax: the prefix-list name in the new route-map is misspelled.
undefinedReferencesflags it in CI. Nothing reaches a device. - Design: the route-map sets the higher local preference on every route from that provider, including prefixes of your own customers that should keep using their direct path. The differential analysis shows traffic to those prefixes moving. The fix is a narrower match, found before approval.
- Platform: with both fixes, the canary router shows its BGP process CPU staying elevated after the policy is applied, which the lab (with a far smaller route table) never showed. The vendor's release notes list a known issue for that version. The rollout stops at one router, the timer rolls the router back, and the approval waits for a software fix or a workaround.
Approval rule
I approve when the syntax checks are green, every stated expectation holds in the model, the vendor's known issues for the target version are read, and for any change touching routing or forwarding, a canary and a rollback are planned. A change with a higher blast radius (the amount of the network a mistake can hurt, as with an edge or core device) needs the canary even if every earlier check passed.
Your fleet spans several vendors and some old platforms that only offer a command-line session, while others expose structured APIs and transactions. The team wants one automation workflow for all of them. How do you design it?
Sample Answer
Direct answer
Put one automation workflow on top of a normalized, vendor-neutral data model of intent, and put the vendor differences below it in drivers. Every device then goes through the same stages (render, compare with live state, apply with a safety net, verify). A driver declares what its platform can do (structured API with transactions, or only a CLI session), and the workflow picks the safest available path for each device. The worked example below is a trunk port: its allowed VLANs and native VLAN, which Cisco IOS, Arista EOS and Juniper Junos express differently. A trunk port is a switch port that carries several VLANs (virtual LANs, separate layer-2 broadcast domains) over one cable by adding a numbered tag to each frame (the IEEE 802.1Q standard); the native VLAN is the one VLAN on that trunk whose frames travel untagged, so a frame without a tag is assumed to belong to it.
Layers
| Layer | Responsibility | Examples |
|---|---|---|
| Intent model | What should be true, with no vendor words | {port, mode: trunk, native_vlan: 99, allowed_vlans: [10,11,12,20,30]} |
| Renderers | Intent to the platform's syntax, one per platform family | Jinja2 templates for CLI platforms, structured payloads for API platforms |
| Drivers | Connect, read state, stage, commit, roll back | NAPALM drivers (eos, ios, iosxr, junos, nxos), Ansible connection types (network_cli, netconf, httpapi), Netmiko for CLI sessions, ncclient for NETCONF |
| Capability detection | Pick the safest path per device | NETCONF hello capabilities (the list of optional features a device announces when a session opens), driver feature table |
| Verification | Read back and compare with intent | Structured getters, parsed show output |
NAPALM (a Python library that gives one API over several vendors) already abstracts configuration methods across platforms: its documentation lists merge and replace as supported on every driver, while commit confirm is supported on only EOS, Junos and IOS. ncclient is a Python NETCONF client library. Netmiko drives plain CLI sessions for older platforms; its send_config_set has an error_pattern argument so a CLI error message can be detected instead of ignored.
Capability detection and fallbacks
Everything in this section serves one decision per device: use a transaction (stage, validate, commit) where the device offers one, otherwise build a safety net around CLI commands.
- Structured platforms with transactions. A NETCONF server announces its capabilities in its
<hello>message.urn:ietf:params:netconf:capability:candidate:1.0means a candidate datastore exists (stage changes, then commit);urn:ietf:params:netconf:capability:confirmed-commit:1.1means a commit can roll back by itself unless confirmed (the default confirm timeout is 600 seconds);urn:ietf:params:netconf:capability:validate:1.1lets a candidate be validated before commit. If the device offers candidate plus confirmed-commit, use them. - CLI-only platforms. No transaction exists, so build one: save a backup, arm a timed rollback if the platform offers one (for example IOS XE
configure replace ... timewithconfigure confirm), send only the computed difference, check for CLI errors, verify by reading back, and only then save to startup. - Everything else gets the CLI path with a stricter rollout (smaller phases). Each driver publishes a flag set such as
{candidate, confirmed_commit, structured_get}, so the workflow chooses by flag, not by vendor name.
I verified the capability URNs against RFC 6241 and executed the complete script below in a container; the connection to a device and the device behaviours are not executed here.
Worked example: trunk allowed VLANs and native VLAN
The same intent in three syntaxes (each from the vendor's documentation):
| Intent | Cisco IOS XE | Arista EOS | Juniper Junos (ELS switching syntax) |
|---|---|---|---|
| Mode | switchport mode trunk | switchport mode trunk | set interfaces ge-0/0/1 unit 0 family ethernet-switching interface-mode trunk |
| Allowed VLANs (VLAN: virtual LAN, a separate layer-2 broadcast domain) | switchport trunk allowed vlan add 12,30 and ... remove 40 | same keywords: switchport trunk allowed vlan with an edit action | set ... vlan members 12 and delete ... vlan members 40 |
| Native VLAN | switchport trunk native vlan 99 | switchport trunk native vlan 99 (EOS also has a setting to send the native VLAN tagged; confirm its exact syntax on your release) | set interfaces ge-0/0/1 native-vlan-id 99 (on the physical interface, not under unit 0) |
| Default if you configure nothing | All VLANs 1 to 4094 allowed; native VLAN 1 | All VLANs allowed; native VLAN 1 | The trunk lists its members explicitly (vlan members, which can also be all) |
The EOS cells are the least certain of the three, so confirm the keywords on your EOS release before relying on them.
In a structured model the same facts are OpenConfig YANG (YANG is a language for describing device data as a tree of named, typed values called leaves; OpenConfig is a vendor-neutral set of such models): the module openconfig-vlan adds a switched-vlan container with config under an Ethernet interface, holding the leaves interface-mode (ACCESS or TRUNK), native-vlan, access-vlan and the leaf-list trunk-vlans.
The semantic traps the abstraction must hide:
- Defaults differ. IOS and EOS carry every VLAN on a trunk until you restrict the list, so the intent model always states
allowed_vlansexplicitly and never relies on the default. - Native VLAN must match at both ends. If one end sends untagged frames meaning VLAN 99 and the other end reads untagged frames as VLAN 1, traffic lands in the wrong VLAN. The Cisco guide warns that if the native VLAN differs between the two ends of an 802.1Q trunk, spanning-tree loops might result. Verification therefore compares both ends of the link, not one port.
- List editing is not the same as list setting. Because the CLI edit actions are
addandremove, the renderer computes a difference against live state and emits only that. Setting a list wholesale can drop VLANs that are in use for a moment during the push. - Where the setting lives differs (Junos puts
native-vlan-idon the physical interface), so a renderer, not an engineer, knows the location.
Executed: the diff and render step
Intent: native 99, allowed 10,11,12,20,30. Live state on the Cisco-style port: native 1, allowed 10,11,20,40. Set subtraction gives add = 12 and 30 and remove = 40, and compress turns a set of VLAN numbers into the compact list a switch accepts (a run such as 10,11,12 becomes 10-12). The complete script, which I ran in a container (python:3-slim):
from dataclasses import dataclass
@dataclass
class TrunkState:
port: str
native: int
allowed: set
def compress(vlans):
"""Turn {10, 11, 12, 20} into '10-12,20': runs become ranges, commas separate the rest."""
ordered, parts, i = sorted(vlans), [], 0
while i < len(ordered):
j = i
while j + 1 < len(ordered) and ordered[j + 1] == ordered[j] + 1:
j += 1
parts.append(str(ordered[i]) if i == j else f"{ordered[i]}-{ordered[j]}")
i = j + 1
return ",".join(parts)
def plan_cli_cisco_style(want, have):
lines = [f"interface {want.port}", " switchport mode trunk"]
if want.native != have.native:
lines.append(f" switchport trunk native vlan {want.native}")
add, rem = want.allowed - have.allowed, have.allowed - want.allowed
if add:
lines.append(f" switchport trunk allowed vlan add {compress(add)}")
if rem:
lines.append(f" switchport trunk allowed vlan remove {compress(rem)}")
return lines
def pick_path(capabilities):
"""Choose the safest path from the URNs a NETCONF server sent in its <hello>."""
cand = "urn:ietf:params:netconf:capability:candidate:1.0" in capabilities
conf = any(c.startswith("urn:ietf:params:netconf:capability:confirmed-commit:") for c in capabilities)
return "candidate + confirmed-commit" if cand and conf else "candidate" if cand else "CLI with timed rollback"
want = TrunkState("GigabitEthernet1/0/48", native=99, allowed={10, 11, 12, 20, 30})
have = TrunkState("GigabitEthernet1/0/48", native=1, allowed={10, 11, 20, 40})
print("\n".join(plan_cli_cisco_style(want, have)))
print("--- already converged:")
print("\n".join(plan_cli_cisco_style(want, want)))
print("--- path with candidate + confirmed-commit advertised:")
print(pick_path({"urn:ietf:params:netconf:capability:candidate:1.0",
"urn:ietf:params:netconf:capability:confirmed-commit:1.1"}))
print("--- path with no capabilities:")
print(pick_path(set()))
It printed:
interface GigabitEthernet1/0/48
switchport mode trunk
switchport trunk native vlan 99
switchport trunk allowed vlan add 12,30
switchport trunk allowed vlan remove 40
--- already converged:
interface GigabitEthernet1/0/48
switchport mode trunk
--- path with candidate + confirmed-commit advertised:
candidate + confirmed-commit
--- path with no capabilities:
CLI with timed rollback
Read the output: the first block is the Cisco-style plan, with only the differences emitted. The second block shows the idempotent case: when live state already equals intent, only the interface line and the mode line remain and no VLAN changes are sent. The last two lines show pick_path choosing a transaction when the device announces both candidate and confirmed-commit, and the CLI path with a timed rollback when it announces nothing. A Junos renderer would follow the same diff, emitting set and delete lines for native-vlan-id and members.
Post-change validation
After the push the same driver reads state back and compares it with intent: allowed VLAN list, native VLAN, trunk operational state, and that the neighbor at the other end reports the same native VLAN. A mismatch triggers the rollback path chosen above. Only a passing check marks the device converged.
Trade-offs and pitfalls
- A lowest-common-denominator model hides useful vendor features. Give the model an escape hatch (a per-platform override block, reviewed and tested) rather than growing it for every feature.
- Normalizing output is as hard as normalizing input. Prefer structured getters (NAPALM, NETCONF, gNMI) over parsing text, and for CLI-only devices parse with a tested template library such as TextFSM, which Netmiko's
send_commandcan use throughuse_textfsm. - Do not claim identical safety everywhere: a CLI-only device has weaker rollback than a NETCONF one, so the rollout plan for it is smaller phases, not the same plan.
Your team opens a new branch office every month, and each site needs a router, switches, baseline ACLs and monitoring settings before local IT can use it. How would you automate provisioning from an empty rack to a usable site?
Sample Answer
Direct answer
Treat a new site as data, not as a project. The site is recorded once in a source of truth (SoT, the system that holds the intended state; NetBox is a common open-source choice), configurations are rendered from templates, a small day-0 bootstrap (the minimum configuration to make a new device reachable) gets the router reachable, and an Ansible pipeline applies the full baseline and checks the result before the site is marked live. My commitment: ship the router with zero-touch bootstrap, because local IT cannot be asked to configure it, and fall back to staging in a depot only if first boot at the site cannot reach the internet or a DHCP server.
The flow, step by step
- Record the site. The engineer (or a form) creates the site in the SoT: site code, address, circuit details, the device list with serial numbers, and a prefix allocated from IP address management (IPAM). NetBox's device status values include
planned,stagedandactive; a new site's devices start asplanned. - Generate the plan. CI (continuous integration, an automatic job on every change) renders each device's configuration from the template plus SoT data and runs checks (valid syntax, no overlapping addresses, every required template variable present). A reviewer approves the rendered output.
- Ship hardware to the site with its serial numbers already known.
- Day-0 bootstrap. The router needs just enough configuration to be reachable: management address, SSH keys, authentication server. On Cisco devices, Plug and Play (PnP, Cisco's zero-touch mechanism) finds its server through a short list of methods (the exact order and options depend on the platform and release, so check Cisco's PnP guide for yours): DHCP option 43 (a DHCP option is a labelled field in the DHCP reply; option 43 is the vendor-specific one, which RFC 2132 defines as an opaque block of vendor data, here used by the local DHCP server to tell the device where its PnP server is), a DNS lookup of the name
pnpserverin the domain DHCP returned, and the Plug and Play Connect cloud service as a fallback. The PnP server maps the serial number to the bootstrap configuration. Other vendors have their own mechanisms; the principle (serial number to a minimal config over DHCP) is shared. - Day-1 baseline. (Day 1 is the full standard configuration, as opposed to the day-0 minimum.) Once the device answers, Ansible applies the full rendered configuration: interfaces, baseline access control lists (ACLs), logging, NTP, SNMP or streaming-telemetry settings, and registration with the monitoring system through its API. Apply is idempotent (safe to repeat), so a failed run is simply re-run.
- Verify, then hand over. Automated checks read the device back: the serial number matches the SoT, each uplink is up, the neighbors found by LLDP (Link Layer Discovery Protocol, through which a device reports who is plugged into each port) match the documented cabling, the clock is synchronized, and a test probe from the monitoring system succeeds. Only when all pass does the pipeline set the SoT status to
activeand notify local IT. A failing check leaves the status atstagedand opens a ticket with the failing check named. - Keep it honest afterwards. The same pipeline runs nightly in compare-only mode to detect drift (a device differing from what the SoT says it should be).
Worked example: addressing is computed, not chosen by hand
Suppose all branches draw from 10.64.0.0/16, and each site gets one /24. A /16 contains 256 /24 networks, so this scheme supports 256 sites, which is more than enough for twelve new sites a year (computed with Python's ipaddress module). For site number 14 the pipeline allocates 10.64.14.0/24 and splits it into:
| Purpose | Subnet | Usable hosts |
|---|---|---|
| Users | 10.64.14.0/26 | 62 |
| Voice | 10.64.14.64/26 | 62 |
| Device management | 10.64.14.128/27 | 30 |
| Unallocated (growth) | 10.64.14.160 to 10.64.14.255 | not assigned |
The three allocated blocks use 160 of the 256 addresses, so the rest stays free. Because the SoT hands out the prefix, two sites can never receive the same one.
Trade-offs and what would change the recommendation
- At one site a month, building the zero-touch bootstrap is real engineering for a small number of uses. I still commit to it, because local IT only takes over after provisioning, so nobody at the site can type the day-0 configuration, and staging every router in a depot first adds a physical handling step to every opening. Build order: the rendering, the Ansible apply and the verification first, because those remove the inconsistency; the zero-touch bootstrap next. Until it is ready, stage routers in a depot, the fallback described below for a site with no network path.
- Bootstrap trust. A device enrolls based on its serial number, so a wrong serial in the SoT hands a site's configuration to the wrong box. Keep long-lived secrets out of the bootstrap config (use a temporary credential that day-1 rotates).
- If first boot has no network path (no DHCP, no internet at the new site), stage the device in a depot (a central warehouse or lab where it is configured and tested before shipping) with the day-0 file loaded, or ship a router with an LTE (mobile network) modem so it can reach the internet without local wiring.
- Hardware that fails after shipping is a replacement case: the SoT lets you re-render the same config for the replacement serial number.
Pitfalls
- Hand-edited devices after handover create drift. Make the nightly compare-only run raise a ticket, and fix the SoT or the device deliberately.
- Skipping the readback verification makes "automation succeeded" mean only "commands were accepted".
A cloud platform must give each of about 10,000 tenants its own isolated L3 network on shared Linux hosts. How would you build the host networking, handle addressing and tenant routing, isolate east-west traffic, and keep performance and monitoring manageable?
Sample Answer
Direct answer
Give each tenant its own VRF (virtual routing and forwarding instance: a separate routing table) on every host that runs one of its workloads, and carry tenant traffic between hosts inside VXLAN (Virtual Extensible LAN) tunnels with one VNI (VXLAN Network Identifier) per tenant. Use BGP (Border Gateway Protocol) EVPN (Ethernet VPN, a BGP address family that advertises tenant routes and MAC addresses so hosts do not have to flood and learn: the older switch method of sending unknown traffic everywhere and remembering who answers) as the control plane, run by FRR (an open-source routing daemon) on each host. Plain VLANs cannot do the job: a VLAN ID is a 12-bit number, which allows 4094 usable values (0 and 4095 are reserved), and RFC 7348 says that limit is inadequate for multi-tenant environments, while a VNI is a 24-bit value that allows up to 16 million segments. Ten thousand tenants is already 2.4 times the whole VLAN space.
The core of the design is the two lines of control: a VRF per tenant for routing isolation and a VNI per tenant for overlay isolation. The firewall, MTU, monitoring and offload sections below make that design safe and operable at 10,000 tenants.
The path a packet takes through the host
flowchart LR
W["Workload: network namespace or VM"] --> V["veth into the tenant VRF"]
V --> F["nftables forward hook: default drop"]
F --> B["Bridge + VXLAN device, VNI per tenant"]
B --> U["Underlay IP fabric, MTU 1550 or more"]
U --> R["Remote host: VNI, VRF, workload"]
C["FRR BGP EVPN"] -.-> B
| Layer | Linux object | Job |
|---|---|---|
| Workload | Network namespace (netns) or VM tap (a virtual network port for a VM), connected to the host by a veth (a pair of virtual Ethernet ports, like a cable: what goes in one end comes out the other) | A netns gives the workload its own interfaces, routes, firewall and sysctls. A VRF only gives one shared network stack a second routing table, so the workload gets the netns and the host side gets the VRF. |
| Tenant router | VRF device (ip link add vrf-A type vrf table 1001) | The VRF is the tenant's router on this host. Interfaces are attached with ip link set dev NAME master vrf-A, and their connected routes move into the VRF's table (kernel VRF documentation). |
| Tenant tunnel | Bridge enslaved to the VRF plus a VXLAN device with nolearning | Carries the tenant's L3 VNI. Learning is off because EVPN supplies the MAC and route entries instead of flood-and-learn. |
| Address of the host in the underlay | A VTEP (VXLAN tunnel endpoint) IP such as a loopback | Never visible to tenants and never in a tenant VRF. |
One VXLAN device per tenant would mean up to 10,000 devices on a host. FRR's EVPN documentation describes a single VXLAN device mode (ip link add vxlan0 type vxlan dstport 4789 local <VTEP IP> nolearning external vnifilter, then bridge vni add dev vxlan0 vni 100, plus a bridge VLAN id mapped to each VNI with tunnel_info id, plus a VLAN interface on the bridge enslaved to the VRF). In plain words: external makes one device handle many VNIs; vnifilter plus bridge vni add lists which VNIs this device accepts; and tunnel_info id maps a local bridge VLAN id to a global VNI, so the bridge can translate between the two. Each tenant on the host therefore borrows one local VLAN id, and a VLAN id is a 12-bit number with 4094 usable values. That VLAN id only has local meaning (it never leaves the host; the VNI is what travels), so one host can hold at most 4094 tenants this way, which is far above how many tenants have a workload on one host. Hosts create a tenant's VRF and VNI mapping when its first workload lands and remove it with the last one, so the object count per host follows resident tenants and not the 10,000 total.
Addressing and tenant routing
- Tenant address space. Each tenant picks its own CIDR (for example a /24 per subnet carved from the tenant's block). Overlap between tenants is legal because every VRF is a separate table; the platform's IP address management (IPAM) only enforces uniqueness inside one tenant. The demo below gives tenants A and B the same 10.0.1.0/24 and 10.0.2.0/24.
- Underlay. The fabric's own addresses (VTEP IPs, BGP peering) live in the default VRF from a range tenants never see.
- Routing between hosts. Use symmetric routing: the source host routes the packet inside the tenant VRF, encapsulates it with that VRF's L3 VNI, and the destination host decapsulates and routes again. Symmetric routing means both the sending and receiving host do a routing lookup in the tenant's VRF, with the tunnel in between carrying a single per-tenant VNI. Tenant prefixes travel as EVPN type-5 (IP prefix) routes (a route type that says "this address block is reachable behind this host"). In FRR the VRF-to-VNI mapping is a
vrf vrf1stanza containingvni 100;advertise-all-vnigoes insideaddress-family l2vpn evpnof the defaultrouter bgp <ASN>instance (ASN is the autonomous system number); andadvertise ipv4 unicastgoes insideaddress-family l2vpn evpnof the per-tenantrouter bgp <ASN> vrf vrf1instance. FRR documents that prefixes are not exported as type-5 routes until thatadvertiseline exists. - Route targets. FRR derives route targets automatically: import is the wildcard
*:VNIand export is(AS & 0xFFFF):VNI. A route exported for VNI 5001 is therefore imported only by the VRF that owns VNI 5001, which is the control-plane half of tenant isolation. A route target is a label attached to a BGP route that says which VRFs may import it. Worked example (illustrative AS 65001, VNI 5001):65001 & 0xFFFFis 65001, because 65001 is below 65536 and fits in two bytes, so the export label is65001:5001. Tenant A's VRF, which owns VNI 5001, imports anything labelled*:5001(any AS, VNI 5001), so it accepts this route. Tenant B's VRF owns VNI 5002 and imports*:5002, so it ignores the route. The*wildcard on the AS side is what lets every host import routes from any other host's AS, while the VNI half keeps tenants apart. - Leaving the cloud. Overlapping tenants cannot share one flat internet or on-premises route table, so each tenant's egress goes through a per-tenant NAT or gateway attachment, and any shared service (DNS, metadata, a managed database) is reached by an explicit, audited route import into that one VRF, never by leaking whole tables.
Isolating east-west traffic
Isolation comes from stacking independent layers so one mistake does not expose a tenant:
- Routing. A VRF has no route to another tenant's prefixes. The demo shows tenant A's workload receiving no reply for 10.9.9.2, a prefix that exists only in tenant B.
- Overlay. Frames carry a VNI, and a host only accepts VNIs it has configured.
- Host firewall. An nftables table (nftables is the Linux kernel's packet-filtering framework) hooked on
forwardwithpolicy drop, onect state established,related acceptline, and explicit per-tenant allow rules. Only routed traffic crosses this hook, so two workloads of one tenant that share a bridge need their own bridge-level filtering, and a test should confirm which path same-host traffic takes.
table inet tenant_fw {
chain forward {
type filter hook forward priority 0; policy drop;
ct state established,related accept
iifname "vrf-A" ip daddr 10.0.2.0/24 icmp type echo-request accept
counter comment "dropped by default"
}
}
Pitfall found while testing this on a Linux 7.0 kernel: a first version matched iifname "wv1A" (the host side of the workload veth) and silently dropped everything, including the allowed ping. For VRF-routed traffic the forward hook reported the VRF device (vrf-A) as the input interface. With the rule matching vrf-A, tenant A's ping passed and tenant B's identical ping was dropped. Check the interface name your own kernel presents before trusting a rule.
Performance
- MTU (maximum transmission unit). VXLAN adds about 50 bytes of outer Ethernet, IP, UDP and VXLAN headers (RFC 7348), so tenants that get a 1500-byte MTU need at least 1500 + 50 = 1550 on every underlay link, or a jumbo underlay. The demo sets 1550 on the underlay link, the kernel set the VXLAN device to 1500, and a 1500-byte packet with the do-not-fragment bit passed.
- ECMP spreading. RFC 7348 recommends deriving the outer UDP source port from a hash of the inner packet, which lets the fabric's equal-cost multipath (ECMP) spread tenant flows over all uplinks.
- Offload. Offload means letting the network card (NIC) do work the CPU would otherwise do. The kernel segmentation documentation lists UDP-tunnel segmentation types (such as SKB_GSO_UDP_TUNNEL; GSO is generic segmentation offload, splitting large packets into wire-size ones late or in hardware), so tunnel traffic can be segmented and checksummed in the NIC. Confirm each candidate NIC with
ethtool -kand compare iperf3 throughput with and without the tunnel before committing to hardware. - Control plane. Per-host cost follows resident tenants, as above. Watch the BGP table size with
show bgp l2vpn evpn summaryduring the scale test.
Keeping monitoring manageable
| Signal | Where it comes from | Why it matters |
|---|---|---|
| Per-tenant traffic | ip -s link show vrf-A (receive side only) plus the host-side veth counters of the tenant's workloads, summed per tenant by the collector | Bytes and packets per tenant in both directions, one series per tenant |
| Control plane | show bgp l2vpn evpn summary, show vrf vni, show evpn mac vni (FRR) | Session state, VNI-to-VRF mapping, learned MACs |
| Policy drops | nftables counters | Distinguishes a blocked flow from a broken path |
| Isolation | A scheduled negative probe: tenant A must NOT reach a canary only tenant B owns | Catches a leak that no throughput graph shows |
Four counters (receive and transmit bytes and packets) per tenant is 4 x 10,000 = 40,000 time series for the whole fleet. One measured caveat shapes where those counters come from: in the demo run the VRF device's TX counters stayed at 0 for forwarded traffic and only its RX counters moved, so the VRF device alone gives one direction. The other direction comes from the host-side veth of each workload, where RX is what the workload sent and TX is what was delivered to it. The VXLAN device is not a per-tenant source when one device carries every VNI. The collector reads the per-workload counters locally and exports only the per-tenant sum; per-workload series are exported for one tenant only while debugging.
Worked example, executed
The script builds two hosts as namespaces joined by an underlay link, gives tenants A (VNI 5001) and B (VNI 5002) overlapping addresses, and installs by hand the neighbor, forwarding-database and route entries that BGP EVPN would install. Save the script below as tenants.sh in the current directory and run it as root on Linux with iproute2 and iputils-ping, for example docker run --rm --privileged -v "$PWD":/w debian:stable-slim bash -c "apt-get update -qq && apt-get install -y -qq iproute2 iputils-ping && bash /w/tenants.sh".
#!/usr/bin/env bash
# Run as root in a Linux environment with iproute2 and iputils-ping (for example a privileged container).
# Two hypervisor "hosts" (h1, h2) as namespaces joined by an underlay link.
# Each tenant gets: a VRF, a bridge + VXLAN device carrying its L3 VNI, and one workload namespace per host.
set -euo pipefail
ip netns add h1; ip netns add h2
ip link add u1 type veth peer name u2
ip link set u1 netns h1; ip link set u2 netns h2
ip -n h1 addr add 172.16.0.1/30 dev u1; ip -n h2 addr add 172.16.0.2/30 dev u2
ip -n h1 link set u1 mtu 1550 up; ip -n h2 link set u2 mtu 1550 up # 1500 inner + 50 VXLAN overhead
# tenant_setup <name> <vni> <vrf-table> <router-mac-h1> <router-mac-h2> <transit-net>
tenant_setup() {
local t=$1 vni=$2 tbl=$3 m1=$4 m2=$5 n=$6
for h in 1 2; do
# r is the other host (h=1 gives 2, h=2 gives 1); ${!mac} reads the variable whose NAME is stored in mac (m1 or m2)
local r=$((3 - h)); local mac=m$h rmac=m$r
ip -n h$h link add vrf-$t type vrf table $tbl; ip -n h$h link set vrf-$t up
ip -n h$h link add br-$t type bridge
ip -n h$h link set br-$t address ${!mac} master vrf-$t up
ip -n h$h link add vx-$t type vxlan id $vni dstport 4789 local 172.16.0.$h nolearning
ip -n h$h link set vx-$t master br-$t up
ip -n h$h addr add 10.255.$n.$h/32 dev br-$t
# STAND-IN FOR EVPN (two lines): a type-2/type-5 route would tell this host the remote router's MAC and VTEP.
# neigh add ... nud permanent: a fixed ARP entry (remote router IP -> remote router MAC) that never ages out.
# bridge fdb add ... dst: forwarding entry saying frames for that MAC go inside VXLAN to the remote VTEP IP.
ip -n h$h neigh add 10.255.$n.$r lladdr ${!rmac} dev br-$t nud permanent
bridge -n h$h fdb add ${!rmac} dev vx-$t dst 172.16.0.$r self static
# tenant workload: veth into the VRF, namespace w<h>-<t>
ip netns add w$h-$t
ip link add wv$h$t type veth peer name wp$h$t
ip link set wp$h$t netns w$h-$t; ip link set wv$h$t netns h$h
ip -n h$h link set wv$h$t master vrf-$t up
ip -n h$h addr add 10.0.$h.1/24 dev wv$h$t
ip -n w$h-$t addr add 10.0.$h.2/24 dev wp$h$t; ip -n w$h-$t link set wp$h$t up; ip -n w$h-$t link set lo up
ip -n w$h-$t route add default via 10.0.$h.1
# STAND-IN FOR EVPN type-5: the other host's subnet, reachable via its router IP, installed in this tenant's VRF only.
# onlink: accept the next hop as directly reachable on br-<tenant> even though no address is configured on that subnet.
ip -n h$h route add 10.0.$r.0/24 vrf vrf-$t via 10.255.$n.$r dev br-$t onlink
done
}
tenant_setup A 5001 1001 02:00:00:00:0a:01 02:00:00:00:0a:02 1
tenant_setup B 5002 1002 02:00:00:00:0b:01 02:00:00:00:0b:02 2
for h in h1 h2; do ip netns exec $h bash -c 'echo 1 > /proc/sys/net/ipv4/ip_forward'; done
# tenant B alone owns a second subnet on h2
ip -n h2 addr add 10.9.9.1/24 dev wv2B
ip -n w2-B addr add 10.9.9.2/24 dev wp2B
ip -n h1 route add 10.9.9.0/24 vrf vrf-B via 10.255.2.2 dev br-B onlink
echo "--- A w1 -> A w2 (same addresses exist in tenant B)"
ip netns exec w1-A ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> B w2"
ip netns exec w1-B ping -c2 -W1 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)"
ip netns exec w1-B ping -c1 -W1 10.9.9.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)"
ip netns exec w1-A ping -c1 -W1 10.9.9.2 2>&1 | grep -o '[0-9]* packets transmitted, [0-9]* received' || true
echo "--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)"
ip netns exec w1-A ping -c1 -W1 -M do -s 1472 10.0.2.2 | grep -o '[0-9]* packets transmitted, [0-9]* received'
echo "--- MTUs on h1"
for dev in u1 vx-A; do ip -n h1 -o link show $dev | awk '{sub(/@.*/, "", $2); sub(/:$/, "", $2); print $2, $4, $5}'; done
echo "--- tenant VRF route tables on h1"
echo "vrf-A:"; ip -n h1 route show vrf vrf-A
echo "vrf-B:"; ip -n h1 route show vrf vrf-B
Output:
--- A w1 -> A w2 (same addresses exist in tenant B)
2 packets transmitted, 2 received
--- B w1 -> B w2
2 packets transmitted, 2 received
--- B w1 -> 10.9.9.2 (prefix that exists only in tenant B)
1 packets transmitted, 1 received
--- A w1 -> 10.9.9.2 (no such route in tenant A's VRF)
1 packets transmitted, 0 received
--- full-size inner packet, DF set (1472 + 28 = 1500 bytes)
1 packets transmitted, 1 received
--- MTUs on h1
u1 mtu 1550
vx-A mtu 1500
--- tenant VRF route tables on h1
vrf-A:
10.0.1.0/24 dev wv1A proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.1.2 dev br-A onlink
vrf-B:
10.0.1.0/24 dev wv1B proto kernel scope link src 10.0.1.1
10.0.2.0/24 via 10.255.2.2 dev br-B onlink
10.9.9.0/24 via 10.255.2.2 dev br-B onlink
How to read the script: tenant_setup runs once per tenant and builds, on each of the two hosts, a VRF, a bridge and VXLAN device carrying that tenant's VNI, and a workload namespace connected by a veth. The shell trick ${!mac} means "the value of the variable whose name is stored in mac", so on host 1 it reads m1, the argument holding host 1's router MAC; r=$((3 - h)) is arithmetic that gives the other host's number. The only steps that stand in for BGP EVPN are the neigh add ... nud permanent line (a fixed ARP entry for the remote router, never expiring), the bridge fdb add line (which tunnel destination to use for that MAC) and the final route add ... onlink line (the remote subnet, inside this tenant's VRF only; onlink accepts the next hop as directly reachable). In a real deployment EVPN advertises those three facts by itself. Everything else is the same.
Reading it: both tenants reach their own 10.0.2.2 through the same address plan, tenant A has no route to the prefix owned by tenant B, and each VRF holds only its own routes. The static entries stand in for EVPN, and the single-VXLAN-device and FRR lines above come from the FRR documentation and were not run here.
Trade-offs and pitfalls
- EVPN on the host versus a central controller. A controller-driven overlay such as OVN (Open Virtual Network, an Open vSwitch based controller that uses Geneve tunnels, a tunnel format similar in purpose to VXLAN) gives distributed firewalling and L2 features with one control point. EVPN on the host keeps the fabric on standard BGP that a network team already operates and avoids a controller as a single dependency. Recommend EVPN when the team is network-led and the model is routed (L3) per tenant; switch to the controller model if tenants need stretched L2 subnets and live migration with identical addresses.
- VRF is not a security boundary by itself. It separates routing only. A host compromise, a kernel bug or a mistaken route import defeats it, which is why the firewall and the negative probe exist.
- MTU black holes. If ICMP "fragmentation needed" is blocked in the underlay, small pings work while large transfers stall.
- Overlapping prefixes at every shared boundary (NAT, peering, shared services) need per-tenant handling, or the first two tenants who overlap collide there.
That is every published Network Automation and Software-Defined Networking question for Systems Engineer so far. Browse the other topics in this category, or practice this one interactively.