Network Automation and Software-Defined Networking Questions
Programmatic and automated network operation: network automation tooling (Ansible network modules, Netmiko, NAPALM, Nornir, Jinja2 templating, ZTP), choosing between CLI-scraping libraries and model-driven programmability (NETCONF, RESTCONF, gNMI, YANG) across multi-vendor fleets, idempotent and declarative-versus-imperative design, a network source of truth (NetBox) and intended-versus-running configuration drift reconciliation, Git-based versioning and review of device configuration and automation code, pre- and post-change validation (for example Batfish) including lab or emulated test environments, CI/CD and staged, canary or rollback-safe rollout of configuration to large device fleets with concurrency control and credential and secret handling for automation, automated site and leaf-switch provisioning, and event-driven remediation from streaming telemetry. Also software-defined networking: control and data-plane separation, controller architecture and state consistency, controller-managed overlays and tenant network virtualization, SD-WAN controller orchestration and programmable data planes (P4). Boundary: general scripting and Terraform practice, shell craft, IAM, incident response, protocol design, data-centre fabric and WAN topology design (including choosing EVPN-VXLAN or SD-WAN), and network fault diagnosis are covered elsewhere.
Tell me about a time an automated network change caused an unexpected problem in production, or nearly did. What was your role, how did you contain the impact, and what did you change in the automation or rollout process afterward?
Sample Answer
Direct answer
A strong answer is one real incident told in order: the change, the signal that something was wrong, what you did in the first minutes to limit the damage, how you restored service, and the specific changes to the automation and rollout that now prevent a repeat. The example below is an illustrative story with arithmetic you can check; use your own facts.
Story skeleton
Situation. I owned a playbook that kept the VLAN list on 40 access switches in line with our source of truth (the repository that holds the intended state). A cleanup moved VLAN data between two files and I ran the playbook with ios_vlans in state: overridden, which replaces all VLANs on the device with the configuration supplied. The moved file was missing one site's voice VLAN (the separate VLAN, a virtual LAN that splits one physical switch into isolated networks, that carries IP phone traffic; without it on a switch, the phones plugged into it cannot register with the phone system).
What happened. The first batch was a single canary switch (one device changed first, as an early warning, before the rest). The task itself finished green, but a check I had added after an earlier near-miss, a comparison of the VLAN count before and after the change, flagged that the canary had lost a VLAN. Because I had rolled out in batches (1, 2, 3, ...), the other 39 switches were untouched. serial is the Ansible play keyword that sets how many hosts are changed at a time; with it set to a list of increasing batch sizes, a failure scopes to the batch, not the whole host list (Ansible documentation).
Containment (my role).
- Stopped the run, and told the on-call and the site owner what had changed and on which switch.
- Restored the missing VLAN on the one affected switch from the pre-change configuration backup, then confirmed the phones registered again.
- Ran a read-only
state: gatheredpass (theios_vlansstate that reads the device's current VLANs back as data and changes nothing) on all 40 switches and compared each switch's VLANs with the source of truth, to confirm no other switch was missing a VLAN or held one the data did not list.
Impact arithmetic. 1 affected switch out of 40 is 1/40 = 2.5 percent of the fleet; 39 of 40 (97.5 percent) were never touched because the run stopped after the first batch. If I had run all 40 at once, every switch with the same data would have been hit (the story's switch counts, not measured figures).
What I changed afterwards.
- Guard: keep
serialin growing batches and add a failure threshold (max_fail_percentage, the percentage of hosts in a batch allowed to fail before Ansible aborts the play; at 0, one failure aborts it), so the first bad batch ends the run by itself instead of depending on me to notice and stop it (Tested on six hosts withserial: [1, 2, 3]: a failing host in the one-host first batch ended the play with or without this setting, because every host in that batch failed. The setting earns its place in the later, larger batches: when one of the two hosts in the second batch failed, the default let the third and later batches run, whilemax_fail_percentage: 0stopped the play before them.) - Pre-flight: a read-only check that runs before anything is changed; it reads each switch's current VLANs with
state: gathered, compares them with the intended data, and fails the pipeline if the run would delete any VLAN not named in the pull request (therenderedstate cannot do this, because it only turns your data into commands and never looks at the device). - Scope: prefer
mergedfor routine runs; useoverriddenonly through a reviewed change with a diff. - Review: the data files are validated for required entries in CI (continuous integration) before a merge.
- Blameless review (written without assigning fault to a person), shared with the team.
What I would do differently. Build the pre-flight diff before the first production use of overridden, not after this near-miss.
What interviewers look for
- A clear statement of your role and decisions, with timing of detection (how you found out) and of containment.
- Containment first, root cause second.
- Process fixes aimed at the pipeline (canary, batching, diff gate, scope), not a promise to be more careful.
Pitfalls
- Do not blame the tool or a colleague; describe the gap in the process.
- Do not invent precise outage durations you did not measure; use the ones you recorded.
- Include the contributing factor that made the blast radius larger or smaller (batching, backups, monitoring).
You must roll out a complex BGP route-map change across edge routers with minimal risk to routing stability. Design the canary plan: how you pick the canaries, what you watch, when you stop, and how you roll back.
Sample Answer
Direct answer
Roll the route-map change (a route-map is an ordered list of rules that decide, route by route, whether to accept or reject it and how to alter it) out in widening phases, starting with the router whose failure hurts least, and decide the stop conditions and the rollback before the first command. Each phase must pass the same checks: no BGP (Border Gateway Protocol) session resets, prefix counts that match what the change was predicted to do, and healthy traffic and probes. Rollback means redeploying the previous reviewed version of the route-map from Git and soft-resetting the affected sessions, with a device-side timer as a safety net. The commands below are Cisco IOS XE; other platforms have equivalents.
What a route-map is, with an example
BGP (Border Gateway Protocol) is how routers exchange reachability: each router tells its peers (the neighbouring routers it exchanges routes with) which prefixes it can reach, where a prefix is a block of addresses such as 198.51.100.0/24. A route-map is the filter-and-edit step applied to those advertisements. Each numbered entry says: if the route matches this condition, permit or deny it and optionally set attributes. Cisco documents that a route that matches no entry is ignored: not accepted on inbound, not advertised on outbound. An illustrative inbound example (names and addresses are made up):
ip prefix-list CUSTOMER seq 5 permit 198.51.100.0/24
!
route-map FROM-PEER permit 10
match ip address prefix-list CUSTOMER
set local-preference 200
route-map FROM-PEER permit 20
!
router bgp 64496
neighbor 203.0.113.1 route-map FROM-PEER in
Read it top to bottom: a route for 198.51.100.0/24 matches entry 10 and is accepted with local-preference 200 (higher is preferred, so the router will choose this path over other paths for that prefix); any other route skips entry 10, hits entry 20, which has no match line and so matches everything, and is accepted unchanged. If entry 20 were missing, every other route from that peer would be dropped, which is exactly the kind of mistake a canary exists to catch. in means the policy runs on routes received from the peer; out would run it on routes sent to the peer.
Choosing the canaries
A canary is the first device that gets a change so that a failure is small and informative. For edge routers I pick for impact and for coverage:
- Low impact first: a router at a site with redundant paths, off the busiest transit or customer links, so traffic can shift away if it misbehaves.
- Coverage of the differences: at least one router per hardware model and software version in the fleet, and one that exercises each kind of peer the policy touches (a transit provider sells you access to the whole internet, an internet exchange peer is a network you meet at a shared exchange point, a customer buys access from you), because the policy may behave differently by peer type.
- Out-of-band access: a canary whose management path does not depend on the policy being changed. A bad inbound policy can remove the route you use to reach the router.
For a fleet of 40 edge routers, a sensible plan is cumulative phases of 1, 4, 20 and 40 routers, which is incremental groups of 1, 3, 16 and 20 (1 + 3 + 16 + 20 = 40). Between phases there is a bake period (a deliberate wait, with monitoring on, before widening) long enough to include a peak-traffic period and one routine BGP table churn, since many policy faults only show under load.
Before the change: capture the baseline
On each router in the phase, save to files for later comparison:
show ip bgp summary: per-peer state, theUp/Downtimer and theState/PfxRcdcolumn. An illustrative capture (the columns follow the layout in FRR's documentation; the values are made up):
Neighbor V AS MsgRcvd MsgSent TblVer InQ OutQ Up/Down State/PfxRcd
203.0.113.1 4 64500 10234 9871 42 0 0 3d04h 812
203.0.113.9 4 64501 5120 5001 42 0 0 00:02:10 0
198.51.100.9 4 64502 0 0 0 0 0 00:00:41 Active
How to read it: Neighbor is the peer, AS its autonomous system (network) number, and Up/Down how long the session has been in its current state. State/PfxRcd shows the prefix count received from the peer once the session is Established (up and exchanging routes), and the state name (here Active, meaning it is still trying to connect) while it is not. Row 1 is healthy: up for three days and 812 prefixes received. Row 2 reads 00:02:10, so that session went down and came back two minutes ago, which in a canary phase is a stop signal. Row 3 is not Established at all.
show ip bgp neighbors <peer> advertised-routes: what you send each peer. For routes received, usereceived-routes(needsneighbor <peer> soft-reconfiguration inbound, otherwise IOS answers that inbound soft reconfiguration is not enabled) orroutesfor accepted ones only. Soft reconfiguration stores every received update unmodified and the Cisco guide calls that approach memory-intensive, so use it only on the canary peers, or rely on route refresh (RFC 2918) where the peer supports it.show ip bgp neighborsshows "Received route refresh capability from peer" when it does.- Interface counters and latency probes on the links to those peers.
Also predict the effect offline. Take the old and new route-map and the set of routes you expect each peer to see, and work out which prefixes should change attribute, be filtered or be unaffected. Tools such as Batfish (an open-source network configuration analyzer that works from configuration files without touching devices) can evaluate route policies against sample routes, but the point holds with any method: you must know the expected change before you look at the actual change.
Applying it to each phase
- Take an archive of the running config (
archive config) so there is a known restore point. - Push the new route-map through the normal pipeline. A route-map edit does not affect sessions that already exist until they are soft-reset: the Cisco soft-reset guide states that on a policy change the BGP session must be soft-cleared for the new policy to take effect. Use
clear ip bgp <peer> soft outfor outbound policy andclear ip bgp <peer> soft infor inbound. Thesoftkeyword re-evaluates policy without tearing the session down (a soft reset), unlike a hard reset, which drops the session and every route learned over it. Do one peer first, not all peers. - Compare the same captures with the baseline.
What I watch and when I stop
| Signal | Stop and roll back if |
|---|---|
Session state and Up/Down timer | Any session in the phase resets or leaves Established |
State/PfxRcd and advertised-route counts | The result differs from the offline prediction for that peer (for example, a peer that should see no change now receives fewer prefixes) |
| Advertised route content | Any prefix that should be filtered is sent, or any intended prefix is missing |
| Traffic and probes | Loss or latency worsens, or traffic shifts onto a link the change should not have touched |
| Router health | CPU on the BGP process or memory climbs and stays up after the soft reset settles |
The stop rule is written down in advance: one violated row halts the rollout, and nobody argues about it while the change is live.
Rollback
- Normal rollback: redeploy the previous Git revision of the route-map through the same pipeline, then soft-reset the same peers. Using the rollout path for rollback means it is tested.
- Safety net for lost management access: IOS XE can restore an archived config by itself:
configure replace <file> time <minutes>applies a configuration and reverts unless you runconfigure confirmin time. Note that it replaces the whole running configuration with the archived file, so nothing else may be changed on that router during the window. - A failure on the canary also stops later phases automatically. Resume only after the cause is understood.
Pitfalls
- A canary that passes proves little if it has no peer of the type the policy affects, or if it carried no traffic. One clean router rules out common faults, not rare ones.
- Checking only that sessions are up misses a policy that filters correctly formed prefixes. Compare prefix content, not just counts.
- Changing the route-map and other things in the same window makes the cause of a failure unprovable.
Design a closed-loop system where streaming telemetry from network devices triggers automated remediation for problems such as interface flaps or BGP session loss. Describe the components and how you decide that an event is real and a fix is safe to run.
Sample Answer
Direct answer
Build it as a pipeline with a decision gate (a check that must pass before the next step) and a brake (a limit that stops the system acting too often) in front of every action: devices stream state over gNMI (the gRPC network management interface) into a bus (a message queue that decouples producers from consumers); a normalizer and correlator (a component that groups events sharing one cause) turn raw updates into one incident per cause; a policy engine decides whether the incident is real and whether a known-safe runbook (a written, tested fix procedure for one kind of problem) applies; an executor runs it under locks and rate limits; and a verifier confirms the symptom cleared, otherwise it rolls back and pages a human. Automatic action is allowed only for event classes with a tested runbook, never for "anything that looks bad".
Components
- Telemetry collection. A gNMI subscription in STREAM mode delivers long-lived updates. Only two values of each state matter for this design. Interface
oper-status(a leaf under/interfaces/interface/state) has several values, but the loop cares about DOWN and UP; the BGPsession-stateleaf walks through IDLE, CONNECT, ACTIVE, OPENSENT and OPENCONFIRM while a session is coming up, and the loop cares about ESTABLISHED (healthy) versus anything else. For state that changes rarely like these, use ON_CHANGE (the device sends an update only when the value changes, so a flap shows up as a burst of updates), optionally with aheartbeat_intervalso the value is re-sent periodically and a silent stream is noticed. For counters such as errors, which change constantly, use SAMPLE with asample_interval(the device sends the current value every interval, for example every 10 seconds). Thesync_responsemessage marks that the initial state of every subscribed path has been sent, so the pipeline knows its picture is complete. In time order: the subscription starts, the device sends the current value of every path,sync_responsemarks the end of that first dump, then ON_CHANGE updates arrive as things happen, SAMPLE updates arrive every interval, and a heartbeat re-sends an unchanged value so silence can be told apart from stability. If a platform does not offer ON_CHANGE for a leaf, subscribe with SAMPLE and accept the slower detection. - Stream bus and normalizer. Events are stamped with device, path and time and mapped to a common schema (device, object, kind, value).
- Correlator. Groups events that share a cause. A BGP session loss within seconds of an interface going down on the same device is a child of that interface incident; it does not trigger its own remediation.
- Policy engine. Checks the gates below and picks a runbook from an allowlist keyed by event class.
- Executor. Runs the runbook (for example an Ansible job) with a per-device lock, a rate budget, and a pre-check of current state.
- Verifier and audit. Re-reads telemetry after the action, closes the incident or reverts and pages, and logs the event, the decision and the evidence.
Deciding that an event is real
- Persistence: the condition holds past a debounce window (a short wait during which brief blips are ignored), or it repeats (three DOWN updates in 60 seconds is a flap).
- Corroboration: a second independent signal (the neighbor side also reports the link down, or the counters agree). A telemetry stream that goes quiet is a collector problem, not a link problem; the heartbeat separates the two.
- Not caused by us or by a change: the device is not in a maintenance window and the event is not within the hold-down (a quiet period after our own action during which the resulting events are ignored) of our own previous action.
Deciding that a fix is safe
- The event class has a runbook that has been tested and an owner.
- Blast radius (how much of the network a mistake can hurt) is bounded: one device or link per incident, and a global budget (here 3 actions per hour) after which the system stops acting and pages.
- Preconditions pass (for an interface shut: an alternate path with enough headroom to take the traffic), and the runbook is idempotent (running it twice leaves the same result as running it once).
- A kill switch (a manual off control) stops all automatic action instantly, and a circuit breaker (an automatic off control, like an electrical one) stops it when two consecutive actions fail verification.
- Actions with no known-safe fix, such as a lone BGP session loss with no interface event, become a ticket with the evidence attached.
Worked example
A tiny simulator of the correlation and gating logic, with fixed events (times in seconds). Each event is (time, device, kind, object, new value), where kind if is an interface and bgp is a session. Three constants set the rules: FLAP_WINDOW is how many seconds of history count toward a flap (60), FLAP_MIN is how many DOWN events inside that window make it a flap (3), and CORRELATE_S is how close in time a BGP drop must be to an interface DOWN on the same device to be treated as its consequence (10 seconds):
from collections import defaultdict
events = [
(0, "leaf1", "if", "Ethernet1", "DOWN"),
(2, "leaf1", "bgp", "10.1.0.0", "IDLE"), # session rides that interface
(9, "leaf1", "if", "Ethernet1", "UP"),
(14, "leaf1", "if", "Ethernet1", "DOWN"),
(21, "leaf1", "if", "Ethernet1", "UP"),
(30, "leaf1", "if", "Ethernet1", "DOWN"),
(31, "leaf1", "bgp", "10.1.0.0", "IDLE"),
(60, "leaf2", "bgp", "10.2.0.0", "IDLE"), # lone BGP loss, no interface event
]
FLAP_WINDOW, FLAP_MIN, CORRELATE_S = 60, 3, 10
MAX_ACTIONS_PER_HOUR = 3
downs = defaultdict(list); incidents = []
for t, dev, kind, key, val in events:
if kind == "if" and val == "DOWN":
downs[(dev, key)].append(t)
recent = [x for x in downs[(dev, key)] if t - x <= FLAP_WINDOW]
if len(recent) >= FLAP_MIN and not any(i["root"] == (dev, key) for i in incidents):
incidents.append({"root": (dev, key), "t": t, "children": []})
if kind == "bgp" and val != "ESTABLISHED":
parent = next((i for i in incidents if i["root"][0] == dev and t - i["t"] <= CORRELATE_S), None)
near_if = any(d == dev and 0 <= t - x <= CORRELATE_S for (d, _), xs in downs.items() for x in xs)
if parent: parent["children"].append(key)
elif not near_if: incidents.append({"root": (dev, key), "t": t, "children": []})
actions = 0
for i in incidents:
kind = "interface flap" if i["root"][1].startswith("Ethernet") else "bgp session loss"
if actions >= MAX_ACTIONS_PER_HOUR: verdict = "BUDGET EXHAUSTED: page a human"
elif kind == "interface flap": verdict = "ELIGIBLE: shut interface, then verify"; actions += 1
else: verdict = "NOT ELIGIBLE: no known-safe fix, open ticket with context"
print(i["root"], kind, "children", i["children"], "->", verdict)
Output:
('leaf1', 'Ethernet1') interface flap children ['10.1.0.0'] -> ELIGIBLE: shut interface, then verify
('leaf2', '10.2.0.0') bgp session loss children [] -> NOT ELIGIBLE: no known-safe fix, open ticket with context
Reading the BGP branch line by line: parent looks for an already-open incident on the same device that began at most 10 seconds ago; near_if is true if any interface on this device went DOWN within the last 10 seconds (the nested loop walks every recorded DOWN time x for every interface); if a parent exists the drop is recorded as its child, and only if there is neither a parent nor a nearby interface DOWN does the drop open its own incident.
Eight events become two incidents. The third DOWN at t=30 completes three DOWN events inside 60 seconds (at 0, 14 and 30), which opens the interface incident; the BGP drop at t=31 is attached to it as a child rather than acting separately. The BGP drop at t=2 happened before the incident existed but within 10 seconds of an interface DOWN, so it was not opened as its own incident either; it is the same event the flap explains. The lone loss on leaf2 has no interface cause, so the only safe automatic outcome is a ticket.
Pitfalls
- Remediation that triggers its own alert: the shut interface generates DOWN events. Mark the target under remediation first and suppress its events for a hold-down period.
- Acting on a stale picture: if the stream lagged, re-read the current state just before the action.
- Acting on the symptom the telemetry shows while the cause is upstream, for example shutting leaf uplinks because a spine is failing. Gate on correlation across devices and cap the blast radius, so a fabric-wide event stops the automation and pages a human.
Write a reusable Jinja2 template that renders an interface configuration block for Cisco IOS from structured data (interface name, description, IP address and mask, MTU), and show how you would render it from an Ansible task.
Sample Answer
Direct answer
A Jinja2 template (a text file with placeholders and loops that a templating engine fills from data) turns a list of interface dictionaries into IOS configuration. In Ansible you render it with the ansible.builtin.template lookup and hand the result to cisco.ios.ios_config through its content parameter. Run with --check -v first: in check mode the module sends nothing and returns, in its commands result, the lines it would send; then run for real.
The template
File templates/interfaces.j2:
{% for intf in interfaces %}
interface {{ intf.name }}
description {{ intf.description }}
ip address {{ intf.ip }} {{ intf.mask }}
ip mtu {{ intf.mtu | default(1500) }}
no shutdown
!
{% endfor %}
interfacesis a list, so one template serves a router with two interfaces or twenty.- The names
name,description,ip,maskare required: if one is missing, Ansible stops with an undefined-variable error instead of pushing a half-written line.mtuis optional and falls back to 1500 through thedefaultfilter. ip mtusets the IPv4 MTU (maximum transmission unit, the largest IP packet in bytes) on a routed interface. The separate interface commandmtusets the port MTU and its allowed range depends on the platform; on Catalyst 9000 switches the Cisco guide showsmtuper physical port or port-channel. Choose the command your platform and design call for. This template usesip mtu.- The mask is written in dotted form (
255.255.255.252), which is what theip addresscommand takes. Keeping the mask in the data avoids converting prefix lengths inside the template.
The data
File inventory/host_vars/edge1.yml (Ansible loads host_vars/<hostname>.yml for that host automatically):
interfaces:
- name: GigabitEthernet0/0/0
description: "UPLINK to core-sw1 Gi1/0/48"
ip: 10.10.0.2
mask: 255.255.255.252
mtu: 1500
- name: GigabitEthernet0/0/1
description: "LAN users"
ip: 10.10.1.1
mask: 255.255.255.0
The Ansible task
- name: Apply interface configuration
hosts: edge_routers
gather_facts: false
vars:
ansible_connection: ansible.netcommon.network_cli
ansible_network_os: cisco.ios.ios
tasks:
- name: Push rendered interface stanzas
cisco.ios.ios_config:
content: "{{ lookup('ansible.builtin.template', 'interfaces.j2') }}"
backup: true
save_when: modified
Why these choices:
ansible.netcommon.network_cli(CLI over SSH) andansible_network_osare what Ansible's network modules need to know the platform. Network modules run on the control node, not on the device, because most devices cannot run Python.contenttakes already-rendered configuration text. The older way,src: interfaces.j2, still works today, but the module's documentation sayssrcwill stop processing Jinja2 in January 2028 and recommends exactly thelookup('ansible.builtin.template', ...)pluscontentform used here.ios_configcompares the lines with the running configuration and sends only the lines it cannot find there, so lines that appear in the running configuration exactly as written (including indentation) are not sent again. Lines IOS does not normally display are the exception:no shutdownon an enabled interface and anip mtuequal to the default usually do not appear inshow running-config, so the module cannot find them and will send them and report the task as changed on every run. The module itself warns that the input lines should look like the running configuration. For a clean second run, render those two lines only when the data asks for a non-default value, or accept a changed result each time.backup: truesaves the running config to the control node first.save_when: modifiedcopies running to startup only if the run changed something.- The module supports check mode, so
ansible-playbook -i inventory/hosts.yml push.yml --check -vprints thecommandslist, the lines that would be sent, without sending them.--diffalone adds nothing for this module: in the collection's source it builds a diff only whendiff_againstis set.
Executed
I executed everything that does not need a router, inside a python:3.12-slim container with ansible-core: the template rendered through both ansible.builtin.template and the lookup, and this was the output for the data above:
interface GigabitEthernet0/0/0
description UPLINK to core-sw1 Gi1/0/48
ip address 10.10.0.2 255.255.255.252
ip mtu 1500
no shutdown
!
interface GigabitEthernet0/0/1
description LAN users
ip address 10.10.1.1 255.255.255.0
ip mtu 1500
no shutdown
!
The second interface had no mtu in its data, so it got the default of 1500. The call that needs a live device is ios_config itself, which was not run; its parameters were checked against the module documentation and the installed collection's argument specification (it accepts content and supports check mode).
Trade-offs and pitfalls
- A template only renders what it is given. Validate the data (a valid address, a mask that matches, a unique IP per device) in CI before it reaches the template, or the template will faithfully render a mistake.
- Line-by-line merging never removes a stale line. If an old
ip helper-addressshould disappear, the template has to sayno ...for it, or the push has to use a replace-style workflow. This is the limit of an additiveios_configpush. no shutdownin the template will bring up an interface that someone shut down on purpose. Put the admin state in the data (enabled: true) if that matters.- Quote description strings in YAML. A colon or
#in an unquoted value changes how YAML parses it.
An audit shows several devices no longer match their intended configuration, but some differences came from emergency changes during incidents and are now approved. How would you design the reconciliation process so automation corrects real drift without undoing legitimate changes?
Sample Answer
Direct answer
Automation cannot tell from a device alone whether a difference is legitimate, so legitimacy must be recorded somewhere it can read: either the intended configuration (Git or the source of truth) or a register of approved exceptions with a ticket and an expiry. The reconciler pulls each device's running config, compares only the sections you manage, and sorts each difference into a class. Differences with a valid approval are adopted into intent through a reviewed change, not reverted. Differences with no approval are auto-corrected only for low-risk sections and are held for a human for routing, ACL and interface changes. Run it report-only first, then enable correction section by section.
Handling the differences the audit found today
Treat each audit finding as a ticket with three possible outcomes:
- The emergency change was approved and is still wanted: open a merge request that adds it to intent, link the incident ticket, and merge it. After that the difference no longer exists.
- The emergency change was a workaround that should go: schedule its removal as a normal change.
- Nobody can account for it: treat it as a possible security event, escalate, and correct it only after the owner or security signs off.
A reverse sync (copying what is on the device back into intent) is a process and not a button, and outcomes 1 and 2 are the reason: the right action depends on a human decision that the tooling records.
Detection: polling versus event-driven
| Approach | Strength | Weakness |
|---|---|---|
| Scheduled pull and diff | Catches every change however it was made, including ones that never generated an event; simple to reason about | Delay up to the interval; load on devices |
| Event-driven (a device configuration-change notice triggers a pull of that device) | Fast, targeted | Events can be lost (syslog, the standard protocol devices use to send log messages, over UDP has no delivery guarantee) or not generated for every kind of change |
Recommendation: a scheduled full pull as the authority (nightly for the fleet, hourly for critical tiers) with event triggers as an accelerator that starts an immediate targeted check. Because the schedule is the backstop, a lost event costs only latency. Snapshot the configs into Git after every approved change and nightly otherwise, so the repository holds both the intended state and a history of observed state.
Diffing approach
- Normalise before comparing: drop volatile lines (timestamps, the build header, counters), mask secrets before storage, and compare only managed sections. Without this the report is full of noise.
- Compare structured sections against intent, not the whole file against a template, so a line the template does not model is out of scope.
- Tools: NAPALM (a Python library giving one API across vendors) provides
get_config(retrieve='running')to read the config and, withload_replace_candidatefollowed bycompare_config(), returns the difference between running and the candidate; its IOS driver supports replace and compare. On IOS itselfshow archive config differences <file1> <file2>marks lines in the first file only with-and lines in the second only with+. For Ansible,--checkwithios_configreports in itsupdatesresult the commands it would send;--diffadds a before and after only when the task setsdiff_againsttostartuporintended(runningis not available in check mode), so a bare--check --diffshows no diff.
Classifying and acting
The script below compares intended and running lines for one device and applies the rules. The approved register carries a ticket and expiry per line, and only these low-risk sections are auto-corrected: NTP (Network Time Protocol) servers, SNMP (Simple Network Management Protocol) communities and logging hosts.
"""Classify drift between intended and running config lines for one device."""
import difflib
from datetime import date
intended = """hostname sw-ams-01
ntp server 10.0.0.1
ntp server 10.0.0.2
logging host 10.0.9.9
ip route 0.0.0.0 0.0.0.0 10.1.1.1
""".splitlines()
running = """hostname sw-ams-01
ntp server 10.0.0.1
ntp server 10.0.0.7
logging host 10.0.9.9
ip route 0.0.0.0 0.0.0.0 10.1.1.1
ip route 198.51.100.0 255.255.255.0 10.1.1.9
snmp-server community public RO
""".splitlines()
# Approved emergency changes: line -> (ticket, expires). Only lines listed here are adopted.
approved = {"ip route 198.51.100.0 255.255.255.0 10.1.1.9": ("INC-5120", date(2026, 12, 31))}
auto_fix_prefixes = ("ntp server", "snmp-server community", "logging host") # low-risk classes
today = date(2026, 10, 5)
for d in difflib.unified_diff(intended, running, "intended", "running", lineterm="", n=0):
if d[0] not in "+-" or d.startswith(("+++", "---")):
continue
side, line = d[0], d[1:]
if side == "+": # on the device, not in intent
if line in approved and approved[line][1] >= today:
t = approved[line]
print(f"ADOPT {line!r}: open PR to add to intent ({t[0]}, review by {t[1]})")
elif line.startswith(auto_fix_prefixes):
print(f"REMOVE {line!r}: low-risk class, push intended state")
else:
print(f"HOLD {line!r}: not approved, page the owner")
elif line.startswith(auto_fix_prefixes): # in intent, missing from device, low-risk class
print(f"RESTORE {line!r}: add back from intent")
else: # in intent, missing from device, routing or other risky class
print(f"HOLD {line!r}: missing from device, propose the diff for approval")
Run python drift.py and it prints:
RESTORE 'ntp server 10.0.0.2': add back from intent
REMOVE 'ntp server 10.0.0.7': low-risk class, push intended state
ADOPT 'ip route 198.51.100.0 255.255.255.0 10.1.1.9': open PR to add to intent (INC-5120, review by 2026-12-31)
REMOVE 'snmp-server community public RO': low-risk class, push intended state
In the diff the script reads (a unified diff), a line starting - exists only in the first file (intent) and a line starting + exists only in the second (the running config), which is why + means "on the device, not in intent" and - means "missing from the device". The same class rule applies in both directions: if the default route were missing from the running config, the script would print HOLD 'ip route 0.0.0.0 0.0.0.0 10.1.1.1': missing from device, propose the diff for approval, because re-adding a route unattended is a routing change. Reading the output: the NTP drift is corrected both ways (the stray server removed and the intended one restored), the open SNMP community is removed as a low-risk correction, and the extra static route is adopted because INC-5120 approves it and has not expired. A static route with no approval would print HOLD and page the owner, because deleting a route that carries live traffic is an outage.
Emergency changes
Allow break-glass changes (emergency changes made directly on a device outside the normal process) during an incident; they are legitimate. Require a ticket reference and a back-port deadline (for example five business days) when the incident closes. Each entry in the register expires: when it does, it is either merged into intent or scheduled for removal, so exceptions cannot become permanent undocumented configuration. Report the age of open exceptions to the owning team.
Minimising disruption on a mixed fleet
- Auto-correct by section risk: low-risk classes fully automatic, everything touching routing, ACLs (access control lists) or interfaces becomes a proposed diff for approval.
- Correct with the smallest delta (the specific lines) rather than replacing the whole configuration, and use a full replace only where the platform supports it, with a timed revert.
- Use a capability record per device (platform, software version, supported method) so the engine picks line-level push for older devices and replace for newer ones, with one driver layer such as NAPALM or per-vendor Ansible modules.
- Batch by blast radius (how much can break if one change goes wrong): the same waves, one member of a redundant pair at a time, halt on a failed check.
- Measure before enforcing: run report-only for two weeks and count false positives, then enable correction for one section at a time.
Pitfalls
- Reverting everything that differs: that undoes an emergency fix and recreates the incident.
- Comparing unnormalised files: the noise trains people to ignore the report.
- Event-only detection with no scheduled backstop.
- Approvals that never expire.
Unlock Full Question Bank
Get access to all 43 Network Automation and Software-Defined Networking interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.