Network Automation and Software-Defined Networking Questions
Programmatic and automated network operation: network automation tooling (Ansible network modules, Netmiko, NAPALM, Nornir, Jinja2 templating, ZTP), choosing between CLI-scraping libraries and model-driven programmability (NETCONF, RESTCONF, gNMI, YANG) across multi-vendor fleets, idempotent and declarative-versus-imperative design, a network source of truth (NetBox) and intended-versus-running configuration drift reconciliation, Git-based versioning and review of device configuration and automation code, pre- and post-change validation (for example Batfish) including lab or emulated test environments, CI/CD and staged, canary or rollback-safe rollout of configuration to large device fleets with concurrency control and credential and secret handling for automation, automated site and leaf-switch provisioning, and event-driven remediation from streaming telemetry. Also software-defined networking: control and data-plane separation, controller architecture and state consistency, controller-managed overlays and tenant network virtualization, SD-WAN controller orchestration and programmable data planes (P4). Boundary: general scripting and Terraform practice, shell craft, IAM, incident response, protocol design, data-centre fabric and WAN topology design (including choosing EVPN-VXLAN or SD-WAN), and network fault diagnosis are covered elsewhere.
How would you manage network device configurations and automation code in Git? What would you version, how would you handle branching, releases tied to change windows, and anything sensitive?
Sample Answer
Direct answer
Treat the Git repository as the statement of what the network should look like: version the inputs that generate configs (templates, variables, inventory) and the automation code that applies them, never plaintext secrets or generated output. Use short-lived branches (separate lines of work that are merged back within days) merged to a protected main (a branch that nobody can change except through a reviewed merge request, the Git host's pull-request equivalent), tag each change window, and keep a separate repository of device backups as evidence of what is actually running. Rollback is a revert plus a redeploy of the last good tag, with a device-side safety net behind it.
What to version, and what not to
| Version in Git | Keep out of Git |
|---|---|
Jinja2 templates (config text with {{ placeholders }}) and the variable files (group and host variables) that fill them | Plaintext passwords, SNMP communities, API tokens, private keys |
| Playbooks, roles and the Python that drives devices | Generated per-device configs (rebuild them from inputs; store as pipeline artifacts) |
| Tests and the pipeline definition file | Logs and ad-hoc output |
| Pinned dependency lists (Python packages, Ansible collections) so a rebuild next year behaves the same | Large device images (store in an artifact repository) |
| An inventory, or the exported view of the source of truth that feeds it |
A template and its variables turn into a config like this (illustrative values, rendered with Jinja2 in a container):
Template: interface {{ name }}\n description {{ description }}\n ip address {{ ip }} {{ mask }}
Variables: name: GigabitEthernet0/1, description: uplink to core, ip: 192.0.2.1, mask: 255.255.255.252
interface GigabitEthernet0/1
description uplink to core
ip address 192.0.2.1 255.255.255.252
Git holds the template and the variables, and the output is rebuilt whenever either changes, so a review shows the one-line variable change instead of a whole generated file.
Device backups (the show running-config captures) go in a second, private repository. They answer "what was on this box last Tuesday" and show drift, but they are outputs, and mixing them with intent makes every diff noisy.
Sensitive data
Encrypt secrets that must live beside the code with Ansible Vault (AES256), or fetch them at run time from a secrets manager. Encrypted vault files carry a recognisable header, which also lets a pre-commit hook reject unencrypted ones:
echo secretpw > vault_password.txt
printf 'ansible_password: placeholder-pw\n' > plain.yml
ansible-vault encrypt --vault-password-file vault_password.txt --output group_vars/all/vault.yml plain.yml
head -c 25 group_vars/all/vault.yml
$ANSIBLE_VAULT;1.1;AES256
Add the vault password file to .gitignore, supply it in CI (continuous integration, the automated pipeline) through a protected, masked pipeline variable (a CI setting that is only exposed to protected branches and hidden from job logs), and run a secret scanner (a tool that searches commits for password- and key-shaped strings) in the pipeline. Secrets here include SNMP communities, the shared strings that act as passwords for SNMP monitoring. If a secret does reach history, rotate it: deleting the commit does not un-leak it.
Branching and releases tied to change windows
- Trunk-based (everyone merges small changes into one main line instead of keeping long-lived branches):
mainis always deployable and always equals the intended production state. Work happens on short-lived branches (days, not weeks) so two engineers do not diverge on the same device group for a month. - A merge request (a proposal to merge a branch into
main, with review and automated checks attached) needs a reviewer from the owning team (a code owners file maps paths to teams, so a change underdatacenter/requires a datacenter reviewer), and the change-record number goes in the commit message so audit can trace a line to an approval. - Merging is not deploying. Deployment happens in the approved window, from an annotated tag (a named, permanent marker on one commit, with a message and an author) named for the window. The sequence below was run in a real Git repository in a container, with a file
ntp.cfgthat already held one NTP line when the tag was made:
git tag -a window-2026-10-17 -m "change window 2026-10-17 content"
echo "ntp server 10.0.0.2" >> ntp.cfg
git commit -am "add second NTP server"
git diff window-2026-10-17 HEAD --stat
git revert --no-edit HEAD
git describe --tags
ntp.cfg | 1 +
1 file changed, 1 insertion(+)
window-2026-10-17-2-gfea054f
Reading it: right after the tag, HEAD (the newest commit) is the tagged commit, so a diff would be empty. The commit that adds the second NTP server is the first commit after the tag, and git diff <tag> HEAD --stat then shows exactly what would go out beyond the last window: one file changed, one line added. git revert does not erase that commit; it creates a second new commit that undoes it, keeping history. git describe --tags names the current commit relative to the nearest tag as <tag>-<commits since the tag>-g<short hash>, so two commits after the tag (the change and its revert) prints window-2026-10-17-2-g.... The hash at the end differs on every run.
- Emergency changes use a fast-track branch with one reviewer, and any change made directly on a device must be pulled back into the repository afterward so the repo does not silently stop being true.
Where device features fit in the workflow
The change flow is: merge request, lint and render checks, a dry run posted on the request, approval, merge, tag at the window, apply, verify.
- Dry run and compare:
ansible-playbook --checkwithios_configsends nothing and reports in itsupdatesresult the commands it would send (visible with-v).--diffadds a before and after only when the task setsdiff_againsttostartuporintended(withintended_config);diff_against: runningis not available in check mode, where the module warns "unable to perform diff against running-config due to check mode", so a bare--check --diffprints no diff. NAPALM (a Python library that gives one API over many vendors) offersload_merge_candidateorload_replace_candidatefollowed bycompare_config(), which returns the difference between running and candidate, anddiscard_config()to drop it. - Transactional safety on the device: on IOS,
configure replace <file> time <minutes>reverts automatically unless you enterconfigure confirm, which is the device-side equivalent of a Git revert and protects you if the change cuts your own management path. NAPALM'scommit_config(revert_in=...)exposes a timed revert on supported platforms. - Drift detection (drift is a device running something different from what Git says it should): a scheduled job that compares each device with the intended config and opens a merge request for differences keeps the repository and the network from drifting apart.
Pitfalls
- Squashing history so thoroughly that a window's content cannot be reconstructed: tag every window.
- Long-lived per-region branches: they drift, and the merge itself becomes the risky change.
- Storing the vault password in the same repository as the vault file.
How would you manage credentials and secrets for a network automation system that logs in to thousands of devices? Think through storage, how long credentials live, who can use them, and how access is audited.
Sample Answer
Direct answer
Keep credentials in a secrets manager (a service such as HashiCorp Vault, not the inventory file or the git repo), hand the automation a short-lived credential for each job, give each kind of job its own least-privilege account (one that can do only what that job needs), and make every login traceable to a job ID and a change ticket. Where devices cannot take short-lived credentials, use a rotated per-function service account and rotate it with the same automation.
Storage
- Secrets live in the secrets manager and are fetched at job start, held in memory, and never written to logs, playbook output or the repository. Ansible's own guidance is never to keep passwords in plain text and to use Ansible Vault (its built-in file encryption) to keep them out of playbooks and source control, with the
no_logtask setting to keep them out of job output. - Ansible Vault protects the file, but its password is one static secret every playbook runner shares. That is why a dedicated secrets manager with its own authentication for the automation host is the better home at thousands of devices.
- A sealed break-glass local account on each device (an emergency login kept offline, used only when normal access fails), stored with two-person access, covers the case where the AAA server (authentication, authorization and accounting; a central server that devices ask who may log in and what they may do, for example over TACACS+, a protocol for that conversation) or the secrets manager is down.
How long credentials live
| Credential | Lifetime | Mechanism |
|---|---|---|
| Per-job SSH credential | Minutes | Vault's SSH secrets engine can act as a certificate authority and sign an SSH key with a ttl (time to live, how long the certificate stays valid; its docs show 30 minutes); the target trusts the CA's public key. Whether a given network OS accepts SSH user certificates is vendor specific, so confirm that before designing around it. |
| Service account password where certificates are not possible | Days to weeks, rotated by automation | Add the new credential, test a login with it, switch the secrets manager to it, remove the old one. |
| SNMP communities, API tokens | Rotated on a schedule, separate from login credentials | Same overlap pattern. |
How a signed key replaces a password: the automation host has an ordinary SSH key pair and asks Vault to sign the public half. Vault checks the job is allowed and returns a certificate, which is that key plus a signature made with the CA's private key and an expiry time. Each device is configured once with the CA's public key (in OpenSSH, TrustedUserCAKeys), so it accepts any unexpired certificate the CA signed, and no password or per-user key has to be stored on the device. When the ttl passes, the certificate stops working and there is nothing to revoke.
Why it matters, computed: a 90-day static password is exposed for 90 x 24 = 2,160 hours; a 30 minute certificate for 0.5 hours. That is 2,160 / 0.5 = 4,320 times less exposure per leaked credential.
Who can use them
Split by function, not by person: a read-only collector account for polling and facts, a config-push account for approved changes, and no shared admin account. On the device side, use per-command authorization from the AAA server (it approves or denies each command a user types) so the read-only account cannot enter configuration mode even if the job misbehaves. On the platform side, the automation tool's role-based access decides which humans can launch which job against which device group, and production pushes require the approved change record.
How access is audited
Three records must join on one job ID:
- The platform log: which human or pipeline launched which job, with which ticket.
- The secrets manager audit log: which job fetched which credential and when.
- The device-side log: TACACS+ accounting (the accounting function of the protocol, RFC 8907) records the commands run under the service account. The device only sees the service account, so the job ID has to come from steps 1 and 2.
TACACS+ protects its packets with obfuscation, not real encryption, and RFC 8907 requires deploying it over networks that ensure privacy and integrity, so run it on the management network.
Worked example: rotating a service account on 3,000 devices
Assume 6 seconds per device and 30 parallel workers (both assumed figures). Time = 3,000 x 6 / 30 = 600 seconds, 10 minutes. The rotation job records a result per device (rotated, failed login with new credential, unreachable). Because the old credential is removed only after the new one is verified on that device, a failed device still accepts the old one and can be retried, instead of being locked out of the automation.
The table's steps (add, test, switch the secrets manager, remove) apply per device, not once for the fleet. The secrets manager keeps both versions, and the rotation record says which one to present to each device: the old one until that device is verified on the new, the new one afterwards. Switching one shared secret for the whole fleet before every device has the new credential would lock the automation out of the stragglers, and so would removing the old credential from a device while the manager still presents the old one. Retire the old version in the secrets manager only after the last device is verified.
Pitfalls
- A single shared "automation" account with enable-level rights means one leaked secret is a full-fleet compromise, and the accounting log cannot tell the jobs apart.
- Secrets in CI environment variables echoed by a debug step, or in the device config backups the system takes, leak quietly. Mask them and scrub backups.
- Rotating without the overlap step is how a fleet gets locked out of automation.
A network team wants every configuration change reviewed in Git before deployment, with per-device values, generated interface stanzas and safe credential handling. What goes in version control, what is rendered or injected at deploy time, and how do you structure the review?
Sample Answer
Direct answer
Version-control the intent and the code that turns intent into configuration: templates, per-device and per-group data, playbooks, tests and review rules. Generate the actual interface stanzas at review time and at deploy time from that source, and inject credentials only at run time from a secret store. Reviewers then approve two things on one pull request: the change to the source, and the exact configuration lines it produces, shown as a diff.
What goes where
| Item | In Git? | Notes |
|---|---|---|
| Jinja2 templates | Yes | Reviewed by the design owners |
Per-device and per-group values (host_vars, group_vars) | Yes | Plain YAML, one file per device, so a diff names the device |
| Playbooks and the shared library | Yes | Same review as templates |
| Tests, lint configuration, CI pipeline definition | Yes | The review rules are themselves versioned |
| Rendered device configuration | Generated, not hand-edited | Produced by CI for the review diff; not committed by hand |
| Passwords, SNMPv3 keys, API tokens | Never in plaintext | Encrypted values, or fetched at run time |
| The vault password or secret-store token | Never in the repo | Held by the CI system and the runner |
Rendered output is kept out of the hand-edited tree so nobody "fixes" a generated file and loses the change at the next render.
Repository layout
network-config/
CODEOWNERS
inventory/
hosts.yml
group_vars/ # per role or region values
host_vars/ # nyc-edge1.yml, one file per device
templates/ # interfaces.j2, aaa.j2 ...
playbooks/ # render.yml, deploy.yml
tests/ # sample data and expected output
ci/ # pipeline definition
Reading the tree: CODEOWNERS is the review-rule file; inventory/ holds the data (hosts.yml lists devices, group_vars/ values shared by a role or region, host_vars/ one file per device); templates/ holds the Jinja2 files that turn data into configuration; playbooks/ holds the render and deploy entry points; tests/ holds sample data with the expected output; ci/ defines the pipeline that runs the checks. Nothing in it is a rendered config or a secret.
Credentials
- Preferred: device credentials live in a secret manager and the runner reads them when the job starts, so no secret exists in Git history at all and rotation needs no commit.
- Smaller teams: Ansible Vault.
ansible-vault encrypt_stringturns one value into an inline block that is safe to commit, for example a lineios_enable_secret: !vault |followed on the next, indented line by$ANSIBLE_VAULT;1.1;AES256and then the encrypted text (I ran the command; the output began that way, over several lines). The vault password is supplied to the job by CI through a password file or script, and the playbook should useno_log: true(an Ansible task setting that hides the task's arguments and results from logs) on tasks that handle the secret so it never prints. - Encrypted-in-Git still means anyone with the vault password can read every secret and history keeps old ciphertext, so rotate the vault password and the secrets together if either leaks.
Structuring the review
- Branch and pull request for every change, one logical change per request.
- Automatic checks first, before any person reads it: YAML lint, Ansible syntax check, data checks (valid IPs, no overlaps, required fields), and rendering for every affected device.
- Rendered diff on the pull request. CI renders the target branch and the proposed branch and shows the unified diff. I executed this in a container: changing one host's address from
10.20.1.1to10.20.1.254produced
- ip address 10.20.1.1 255.255.255.0
+ ip address 10.20.1.254 255.255.255.0
with every other line unchanged. A reviewer who sees a three-line template change cause 400 lines of diff across 60 devices knows to look closer.
4. Review by ownership. A CODEOWNERS file (a file in the repository that maps paths to the people or teams who must review changes to them) maps paths to people or teams (for instance templates/ to the design team and inventory/host_vars/nyc-* to the New York site owners). With branch protection (a repository setting that blocks merging until rules such as required reviews and passing checks are met) set to require code-owner review, GitHub needs approval from at least one code owner of each changed path.
5. Merge triggers deployment of the reviewed commit, pinned by its commit identifier, to a canary set (a small first group of devices that gets the change before the rest) first, then wider. The reviewed state and the deployed state must be the same commit.
6. Compare after deploy. A scheduled job compares each device's running configuration with the render of main and reports differences.
Trade-offs and pitfalls
- Template review without rendered output is the commonest mistake: people approve a loop they cannot mentally expand.
- Check mode against live devices is a useful pre-deploy gate but needs device access, so run it from a runner with network reach, not from the pull-request CI. For
cisco.ios.ios_config,--check -vprints thecommandslist it would send;--diffalone prints nothing for that module unlessdiff_againstis set (from the collection's source). - Emergency changes will happen. Provide a fast path (pair review, deploy, then back-port to Git) rather than banning it, or engineers will bypass the whole process.
- Never print credentials in CI logs; test that the pipeline output contains no secret material.
Tell me about a time an automated network change caused an unexpected problem in production, or nearly did. What was your role, how did you contain the impact, and what did you change in the automation or rollout process afterward?
Sample Answer
Direct answer
A strong answer is one real incident told in order: the change, the signal that something was wrong, what you did in the first minutes to limit the damage, how you restored service, and the specific changes to the automation and rollout that now prevent a repeat. The example below is an illustrative story with arithmetic you can check; use your own facts.
Story skeleton
Situation. I owned a playbook that kept the VLAN list on 40 access switches in line with our source of truth (the repository that holds the intended state). A cleanup moved VLAN data between two files and I ran the playbook with ios_vlans in state: overridden, which replaces all VLANs on the device with the configuration supplied. The moved file was missing one site's voice VLAN (the separate VLAN, a virtual LAN that splits one physical switch into isolated networks, that carries IP phone traffic; without it on a switch, the phones plugged into it cannot register with the phone system).
What happened. The first batch was a single canary switch (one device changed first, as an early warning, before the rest). The task itself finished green, but a check I had added after an earlier near-miss, a comparison of the VLAN count before and after the change, flagged that the canary had lost a VLAN. Because I had rolled out in batches (1, 2, 3, ...), the other 39 switches were untouched. serial is the Ansible play keyword that sets how many hosts are changed at a time; with it set to a list of increasing batch sizes, a failure scopes to the batch, not the whole host list (Ansible documentation).
Containment (my role).
- Stopped the run, and told the on-call and the site owner what had changed and on which switch.
- Restored the missing VLAN on the one affected switch from the pre-change configuration backup, then confirmed the phones registered again.
- Ran a read-only
state: gatheredpass (theios_vlansstate that reads the device's current VLANs back as data and changes nothing) on all 40 switches and compared each switch's VLANs with the source of truth, to confirm no other switch was missing a VLAN or held one the data did not list.
Impact arithmetic. 1 affected switch out of 40 is 1/40 = 2.5 percent of the fleet; 39 of 40 (97.5 percent) were never touched because the run stopped after the first batch. If I had run all 40 at once, every switch with the same data would have been hit (the story's switch counts, not measured figures).
What I changed afterwards.
- Guard: keep
serialin growing batches and add a failure threshold (max_fail_percentage, the percentage of hosts in a batch allowed to fail before Ansible aborts the play; at 0, one failure aborts it), so the first bad batch ends the run by itself instead of depending on me to notice and stop it (Tested on six hosts withserial: [1, 2, 3]: a failing host in the one-host first batch ended the play with or without this setting, because every host in that batch failed. The setting earns its place in the later, larger batches: when one of the two hosts in the second batch failed, the default let the third and later batches run, whilemax_fail_percentage: 0stopped the play before them.) - Pre-flight: a read-only check that runs before anything is changed; it reads each switch's current VLANs with
state: gathered, compares them with the intended data, and fails the pipeline if the run would delete any VLAN not named in the pull request (therenderedstate cannot do this, because it only turns your data into commands and never looks at the device). - Scope: prefer
mergedfor routine runs; useoverriddenonly through a reviewed change with a diff. - Review: the data files are validated for required entries in CI (continuous integration) before a merge.
- Blameless review (written without assigning fault to a person), shared with the team.
What I would do differently. Build the pre-flight diff before the first production use of overridden, not after this near-miss.
What interviewers look for
- A clear statement of your role and decisions, with timing of detection (how you found out) and of containment.
- Containment first, root cause second.
- Process fixes aimed at the pipeline (canary, batching, diff gate, scope), not a promise to be more careful.
Pitfalls
- Do not blame the tool or a colleague; describe the gap in the process.
- Do not invent precise outage durations you did not measure; use the ones you recorded.
- Include the contributing factor that made the blast radius larger or smaller (batching, backups, monitoring).
Design a CI/CD pipeline for network configuration changes, from a commit in Git to production devices across about 50 sites. Walk through each stage, how you would gate and approve changes, how a failure triggers rollback, and how secrets are kept out of the pipeline.
Sample Answer
Direct answer
A merge request runs fast, offline checks (lint, render, unit tests, a throwaway virtual lab). After merge, the pipeline takes a pre-change snapshot (a saved copy of each device's running configuration) of every target device, shows a dry-run diff (a preview of what would change, with nothing sent), and then waits for a human to start a canary site (the first, low-risk site that gets the change), then waves of sites (groups of sites updated together, one group at a time), with a verify run inside each deploy job and a whole-fleet verify stage at the end. Any failed stage triggers a rollback job that restores each device's pre-change snapshot with configure replace. Secrets never enter the repository: they live in protected, masked CI variables and an encrypted Ansible Vault. This is a CI/CD (continuous integration and continuous delivery) design, shown in GitLab CI and equally implementable in GitHub Actions or Jenkins; the stage ideas are what matter.
Stages and gates
| Stage | What runs | Gate that must pass |
|---|---|---|
| lint | yamllint, ansible-lint, playbook syntax check | Merge request blocked on failure |
| render | Build every device's config from templates and variables, then run unit tests (VLAN ids in range, subnets do not overlap, no allowed vlan all) | Merge request blocked |
| lab | Start a virtual copy of one site with containerlab (a tool that runs network devices as containers), apply the change, run the verify playbook, destroy the lab | Merge request blocked. Routing-protocol changes (BGP, the Border Gateway Protocol, or OSPF, Open Shortest Path First) get their own lab topology here, with a neighbour that must keep its session and routes after the change |
| plan | --check -v against production, saved as an artifact (a file the pipeline keeps from a job); each ios_config task reports in updates the commands it would send | The person about to start the canary reads the pending commands in the job output (this job runs after merge, on the default branch) |
| snapshot | Back up every target's running config to the central store | Fails closed (on error it blocks the deploy rather than letting it through): no snapshot, no deploy |
| deploy-canary | Apply to one low-risk site | Manual start, one person |
| deploy-wave | Apply to site groups wave 1, 2, 3 | Manual start per wave; one wave at a time |
| verify | Check device state against expectations | Failure triggers rollback |
| rollback | Restore snapshots, then verify again | Runs only after a failure |
Approval has two layers. Peer review and approvals on the merge request approve the content. The manual start of each deploy job approves the timing (people start it inside the change window with the change-record number in the commit). Separately from approval, a resource group (resource_group) serialises production jobs so two pipelines cannot push at once.
The pipeline definition
Keywords used below (from the GitLab CI documentation): stages lists the stage names in order; rules decides whether a job runs (here: on a merge request, or on the default branch); needs lets a job start once the named jobs finish, instead of waiting for the whole previous stage; interruptible says whether a newer pipeline may cancel the job (false for anything touching production); resource_group allows one job at a time in that group; parallel: matrix creates one copy of the job per listed variable value (here, one per wave); when: manual waits for a person to start the job; when: on_failure runs the job only if an earlier job failed; artifacts keeps files from a job; after_script runs commands after the job body even if it failed; allow_failure: false means a failure fails the pipeline; environment names what is being deployed to.
stages: [lint, render, lab, plan, snapshot, deploy-canary, deploy-wave, verify, rollback]
default:
image: registry.example.com/netauto/runner:2026-10
interruptible: true
variables:
ANSIBLE_VAULT_PASSWORD_FILE: ci/vault_pass.sh
lint:
stage: lint
script:
- yamllint .
- ansible-lint playbooks/
- ansible-playbook -i inventory/prod.ini playbooks/deploy.yml --syntax-check
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
render:
stage: render
script:
- python3 ci/render_all.py --out build/
- python3 -m pytest tests/ -q
artifacts:
paths: [build/]
expire_in: 30 days
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
lab:
stage: lab
script:
- containerlab deploy -t lab/campus.clab.yml
- ansible-playbook -i lab/inventory.ini playbooks/deploy.yml
- ansible-playbook -i lab/inventory.ini playbooks/verify.yml
after_script:
- containerlab destroy -t lab/campus.clab.yml
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
plan:
stage: plan
script:
- ansible-playbook -i inventory/prod.ini playbooks/deploy.yml --check -v | tee plan.txt
artifacts:
paths: [plan.txt]
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
snapshot:
stage: snapshot
script:
- ansible-playbook -i inventory/prod.ini playbooks/backup.yml
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
deploy-canary:
stage: deploy-canary
environment: production
resource_group: production-network
interruptible: false
script:
- ansible-playbook -i inventory/prod.ini playbooks/deploy.yml --limit canary
- ansible-playbook -i inventory/prod.ini playbooks/verify.yml --limit canary
needs: [render, plan, snapshot]
artifacts:
when: always
paths: [touched/]
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
allow_failure: false
deploy-wave:
stage: deploy-wave
environment: production
resource_group: production-network
interruptible: false
script:
- ansible-playbook -i inventory/prod.ini playbooks/deploy.yml --limit "wave_${WAVE}"
- ansible-playbook -i inventory/prod.ini playbooks/verify.yml --limit "wave_${WAVE}"
parallel:
matrix:
- WAVE: ["1", "2", "3"]
artifacts:
when: always
paths: [touched/]
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: manual
needs: [render, deploy-canary]
verify:
stage: verify
environment: production
script:
- ansible-playbook -i inventory/prod.ini playbooks/verify.yml
needs: [deploy-wave]
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
rollback:
stage: rollback
environment: production
resource_group: production-network
interruptible: false
script:
- |
if [ -z "$(find touched -type f 2>/dev/null)" ]; then
echo "No device was touched; nothing to roll back"
exit 0
fi
hosts=$(find touched -type f -printf '%f\n' | sort | paste -sd, -)
ansible-playbook -i inventory/prod.ini playbooks/rollback.yml --limit "$hosts"
ansible-playbook -i inventory/prod.ini playbooks/verify.yml --limit "$hosts"
rules:
- if: $CI_COMMIT_BRANCH == $CI_DEFAULT_BRANCH
when: on_failure
deploy.yml pushes the rendered snippet to five devices at a time and stops everything if any device in a batch fails. Before each batch changes anything, its pre-task drops an empty file named after each device into touched/, which is the record the rollback job reads. max_fail_percentage applies per batch and the percentage must be exceeded, so with serial: 5 and a value of 0 one failure (20 percent) halts the run:
---
- name: Apply the rendered configuration
hosts: all
gather_facts: false
connection: ansible.netcommon.network_cli
serial: 5
max_fail_percentage: 0
vars:
ansible_network_os: cisco.ios.ios
pre_tasks:
- name: Make sure the touched-hosts directory exists
ansible.builtin.file:
path: "{{ playbook_dir }}/../touched"
state: directory
mode: "0755"
delegate_to: localhost
run_once: true
- name: Record that this device is about to be changed
ansible.builtin.copy:
content: ""
dest: "{{ playbook_dir }}/../touched/{{ inventory_hostname }}"
mode: "0644"
delegate_to: localhost
tasks:
- name: Push the rendered snippet
cisco.ios.ios_config:
src: "../build/{{ inventory_hostname }}.cfg"
How to read the file: the first four jobs (lint, render, lab, plan) run on a merge request or main and never touch a device except plan, which only previews. snapshot saves a copy of every device's config. deploy-canary and deploy-wave are the only jobs that change devices: both are manual, both share resource_group: production-network so only one runs at a time, and the wave job fans out into wave_1, wave_2, wave_3 through the matrix. Each of those jobs runs verify.yml for its own group straight after the deploy, so a canary or wave that fails its checks fails the job, and that failure is what the rollback job's when: on_failure condition reacts to. Waves 2 and 3 are still unstarted manual jobs in the same stage, and whether rollback starts before they are decided is GitLab pipeline-processing behaviour to confirm on your version, so the rule for the person starting jobs is not to start the next wave until the failure has been handled (the rollback job has run and its verify step is green). The resource group stops two jobs overlapping but its default process mode is unordered, so the order canary, wave 1, wave 2, wave 3 is a rule for the person starting the jobs, not something GitLab enforces. The final verify job checks the whole fleet once the waves are done, and rollback fires when any earlier job in the pipeline has failed: GitLab runs a when: on_failure job after a failure in any earlier stage, including lint, render, plan or snapshot, when no device has been touched yet. For that reason the rollback job must restore only the devices a deploy job actually started on. That is what the touched/ directory is for: deploy.yml writes one empty file per device (named after the device, so files from different deploy jobs never collide) before the batch runs, each deploy job keeps the directory with artifacts: when: always, and the rollback job builds its --limit list from the file names. GitLab uploads a job's artifacts only when the job succeeds unless artifacts: when: always is set (the default is on_success), so without that setting a failed deploy leaves rollback with nothing to read. A job with no needs, like rollback, downloads the artifacts of earlier-stage jobs by default. If touched/ is absent or empty (a lint or plan failure), the job says so and exits without contacting any device; without that guard an empty --limit would not restrict the run, and a lint failure on the default branch would push whatever old snapshot sits in the store onto healthy devices and silently undo emergency changes made since.
Idempotence, dry run and canarying
- Idempotent (running it twice equals running it once): templates describe full desired state for a section, and tasks compare with the device before changing it, so a retried job after a timeout does no harm.
- Dry run:
ansible-playbook --check -vmakesios_configreport inupdatesthe lines it would send without sending them, so an emptyupdatesfor a device means no change will be sent. (--diffadds a before and after only when the task setsdiff_againsttostartuporintended;runningis not available in check mode, so a bare--check --diffprints no diff.) - Canary: wave order is canary site, then three waves of the remaining sites, each started by a person after reading the verify output that the previous job printed at the end of its own run. Choose the canary for being representative but non-critical, for example a mid-size site with the same hardware models as the big ones.
- Routing changes: the lab stage proves the syntax and the neighbour behaviour on virtual routers. In production, change one device of each redundant pair per wave, so the other path carries traffic while you watch sessions, and compare the number of established sessions and received prefixes per neighbour with their values before the change, before moving on.
Failure and rollback
Three layers, from fastest to broadest:
- Stop the blast:
serialplusmax_fail_percentage: 0halts the play at the first failed batch, so a bad change reaches at most five devices. - Restore exactly: the
rollbackjob runs only when an earlier stage failed on the default branch (when: on_failureinside its rule), limited to the devices listed intouched/.rollback.ymlcopies each snapshot from the store to the device flash (thenet_putmodule, which sends a file to the device, here over SCP, file transfer over SSH) and runsconfigure replacewith that file as the target. The play has no failure threshold, so one unreachable device does not stop the others from being restored. The job then runsverify.yml. Replace is used instead of re-pushing the old Git version because a merge-style push only adds lines: it cannot remove a line the bad change introduced. For example, if the bad change added the lineip route 0.0.0.0 0.0.0.0 10.9.9.9, pushing the old file (which does not contain that line) leaves the route in place, because nothing in the push says to delete it.configure replacemakes the running config equal to the file, so the line is removed.
---
- name: Restore the pre-change snapshot with configure replace
hosts: all
gather_facts: false
connection: ansible.netcommon.network_cli
serial: 5
vars:
ansible_network_os: cisco.ios.ios
tasks:
- name: Copy the snapshot from the store to the device flash
ansible.netcommon.net_put:
src: "/srv/netbackups/{{ site }}/{{ inventory_hostname }}.cfg"
dest: "flash:{{ inventory_hostname }}.cfg"
protocol: scp
- name: Replace the running config with the snapshot on flash
cisco.ios.ios_command:
commands:
- "configure replace flash:{{ inventory_hostname }}.cfg list force"
- Device-side timer for risky changes:
configure replace <file> time <minutes>reverts by itself unlessconfigure confirmis entered, so a change that cuts off the management path undoes itself with nobody needed. Replace needs a complete configuration file as the target; no archive orpathsetting is required (Cisco: no prerequisite configuration is needed to useconfigure replace). In the rollback command,listprints the lines applied andforceskips the confirmation prompt.
A snapshot is only as current as the run that took it: snapshot runs when the merge lands, but a person may start the canary hours or days later, and a rollback to an old snapshot would silently undo any emergency change made in between. Take the snapshot inside the change window, or re-run the snapshot job just before starting the canary. The snapshot job and the rollback job must read and write the same store: if the runner uses throwaway containers, a snapshot written inside the snapshot job's container is gone when the rollback job starts, so mount the store as a persistent volume or have both jobs clone the same store repository. After any rollback, revert the commit in Git so intent matches the network again.
Secrets
- No credentials in the repository: the vault password comes from a protected, masked pipeline variable and is supplied to Ansible by a tiny script named in
ANSIBLE_VAULT_PASSWORD_FILE(a vault password file may be a script). - Device passwords sit in an Ansible Vault file (AES256-encrypted) or in a secrets manager, read at run time.
- Job logs are artifacts too: do not print variables, and keep
no_logon tasks that handle credentials. - Use a dedicated automation account on devices with a command set limited to what the pipeline needs, so a stolen pipeline credential cannot do everything an engineer can.
Audit trail
Every production change maps to a merge request (who wrote it, who approved), a pipeline run (who started each deploy job and when), the plan artifact (what was expected), the snapshot (what it was before) and a tag per window. Retaining these for as long as your change policy requires answers an auditor's question without anyone's memory.
Pitfalls
- A rollback that was never exercised: run the rollback job against the lab stage on a schedule.
- Waves defined by region when the redundant pair spans regions: define waves by pair membership so a wave never takes both devices of a pair.
- Letting the manual job be started by anyone with repository access: restrict who may start production jobs.
Unlock Full Question Bank
Get access to all 6 Network Automation and Software-Defined Networking interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.