System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
Executives have cut maintenance windows to almost nothing. How do you keep hosts patched and hardened without regular downtime, and what do you tell the business about the risk that remains?
Sample Answer
Direct answer
Separate the two things reboots are doing in the plan: applying fixes, and restarting the machine. Reduce how often a machine must restart (live kernel patching, Windows hotpatching where supported, replacing instead of patching, rolling updates over redundant nodes), and be honest with the business that the rest cannot be eliminated, so give them a number for the remaining exposure and ask for the smallest guaranteed window that bounds it. Never promise "patched with no downtime and no risk".
Step 1: Sort the estate by how it can be updated without downtime
Each row below is a method for a different kind of host, so they are not alternatives for one machine; an estate usually uses several. (Hardening baselines, meaning agreed secure settings, come in Step 2. The word "baseline" in the hotpatch row means something else: a full cumulative update that does require a restart.)
| Host type | No-downtime method | What it does not cover |
|---|---|---|
| Stateless web or application tier behind a load balancer | Replace instances with a freshly built patched image, one batch at a time, draining each batch first | Nothing special; this is the cleanest answer |
| Clustered roles on Windows | Cluster-Aware Updating (CAU), Windows Server's built-in tool for updating a failover cluster (a group of servers that can take over each other's roles): puts each node into maintenance mode (a state in which the node is drained of work), moves the clustered roles off it, installs updates, restarts if needed, brings roles back, then does the next node | Many roles trigger a planned failover, which can cause a brief interruption for connected clients; only continuously available workloads (those designed to keep serving clients while a node moves, such as Hyper-V with live migration or a file server using SMB Transparent Failover) avoid it |
| Linux servers needing kernel fixes | Kernel live patching, which loads a replacement for a faulty kernel function into the running kernel and redirects calls to it, so the fix applies without a restart | Only functions that the kernel's tracing hook (ftrace, the mechanism live patching uses to intercept a call) can intercept can be patched, so some fixes still need a reboot, and user-space libraries need their services restarted |
| Windows Server on Azure, Azure Local (Microsoft's product that extends Azure to your own hardware, managed through Azure Arc) or Azure Arc-connected (servers outside Azure registered for Azure management) | Hotpatch (patches in-memory code without restart), supported only on specific editions: Datacenter: Azure Edition images of Windows Server 2022 and 2025 on Azure and Azure Local, and Windows Server 2025 Standard or Datacenter on Azure Arc-connected machines once Arc hotpatching is enabled (Windows Server 2022 is not supported on Arc) | A new hotpatch baseline (a full cumulative update that resets the cycle) arrives every three months and still requires a restart; non-security updates, .NET and driver or firmware updates are not hotpatched; hotpatch has no automatic rollback |
| Single-instance legacy servers | None | Needs a window, or accepted risk with compensating controls |
The hotpatch facts: the hotpatch baseline refreshes every three months, with hotpatch releases for the following two months, so a planned year is four baselines (restart needed) and eight hotpatch months. An unplanned baseline can replace a hotpatch month if a fix cannot be hotpatched.
Checking one host. After a live patch or hotpatch, confirm it landed on the host itself, because "the console says compliant" is a second-hand view.
- Linux live patch:
uname -rstill prints the old kernel version (a live patch does not change it), so check the live-patch interface instead. The kernel keeps one directory per loaded patch under/sys/kernel/livepatch/, andcat /sys/kernel/livepatch/<patch-name>/enabledreads1while that patch is active (patch names depend on the vendor's package; I could not run this here because a container has no live-patched kernel). - Windows hotpatch: list installed updates with
Get-HotFix, which lists the updates Component Based Servicing installed (columns Source, Description, HotFixID, InstalledBy, InstalledOn), and look for the month's KB number:(Get-HotFix | Sort-Object -Property InstalledOn)[-1]shows the most recent one, andGet-HotFix -Id KB5099999(an illustrative number) returns nothing if that update is absent. For Azure VMs, the VM's Updates page in the Azure portal also shows the hotpatch status.
Step 2: Hardening without downtime
Hardening changes are mostly configuration, not binaries. Roll them out as rolling, per-node changes behind the same redundancy, with the change applied to one node, verified, then repeated. Settings that need a service restart (not a host restart) are handled by draining that node. Build the settings into the golden image so new instances arrive hardened, rather than modifying running hosts.
Step 3: What to tell the business about the remaining risk
State it as a measurable exposure window, an owner, and a decision.
Worked example (illustrative numbers, stated as assumptions): suppose after the techniques above 40 of 500 hosts can only be restarted in a single scheduled window per quarter. The longest a fix can wait for such a host is the gap between windows, about 91 days (365 divided by 4, rounded to the nearest day). Live patching or hotpatch narrows that for kernel and Windows security fixes, but not to zero.
| Question the business asks | Answer you give |
|---|---|
| What is still exposed? | 40 hosts, listed, with the services they run |
| For how long? | Up to about 13 weeks for fixes that need a restart; days for those covered by live patching or hotpatch |
| What reduces it? | Compensating controls on those hosts (tight allow-lists, removal of unused services), plus one extra short window per quarter or funding a second node so the host becomes rolling-updatable |
| Who accepts what remains? | A named executive signs a risk acceptance with an expiry date, reviewed each quarter |
Recommendation: ask for the smallest fixed window that bounds exposure, for example two hours monthly for the 40 hosts, rather than none; each shorter gap proportionally shortens worst-case exposure (monthly gives about 30 days, a third of the quarterly gap).
Trade-offs and pitfalls
- Live patching is a mitigation, not a replacement for reboots: a kernel can be live-patched for months and still drift from the tested baseline.
- Replace-instead-of-patch moves risk into the image pipeline; the pipeline must be patched and tested too.
- Do not hide the residual risk in a spreadsheet. If an incident occurs on an accepted host, the signed acceptance is what shows the business chose the exposure knowingly.
- What would change the recommendation: if the 40 hosts include internet-facing ones, the compensating controls alone are not enough and the second node becomes the priority.
Design the patching platform for 20,000 endpoints across regions, including air-gapped segments. What are the components, how do phased rollouts and access control work, and what evidence of compliance does it produce?
Sample Answer
Direct answer
Build the platform as a pipeline with two zones and one control plane. An ingest zone (the only place with internet access) downloads vendor updates and holds them in staging. A distribution layer of regional servers serves approved content to endpoints, which pull from the nearest server. Air-gapped segments (networks with no connection to the internet or to the main network) receive content through a signed offline transfer bundle that is exported, carried across the gap, verified and imported into a server inside the segment, which then distributes it locally. Rings (groups of endpoints patched in a fixed order) control the rollout, separate roles control who can publish, approve and deploy, and a central compliance database records what each endpoint reports.
The assumptions for the sizing below are stated, not measured: 20,000 endpoints in total, of which 1,500 sit in air-gapped segments, four regions, and an average download of 500 MB per endpoint per patch cycle. Replace them with your own inventory.
Components
flowchart LR
V[Vendor update feeds] --> I[Ingest and staging server]
I --> T[Test ring lab]
I -->|approved content| R1[Regional server A]
I -->|approved content| R2[Regional server B]
I -->|signed export bundle| M[Offline media transfer]
M --> X[Import server in air-gapped segment]
X --> S[Segment distribution server]
R1 --> E1[Endpoints region A]
R2 --> E2[Endpoints region B]
S --> E3[Air-gapped endpoints]
E1 -.->|state reports| C[Compliance database]
E2 -.->|state reports| C
The diagram shows two of the four regions, and each regional server box stands for a pair (primary and standby); the other regions repeat the same pattern.
| Component | Job | Why it exists |
|---|---|---|
| Ingest and staging server | Pulls vendor updates, keeps the catalog, is the only internet-facing node | One place to restrict outbound access and to inspect content |
| Test ring lab | Representative machines that install updates first | Finds breakage before real users do |
| Approval service | Records who approved which update for which ring and when | Separation of duties and an audit trail |
| Regional distribution servers (two per region for failover) | Serve approved content to endpoints in that region | Endpoints download locally, so the wide-area links carry one copy per region |
| Offline transfer path | Signed bundle plus checksums manifest, carried on controlled media or through a one-way transfer device | The only route into an air-gapped segment |
| Import server and segment distribution server | Verify and load the bundle, then serve local endpoints | Keeps the air-gapped side self-sufficient |
| Compliance database | Stores per-endpoint patch state and deployment history | Source of the evidence reports |
For Windows endpoints the Windows Server Update Services (WSUS) documentation describes both patterns used here. WSUS is documented by Microsoft as deprecated (no new features, still supported with security and quality updates), so it suits existing estates, and the same pipeline shape applies to newer update tools.
- Multi-server layout: only the topmost server needs internet access, and the others synchronize from it.
- No internet at all: two WSUS servers are used. One with internet access collects updates, they are exported onto removable media, carried across the gap and imported into a second server on the protected network.
- Wire protocol: WSUS sends update metadata over HTTPS (port 8531 by default) and payloads over HTTP (port 8530 by default). Payloads are still protected, because every payload is signed and its hash travels with the metadata over the secure channel; clients verify both before installing.
Linux and other endpoints follow the same shape with internal package repositories that you mirror as dated snapshots, so the same snapshot promoted through the rings reaches every machine.
What the offline bundle and its verification look like
A bundle is the update files plus a manifest (a list of every file and its SHA-256 hash) plus a detached signature over the manifest. The export side creates and signs it; the import side checks the signature first (is this manifest really from us?) and then the hashes (is every file exactly what the manifest says?). Executed in a debian:12-slim container with GnuPG, using a throwaway key (in a real segment the import server holds only the public key, the hashes are shortened here, and the gpg: Signature made ... lines are omitted). The last step shows that after the manifest is rewritten to match the tampered file the checksums pass again, so only the signature check catches it:
$ sha256sum kb-a.cab kb-b.cab > MANIFEST.sha256
2e5ba9e05fd7be2e... kb-a.cab
53bb85fa914a82c5... kb-b.cab
$ gpg --detach-sign --output MANIFEST.sha256.sig MANIFEST.sha256
--- import side: verify the signature, then the checksums
gpg: Good signature from "ingest-signer" [ultimate]
kb-a.cab: OK
kb-b.cab: OK
--- one byte changed in transit
gpg: Good signature from "ingest-signer" [ultimate]
kb-a.cab: OK
kb-b.cab: FAILED
sha256sum: WARNING: 1 computed checksum did NOT match
--- manifest edited to match the tampered file
kb-a.cab: OK
kb-b.cab: OK
gpg: BAD signature from "ingest-signer" [ultimate]
In the middle step the signature still verifies, because the manifest itself was not touched and only the checksum catches the changed file. The last step is why the signature exists: an attacker who swaps a file can also rewrite the manifest hashes, but cannot re-sign it without the private key. The same hash check on a Windows import server, run for real with PowerShell 7 against the same manifest layout (the macOS host produced the manifest with shasum -a 256):
foreach ($line in Get-Content .\MANIFEST.sha256) {
$expected, $name = $line -split '\s+\*?', 2
$actual = (Get-FileHash -Path $name -Algorithm SHA256).Hash
'{0}: {1}' -f $name, $(if ($actual -eq $expected) { 'OK' } else { 'FAILED' })
}
kb-a.cab: OK
kb-b.cab: OK
kb-a.cab: OK
kb-b.cab: FAILED
The first two lines are the intact bundle and the last two are the same bundle after one file was changed. Get-FileHash uses SHA256 by default.
Terms used in this section. A one-way transfer device (a data diode) is hardware that lets data pass in one direction only, so nothing can flow back out of the protected network. Separation of duties means no single person can complete a risky chain alone. A break-glass account is an emergency administrator account whose credentials are sealed and whose every use raises an alert. A privileged access workstation is a hardened computer used only for administration, with no email or browsing.
Phased rollout
Define rings by cumulative share of the 20,000 endpoints. The soak times are example policy.
| Ring | Endpoints | Cumulative endpoints | Cumulative share | Gate before the next ring |
|---|---|---|---|---|
| 0: test lab and IT staff | 200 | 200 | 1% | Install success, no failed services, 2-day soak |
| 1: pilot users across regions | 800 | 1,000 | 5% | Help-desk ticket rate not above baseline, 3-day soak |
| 2: early production | 4,000 | 5,000 | 25% | Same checks, 3-day soak |
| 3: everyone else, one region per wave | 15,000 | 20,000 | 100% | Compliance report shows target reached |
The ring sizes sum to 20,000. Ring 3 is released region by region so a bad interaction in one region does not hit all four at once. Air-gapped endpoints are assigned to rings like everyone else; their updates arrive later because the bundle must be exported and carried, so build that delay into their soak schedule instead of treating them as an exception.
A kill switch matters as much as the rings: the approver role can unapprove or decline an update, which stops further endpoints from receiving it. Endpoints that already installed it need a separate remediation plan (rollback or a fix update).
Access control
- Four roles, held by different people or groups: publisher (ingest and import), approver (decides what each ring may install), deployer (schedules and monitors rollouts), auditor (read-only access to reports). Nobody holds publisher and approver together for production rings, and a second approver confirms an out-of-cycle emergency release.
- Administration goes through privileged access: multi-factor authentication, administrative workstations that are not used for email and browsing, and named accounts, so every action traces to a person.
- Break-glass account: sealed credentials for when the identity system is down, with an alert on any use.
- Air-gap import is a two-person step: one person verifies the bundle's signature and checksums against the manifest, a second confirms and imports. Record both names and the bundle id.
Evidence the platform produces
- Per endpoint: installed, missing or unknown for each required update, with the timestamp of the last report.
- Per update: who approved it, for which rings, and when each ring was released.
- Exceptions: endpoint, reason, approver and expiry date.
- Air-gap: bundle id, contents manifest with checksums, export date, import date, and the people involved.
- Reports: "percentage of endpoints compliant within the policy deadline" computed from raw records. For air-gapped segments, state reports leave the segment the same slow way content enters, so their numbers lag by the transfer interval. Report the age of the last import for each segment so a stale segment is visible.
Capacity check
If every endpoint downloads the same 500 MB set, 20,000 endpoints pull 20,000 x 500 MB = 10,000,000 MB, about 10 TB. Of that, the 18,500 connected endpoints (9,250,000 MB, about 9.25 TB) are served by the regional servers from local networks, and the 1,500 air-gapped endpoints (750,000 MB, about 0.75 TB) are served by their segment distribution servers from the imported bundle, which crosses the gap once per segment by media. The wide-area links between the ingest zone and the four regions carry one copy per region: 4 x 500 MB = 2,000 MB, about 2 GB. Without the regional layer the central site would serve all 9.25 TB of connected demand across the wide area. That is the reason regional servers exist: each update is downloaded once per region over the expensive long-distance links (2 GB in total) and then copied to endpoints over cheap local networks (the 9.25 TB), instead of 9.25 TB crossing the long-distance links. The real figure depends on your update sizes, how many operating system versions you run (each needs its own content) and how much of the payload the endpoints already hold.
Failure modes and trade-offs
| Failure | What happens | Mitigation |
|---|---|---|
| Regional server down | Endpoints in that region cannot fetch content | Second server per region; compliance report shows "unknown", not "ok" |
| Bad update approved | Breakage spreads ring by ring | Rings plus gates, kill switch, rings never skip |
| Bundle corrupted or tampered with in transit | Import fails verification | Reject the bundle, re-export; signature and checksum checks are mandatory |
| Air-gapped segment falls behind | Known vulnerabilities stay open longer | Fixed transfer cadence, an alert on segment age, emergency bundle process |
| Approver account compromised | Attacker can approve a malicious update | Separate publisher and approver, second approver for emergencies, alerts on approval events |
Choosing between one central approval team and regional teams: central approval gives consistency and a single audit trail; regional approval is faster for local issues. Recommendation for this scale: central approval of what is released and when, with regional operators allowed to pause their own region. That keeps one answer to "who approved this" while letting a region stop a rollout that is hurting it.
Write the Ansible tasks that enforce your SSH daemon hardening settings idempotently, restart the service only when something actually changed, and confirm the result before the play can strand you.
Sample Answer
Direct answer
Put the settings in a drop-in file that sshd reads first, write it with the copy module and a validate: command so a bad file never replaces a good one, attach a handler that reloads the service so it runs only when the file really changed, then force that handler to run immediately and prove, from the controller and over a second, brand-new SSH connection, that the automation account can still log in while Ansible's original connection is held open as the way back. Before touching anything, refuse to continue if the account Ansible logs in with has no authorized key, since turning passwords off is the step that strands you. Run the play in small batches that stop on the first failure.
Terms used in the play
- Controller: the machine you run
ansible-playbookfrom. Inventory: the list of hosts and their connection details;linux_serversbelow is a group name from it. - Play and task: a play maps a set of hosts to an ordered list of tasks; a task calls one module (a unit of work such as
copy,statorcommand). - Handler: a task that runs only when another task tells it (
notify) that something changed. ansible_user: the inventory variable naming the account Ansible logs in to each host with.delegate_to: localhost: run this one task on the controller instead of on the target host.- SSH connection reuse: by default Ansible opens one SSH connection to each host and keeps it open for the whole play (the SSH options
ControlMasterandControlPersist;ControlPathis the socket file that shares it), so later tasks do not log in again.
The play
---
- name: Harden the SSH daemon without stranding the play
hosts: linux_servers
become: true
gather_facts: true
serial: [1, "25%"]
max_fail_percentage: 0
vars:
sshd_service: "{{ 'ssh' if ansible_facts['os_family'] == 'Debian' else 'sshd' }}"
sshd_dropin: /etc/ssh/sshd_config.d/00-hardening.conf
sshd_settings:
PermitRootLogin: "no"
PasswordAuthentication: "no"
KbdInteractiveAuthentication: "no"
PubkeyAuthentication: "yes"
X11Forwarding: "no"
MaxAuthTries: "3"
LoginGraceTime: "30"
tasks:
- name: Look up the account this play logs in with
ansible.builtin.getent:
database: passwd
key: "{{ ansible_user }}"
- name: Check that account has an authorized_keys file before passwords go away
ansible.builtin.stat:
path: "{{ getent_passwd[ansible_user][4] }}/.ssh/authorized_keys"
register: login_keys
- name: Stop here if key login is not already possible
ansible.builtin.assert:
that:
- login_keys.stat.exists
- login_keys.stat.size > 0
fail_msg: "{{ ansible_user }} has no authorized_keys; refusing to disable password login"
- name: Apply the hardening drop-in and prove it works
block:
- name: Write the drop-in (sshd -t checks the candidate before it replaces the file)
ansible.builtin.copy:
dest: "{{ sshd_dropin }}"
content: |
# Managed by Ansible. Local edits are overwritten.
{% for key, value in sshd_settings.items() %}{{ key }} {{ value }}
{% endfor %}
owner: root
group: root
mode: "0644"
validate: /usr/sbin/sshd -t -f %s
notify: Reload sshd
- name: Read the effective configuration sshd will use
ansible.builtin.command: /usr/sbin/sshd -T
register: sshd_effective
changed_when: false
check_mode: false
- name: Fail if any setting is overridden by an earlier file
ansible.builtin.assert:
that:
- (item.key | lower ~ ' ' ~ item.value | lower) in sshd_effective.stdout_lines
fail_msg: "{{ item.key }} is not {{ item.value }} in the effective config"
quiet: true
loop: "{{ sshd_settings | dict2items }}"
- name: Reload sshd now, not at the end of the play
ansible.builtin.meta: flush_handlers
- name: Prove a brand-new login works while the old connection is still open
ansible.builtin.command:
cmd: >-
ssh {{ ansible_ssh_common_args | default('') }}
-o BatchMode=yes -o ControlMaster=no -o ControlPath=none -o ConnectTimeout=10
{{ ('-i ' ~ ansible_ssh_private_key_file) if ansible_ssh_private_key_file is defined else '' }}
-p {{ ansible_port | default(22) }}
{{ ansible_user }}@{{ ansible_host | default(inventory_hostname) }} true
delegate_to: localhost
become: false
changed_when: false
check_mode: false
rescue:
- name: Remove our drop-in so the host returns to its previous behaviour
ansible.builtin.file:
path: "{{ sshd_dropin }}"
state: absent
notify: Reload sshd
- name: Reload sshd with the drop-in gone
ansible.builtin.meta: flush_handlers
- name: Fail this host loudly
ansible.builtin.fail:
msg: "SSH hardening failed on {{ inventory_hostname }} and was rolled back"
handlers:
- name: Reload sshd
ansible.builtin.service:
name: "{{ sshd_service }}"
state: reloaded
Reading the play in order
- Header.
become: trueruns tasks as root through sudo.serial: [1, "25%"]andmax_fail_percentage: 0set the batching, explained below.varsholds the service name (sshorsshdby OS family), the drop-in path and the settings as a dictionary. getentasks the host's account database for the login account. On a test host with the account set torootit returned the list["x", "0", "0", "root", "/root", "/bin/bash"]: the password placeholder, user id, group id, description, home directory and shell. Index 4 counts from zero, sogetent_passwd[ansible_user][4]is the home directory (/rootthere), and the next task looks for.ssh/authorized_keysinside it.statandassertcheck that file exists and is not empty, and stop the play before any change if not.block: everything inside is attempted as a unit; if any task in it fails, therescuelist runs instead.copywith the template: the line{% for key, value in sshd_settings.items() %}{{ key }} {{ value }}followed by{% endfor %}loops over the dictionary and writes oneKey valueline per setting. Rendered on a test host with two settings it produced exactly these lines after the comment:PermitRootLogin noandMaxAuthTries 3.validatechecks the candidate file, andnotifyqueues the handler only if the file changed.sshd -Tprints the effective configuration with lower-case keywords.changed_when: falsestops this read-only command being reported as a change, andcheck_mode: falsemakes it run even under--check.- The
assertloop runs once per setting. ForPermitRootLogin: "no"the expressionitem.key | lower ~ ' ' ~ item.value | lowerbuilds the textpermitrootlogin no(~joins strings,| lowerlower-cases), and the assertion is that this exact line appears in the list of output lines fromsshd -T. If an earlier file setPermitRootLogin yes, the linepermitrootlogin yesappears instead, the text is missing, and the task fails. flush_handlersruns the queued reload now. The login test runs thesshcommand on the controller.rescueundoes the change and fails the host.
Why each piece is there
- Drop-in, not an edit of
sshd_config. The OpenSSH manual says anIncludepulls in files "in lexical order" and that "for each keyword, the first obtained value will be used". The distributionsshd_configincludes/etc/ssh/sshd_config.d/*.confnear the top (line 12 on the Ubuntu 24.04 test image, line 15 on AlmaLinux 9), so a file named00-hardening.confis read before the vendor's own drop-ins (AlmaLinux ships50-redhat.conf) and wins. Managing a whole file we own is also what makes the task idempotent (running it a second time changes nothing): the same content produces no change. validate: /usr/sbin/sshd -t -f %s.copywrites the candidate to a temporary path, runs the command with%sreplaced by that path, and replaces the real file only if it exits 0.sshd -tchecks "the validity of the configuration file and sanity of the keys". An invalid value never reaches/etc/ssh.- Why validation is not enough: the
sshd -Tassertion. The temporary file is checked alone, so a conflict with another file is invisible to it.sshd -Tprints the effective configuration (lower-case keyword, then value), and theassertcompares every setting we declared with whatsshdwould really use. It reads files from disk, so it works before the reload and is the guard against the "an earlier file wins" case. - Both
PasswordAuthenticationandKbdInteractiveAuthentication. They are separate switches (the second is the current name ofChallengeResponseAuthentication, a deprecated alias), and both default toyes, so disabling only one leaves a way to be asked for a password. notifyplusmeta: flush_handlers. A handler runs only if the notifying task reportedchanged, which gives "restart only when something changed". The default is to run handlers at the end of the play; flushing runs it now, so the proof happens while the play still has a working connection to fall back on.state: reloadedmakessshdre-read its configuration (the manual describes SIGHUP, the hang-up signal, as makingsshdre-execute itself with its original options) rather than stopping the daemon.- A fresh login from the controller, with the old connection still open. Ansible keeps one SSH connection open and reuses it (the
ControlPersistoption keeps that connection alive between tasks), so a "ping" after the reload would run over the old, already authenticated session and prove nothing. Thecommandtask runs a separatesshprocess on the controller (delegate_to: localhost,ControlPath=noneso it cannot share the master connection (the long-lived connection Ansible opened),BatchMode=yesso it fails instead of prompting) using the same host, user, port, key and extra SSH arguments as the inventory. Because it does not touch Ansible's own connection, that connection is still alive if the new login is refused, and therescuesection can still reach the host to undo the change. (The alternative,meta: reset_connectionfollowed bywait_for_connection, also proves a fresh login, but if the login fails there is no connection left to roll back over.) block/rescue. Any failure inside the block (validation, effective-config mismatch, failed new login) removes our drop-in, reloads, and fails the host loudly, so the host goes back to its previous behaviour instead of staying half-configured.serial: [1, "25%"]andmax_fail_percentage: 0. The first batch is one host (the canary); later batches are 25% of the inventory.max_fail_percentageapplies per batch, and at 0 any failure in a batch stops later batches. I verified that: with four hosts,serial: [1, "50%"]and one failure in the second batch, the fourth host never ran withmax_fail_percentage: 0and did run without it.- Service name.
sshon Debian and Ubuntu,sshdon the RHEL family; thesshd_servicevariable picks one from the OS family.
Results from a test run
The play ran against a real sshd over real SSH (a key-authenticated account with sudo, on 127.0.0.1) in an Ubuntu 24.04 container. That container has no systemd, so the reload handler used the init-script service manager there; on a systemd host the same task uses systemd.
RUN 1: PLAY RECAP target : ok=9 changed=2 failed=0 rescued=0
RUN 2: PLAY RECAP target : ok=8 changed=0 failed=0 rescued=0 (no handler ran)
$ sshd -T | grep -E '^(permitrootlogin|passwordauthentication|maxauthtries|logingracetime|x11forwarding) '
logingracetime 30
maxauthtries 3
permitrootlogin no
passwordauthentication no
x11forwarding no
Before the first run, the same host reported logingracetime 120, maxauthtries 6, passwordauthentication yes, x11forwarding yes.
Four deliberate failures:
| Injected fault | What happened |
|---|---|
PermitRootLogin maybe | copy failed with failed to validate and unsupported option "maybe"; no file was written; rescue ran; host failed |
a file 00-aaa-cloud.conf containing PasswordAuthentication yes | the assert failed on that setting; rescue removed our drop-in and reloaded; host failed |
a valid setting that locks the automation account out (AllowUsers nobody) | validation and the effective-config assert both passed, the reload ran, the fresh login from the controller was refused, and rescue then removed the drop-in and reloaded over the old connection; a re-run afterwards logged in normally |
the login account's authorized_keys removed | the pre-check assertion stopped the play before any change: deploy has no authorized_keys; refusing to disable password login |
Complexity and edge cases
- Work per host is constant (a handful of tasks, one file, one
sshd -T). Wall-clock time depends on the batch plan: for 100 hosts,serial: [1, "25%"]gives batches of 1, 25, 25, 25 and 24, which is five batches. - The login test proves the automation account can log in the way the inventory says it does. A human using a different key, or a
Matchrule that affects only some source addresses, is a different path; for hosts with such rules, add a second test task for that path. If the controller reaches hosts through a bastion or a different connection plugin, thesshcommand must be adapted to match (a bastion inansible_ssh_common_argscarries over). - The authorized-keys check proves a key file exists, not that the key matches the one in use. The login that follows the reload is the real proof; the pre-check only prevents the obvious foot-gun.
- A host with a console or out-of-band access (cloud serial console, hypervisor console) is the recovery path if all else fails; know where it is before the run.
- Settings not included here (
AllowGroups, ciphers and MACs) are deliberate omissions: each can lock out legitimate users or break older clients, so each gets its own tested change.
Design an automated canary patch deployment for a server fleet. Which health signals gate promotion, what triggers an automatic halt or rollback, and how do monitoring and orchestration tooling cooperate to do it?
Sample Answer
Direct answer
Patch in rings: a small canary ring first, then progressively larger rings, with an automated gate between rings. A gate passes only when positive evidence says the patched hosts are healthy compared with unpatched control hosts running the same traffic, and "no data" counts as a failure. Any failed gate halts the rollout automatically: no new hosts are patched and no later ring is started. A failure in the canary ring also triggers an automatic rollback of that ring. A failure found after a larger ring has been patched leaves that ring in place and hands the rollback decision to the change owner, because an automatic mass rollback is itself a risk. The monitoring system holds the health definitions, the orchestration tool holds the sequence, and the orchestrator queries monitoring before every promotion, so the pipeline fails closed if monitoring is unavailable.
Plain meanings of the terms
A soak is the waiting period after a ring is patched, during which you watch for problems. A gate fails closed when it answers "stop" whenever it is unsure, like a door lock that engages when the power goes out; failing open would let the rollout continue while monitoring is blind. Saturation means how full a resource is, for example queue length or connections in use out of the pool size. A synthetic transaction is a scripted fake user (log in, read a known record) run from outside the fleet.
Ring design (worked example for a 200-host fleet)
| Ring | Hosts in ring | Cumulative hosts | Share of fleet | Gate before the next ring |
|---|---|---|---|---|
| 0 canary | 2 | 2 | 1% | health gates clean for the soak period, including one traffic peak |
| 1 | 18 | 20 | 10% | same gates, plus a business-owner check on key transactions |
| 2 | 80 | 100 | 50% | same gates, plus an approval from the change owner |
| 3 | 100 | 200 | 100% | final health report |
Check: 2 + 18 + 80 + 100 = 200. Choose the canary hosts to be representative, not convenient: include each operating system and major version, each hardware or VM type, each data center or availability zone, and at least one stateful node. Canary hosts picked from a dev environment prove nothing about production. Soak times are your decision (for example one business day for ring 0 so that a traffic peak is included), tuned to how long your failure modes take to show up.
Which health signals gate promotion
Compare the patched ring with a control group (hosts of the same role that are not patched yet) over the same period, instead of comparing with a fixed number. Absolute thresholds break on time of day and on seasonal load.
- Host health: the host came back after reboot and reports metrics, no failed system services, disk and memory in a normal range relative to control.
- Service health: request error ratio (5xx), latency at the 99th percentile, and saturation (queue depth, connection pool use), canary against control.
- Dependency health for stateful tiers: replication lag, cluster membership, queue backlog.
- A synthetic transaction that exercises the real user path, run from outside.
- Evidence of traffic: the gate is blind if the canary serves no requests, so the absence of data must fail the gate.
Signals 1, 2 and 5 (host reporting, error ratio against control, proof of traffic) form the minimum gate for any ring; signal 3 adds coverage for stateful tiers and signal 4 for user-facing paths, so add them where the fleet has those tiers.
Gate rules in Prometheus alerting-rule syntax. The ring label, http_requests_total and its code label are application and relabeling choices that you define; up is generated by Prometheus itself for every scrape target, and node_systemd_unit_state comes from the node exporter (the host metrics agent) when its systemd collector is enabled.
groups:
- name: patch-canary-gate
rules:
- alert: CanaryHostDown
expr: up{job="node", ring="canary"} == 0
for: 5m
labels:
severity: halt
annotations:
summary: "Canary host {{ $labels.instance }} stopped reporting after patching"
- alert: CanaryServiceFailed
expr: node_systemd_unit_state{ring="canary", state="failed"} == 1
for: 2m
labels:
severity: halt
annotations:
summary: "{{ $labels.name }} failed on canary host {{ $labels.instance }}"
# Silence must not read as health: no canary traffic means the gate is blind.
- alert: CanaryNoTraffic
expr: |
sum(rate(http_requests_total{ring="canary"}[10m])) < 1
or absent(http_requests_total{ring="canary"})
for: 10m
labels:
severity: halt
annotations:
summary: "Canary ring is receiving under 1 request per second or has no request metric"
- alert: CanaryErrorRateRegression
expr: |
(
sum(rate(http_requests_total{ring="canary", code=~"5.."}[10m]))
/ sum(rate(http_requests_total{ring="canary"}[10m]))
)
>
(
2 * (
(sum(rate(http_requests_total{ring="control", code=~"5.."}[10m])) or vector(0))
/ sum(rate(http_requests_total{ring="control"}[10m]))
) + 0.005
)
for: 10m
labels:
severity: halt
annotations:
summary: "Canary 5xx ratio is more than double the unpatched control plus 0.5 points"
Reading the rules one idea at a time:
up{job="node", ring="canary"} == 0:upis 1 when the last scrape of a target worked and 0 when it failed. The braces filter to canary targets.for: 5mmeans the condition must stay true for five minutes in a row before the alert fires; until then it is "pending".node_systemd_unit_state{..., state="failed"} == 1: the node exporter reports one series per systemd unit and state, and the value 1 means the unit is in that state right now.rate(http_requests_total{ring="canary"}[10m]):http_requests_totalis a counter that only ever goes up, andrateover a 10-minute window turns it into requests per second.sum(...)adds all canary series (every host and status code) together, and< 1means fewer than one request per second.absent(http_requests_total{ring="canary"}): returns 1 only when the selector matches no series at all, which is how "the metric vanished" becomes something a rule can detect.code=~"5..":=~is a regular-expression match, and5..means a 5 followed by any two characters, so 500, 502, 503 and so on.- The regression rule divides the canary 5xx rate by the canary total rate (the canary error ratio) and compares it with
2 * (control error ratio) + 0.005. With the numbers used in the test below, canary 5xx is 60 per minute out of 660 requests per minute in total, so 60 / 660 = 9.09 percent. Control is 6 out of 606, so 0.99 percent, and the threshold is 2 x 0.0099 + 0.005 = 0.0248, or 2.48 percent. 9.09 is above 2.48, so the alert fires; a canary that matches control (0.99 percent) is below it and stays quiet. The control numerator is wrapped inor vector(0)on purpose: a counter series that has never recorded an error may not exist yet, and without the fallback a missing control 5xx series makes the whole right-hand side empty, so the comparison returns nothing and the alert can never fire while the canary is failing. With the fallback, a control with no 5xx series counts as an error ratio of zero and the threshold falls back to the 0.5-point floor. severity: haltis only a label name chosen for this design, and Prometheus does nothing with it. The orchestrator has to be wired to read it: before each promotion it asks the Prometheus HTTP API for firing alerts, for example the queryALERTS{severity="halt", alertstate="firing"}(Prometheus stores every pending or firing alert as that synthetic series), and it stops if any row comes back.
Design points: the regression rule fires when the canary 5xx ratio is more than twice the control ratio plus half a percentage point (the absolute floor stops noise at very low error rates); for: 10m prevents one bad scrape from halting a rollout; the no-traffic rule exists because zero requests divided by zero requests is NaN (not a number), a comparison against NaN is false, and a query over a missing series returns nothing, so without that rule silence would look healthy.
Test series notation: '0+600x40' means start at 0 and add 600, forty times. With interval: 1m each step is one minute, so the canary 200 counter rises by 600 per minute (10 requests per second), the canary 500 counter by 60 per minute, and the control 500 counter by 6 per minute. eval_time: 30m asks what the rule says 30 minutes in, long after the 10-minute for has been satisfied.
Executed check: the rule file passes promtool check rules, and promtool test rules runs the unit test below with pinned input series (the counters rise by a fixed amount per minute, so the results are reproducible). A canary whose error ratio is about 9% against a control near 1% fires the regression alert, a canary matching control stays quiet, a canary with no request metric trips the no-traffic alert, and a fourth case with a control that has no 5xx series at all still fires the regression alert (the same rule without or vector(0) returns nothing for that case). Running the same test with the floor changed from 0.005 to 5 made the regression case fail, which shows the test can detect a broken rule.
rule_files:
- patch-canary-gate.rules.yml
evaluation_interval: 1m
tests:
- name: regression on canary fires
interval: 1m
input_series:
- series: 'http_requests_total{ring="canary", code="200"}'
values: '0+600x40'
- series: 'http_requests_total{ring="canary", code="500"}'
values: '0+60x40'
- series: 'http_requests_total{ring="control", code="200"}'
values: '0+600x40'
- series: 'http_requests_total{ring="control", code="500"}'
values: '0+6x40'
alert_rule_test:
- eval_time: 30m
alertname: CanaryErrorRateRegression
exp_alerts:
- exp_labels:
severity: halt
exp_annotations:
summary: "Canary 5xx ratio is more than double the unpatched control plus 0.5 points"
- name: healthy canary stays quiet
interval: 1m
input_series:
- series: 'http_requests_total{ring="canary", code="200"}'
values: '0+600x40'
- series: 'http_requests_total{ring="canary", code="500"}'
values: '0+6x40'
- series: 'http_requests_total{ring="control", code="200"}'
values: '0+600x40'
- series: 'http_requests_total{ring="control", code="500"}'
values: '0+6x40'
alert_rule_test:
- eval_time: 30m
alertname: CanaryErrorRateRegression
exp_alerts: []
- eval_time: 30m
alertname: CanaryNoTraffic
exp_alerts: []
- name: missing canary metrics trip the blindness alert
interval: 1m
input_series:
- series: 'http_requests_total{ring="control", code="200"}'
values: '0+600x40'
alert_rule_test:
- eval_time: 30m
alertname: CanaryNoTraffic
exp_alerts:
- exp_labels:
severity: halt
ring: canary
exp_annotations:
summary: "Canary ring is receiving under 1 request per second or has no request metric"
- name: control with no 5xx series still lets a canary regression fire
interval: 1m
input_series:
- series: 'http_requests_total{ring="canary", code="200"}'
values: '0+600x40'
- series: 'http_requests_total{ring="canary", code="500"}'
values: '0+60x40'
- series: 'http_requests_total{ring="control", code="200"}'
values: '0+600x40'
alert_rule_test:
- eval_time: 30m
alertname: CanaryErrorRateRegression
exp_alerts:
- exp_labels:
severity: halt
exp_annotations:
summary: "Canary 5xx ratio is more than double the unpatched control plus 0.5 points"
promtool check rules patch-canary-gate.rules.yml
promtool test rules patch-canary-gate.test.yml
Checking patch-canary-gate.rules.yml
SUCCESS: 4 rules found
SUCCESS
What triggers a halt, and what a rollback means
- Automatic halt: any alert with
severity="halt"firing during the soak, or the gate query returning no data. The orchestrator stops promoting and also stops the current batch. - Within a ring, patch in batches and stop on the first failure. With Ansible,
serialsets the batch size andmax_fail_percentage: 0aborts the play when any host in a batch fails. Executed with three hosts andserial: 1, where the second host was made to fail (the playbook and the command are below), the third host was never run:
- hosts: web
gather_facts: false
serial: 1
max_fail_percentage: 0
tasks:
- ansible.builtin.fail:
msg: simulated failure
when: inventory_hostname == 'web02'
ansible-playbook -i inv.ini fail.yml
PLAY [web] *********************************************************************
TASK [ansible.builtin.fail] ****************************************************
skipping: [web01]
PLAY [web] *********************************************************************
TASK [ansible.builtin.fail] ****************************************************
fatal: [web02]: FAILED! => {"changed": false, "msg": "simulated failure"}
NO MORE HOSTS LEFT *************************************************************
NO MORE HOSTS LEFT *************************************************************
PLAY RECAP *********************************************************************
web01 : ok=0 changed=0 unreachable=0 failed=0 skipped=1 rescued=0 ignored=0
web02 : ok=0 changed=0 unreachable=0 failed=1 skipped=0 rescued=0 ignored=0
The output above is complete and unedited (ansible-core 2.16.3). ok=0 for web01 is because fact gathering is turned off in this test, and NO MORE HOSTS LEFT is printed twice by this version.
(web03 does not appear in the recap at all.)
- Halt and rollback are different. Halt is cheap and automatic. Rollback is only clean for some changes: a package downgrade or a Windows update uninstall is usually possible, but a kernel update, a schema change made by a package, or a patch applied to a database node may not be reversible. Prefer rollback by replacement: for virtual machines and cloud instances that are rebuilt from images, the canary hosts are reprovisioned from the previous image; for hosts that are not, pin or downgrade the packages and take the host out of rotation until it is verified.
- Do not let monitoring alone patch or roll back anything. It signals; the orchestrator, which knows about batch state, acts.
How monitoring and orchestration cooperate
- The pipeline starts a ring and writes a marker (labels or a tag) so that monitoring knows which hosts are canary and which are control.
- After the soak, the orchestrator queries monitoring for the gate result and requires two things: zero firing halt alerts and evidence that the required metrics exist. Pulling the result is safer than waiting for an alert to be pushed to a webhook, because a broken push channel looks identical to a healthy fleet.
- A pass promotes to the next ring. A fail halts and notifies the owners. If the failing ring is the canary ring, the orchestrator also starts its rollback; for a later ring the change owner decides.
- Every decision is logged with the query result that justified it.
Concrete tooling (two roles have to be filled: something that runs the sequence, and something that holds the gate definitions; the names are examples and any pipeline runner fills the first role): a pipeline stage per ring in Azure DevOps (or GitHub Actions or Jenkins), with an approval step between ring 2 and ring 3; Ansible playbooks run through the automation controller (the product formerly called Ansible Tower) for the host work; Prometheus with Alertmanager for gates; for Kubernetes worker nodes, kubectl drain --ignore-daemonsets <node> (add --delete-emptydir-data for pods with emptyDir volumes), patch the node, then kubectl uncordon <node>. Drain respects PodDisruptionBudgets (Kubernetes objects that limit how many pods of a workload may be down at once), so it will not evict pods if that would drop a workload below its minAvailable.
Draining stateful and load-balanced hosts
Remove a host from the load balancer and let connections drain before patching, and put it back only after its own health endpoint answers. The play below is the pattern for hosts that sit behind a load balancer; Kubernetes nodes use the drain commands above instead. It was syntax-checked with ansible-playbook --syntax-check (the load balancer URL is a placeholder, so it was not executed):
---
- name: Patch one ring, one batch at a time
hosts: "ring_{{ ring | default('canary') }}"
become: true
serial: "{{ batch | default(1) }}"
max_fail_percentage: 0
tasks:
- name: Disable the host in the load balancer pool
ansible.builtin.uri:
url: "https://lb.example.internal/api/pools/{{ pool }}/members/{{ inventory_hostname }}"
method: PUT
body_format: json
body: { enabled: false }
status_code: 200
delegate_to: localhost
become: false
- name: Wait for in-flight connections to drain
ansible.builtin.pause:
seconds: 60
- name: Apply the patch set
ansible.builtin.package:
name: "*"
state: latest
- name: Wait for the application health endpoint on the patched host
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/healthz"
status_code: 200
register: health
retries: 12
delay: 10
until: health.status == 200
delegate_to: localhost
become: false
- name: Enable the host in the load balancer pool again
ansible.builtin.uri:
url: "https://lb.example.internal/api/pools/{{ pool }}/members/{{ inventory_hostname }}"
method: PUT
body_format: json
body: { enabled: true }
status_code: 200
delegate_to: localhost
become: false
For stateful nodes, patch one replica at a time, wait for replication lag to return to normal before the next, and fail over a primary deliberately (patch the standby first, fail over, then patch the former primary) instead of rebooting it under load.
Multi-data-center ordering, owners and change windows
- Patch one data center or availability zone at a time, starting with the least critical (or the standby side of an active/standby pair). Never patch both sides of a redundant pair in one window.
- Publish the ring calendar so that application owners know when their hosts move. Owners can request a documented exemption with an expiry date, which keeps exceptions visible and temporary. Respect freeze periods (month-end, peak trading), and have the on-call engineer for each owner reachable during their ring.
- Communicate three times: ahead of the window (what, when, what to watch), when each ring starts, and when it finishes with the gate results.
Mixed Windows and Linux
- Same gates, different mechanics. Linux: patch with the package manager from the orchestrator, reboot only when a kernel or core library changed. Windows: patching goes through the update service your estate uses (WSUS, which Microsoft documents as deprecated: no new features, still supported with security and quality updates; or a cloud update service such as Azure Update Manager), reboot windows are explicit, and installs can take longer, so the soak and window sizes are set per operating system.
- Validation tests per OS after the patch: the service starts, the listener answers, the health endpoint returns 200, a synthetic transaction passes, and the host reports metrics again.
Trade-offs
- A bigger canary finds rarer problems but exposes more users; two hosts out of 200 is 1%, which is usually too small to see an error-rate difference at low traffic. That is why the no-traffic rule and a traffic floor matter, and why a canary should be at least as large as needed to receive a meaningful share of requests.
- Comparing against control needs unpatched control hosts at the same time, which is one more reason to patch in rings and not all at once.
A hardened baseline keeps drifting once servers are in production. Design how you would detect drift against it across Windows and Linux hosts, and how you decide per setting between reporting only and fixing it automatically.
Sample Answer
Direct answer
Make the baseline a single versioned source of truth, check it at three points (image build, provisioning, and continuously at runtime), and decide per setting between report-only and auto-fix with a rule: auto-fix only when the correction is deterministic (the same input always gives the same result, with no per-host judgement), idempotent, reversible and cannot take a service or an administrator out; report and ticket everything else. Keep the scan results as dated, tamper-evident evidence (any later edit would be visible) so an auditor can see the history, not only today's state.
Design: three checkpoints and evidence
- Image build: scan every image before it is published and fail the build on violations. On Linux, OpenSCAP (an open-source compliance scanner) can evaluate a profile and write results and an HTML report with
oscap xccdf eval --profile <profile> --results-arf results.xml --report report.html <datastream>. A profile is a named subset of rules, such as a CIS level; a datastream is the single file that bundles the rules and the checks that test them; XCCDF is the XML format those rules are written in and results are reported in; on Windows, build from a hardened image and apply the settings through Group Policy. - Provisioning: enforce at first boot with configuration management (Ansible, Group Policy or an endpoint management tool) so no host starts life drifted.
- Runtime: re-check continuously (for example daily) in detect mode. Ansible can run read-only with
--check(and--diff), where modules that support check mode report the changes they would make. Note the documented limits: modules without check-mode support report nothing, and tasks whose conditions depend on registered results from earlier tasks produce no output in check mode. - Evidence: store each scan result (the XML results file or the exported report) with the host, time, baseline version and operator, in write-once storage (storage that accepts new files but refuses edits and deletions, for example object storage with a retention lock). This is what an auditor asks for.
Deciding per setting: report only or fix automatically
| Setting | Mode | Why |
|---|---|---|
| Disable root SSH login | Fix | Deterministic, no application depends on it, reversible |
| Disable X11 forwarding | Fix | Same |
| Limit SSH authentication tries | Report | Rarely harmful but needs review against automation that retries |
| Reverse-path filtering (a kernel check that drops a packet whose reply would leave by a different network interface than it arrived on) | Report | Can break asymmetric routing (traffic that arrives by one path and returns by another); owner must approve per host |
| Require SMB signing on Windows file servers | Report | Legacy clients lose access; needs a rollout plan |
Rules behind the table:
- Fix automatically when the setting is a pure tightening that no workload legitimately needs and the fix is idempotent (running it twice changes nothing).
- Report only when it can interrupt service or lock people out, needs a reboot, varies by host role, or when the same setting is repeatedly reverted by a person or program. Repeated reverts mean something is fighting your baseline, so raise an alert instead of fixing again.
- Every report-only drift gets an owner and a due date, or a time-limited exception with an expiry.
- Examples of fixes that are not deterministic, so they stay report-only: "set
rp_filterto whatever suits this host's routing" (the right value differs per host), "remove local administrators who look unfamiliar" (who is legitimate is a human judgement), and any change that only takes effect after a reboot.
How the pieces fit: the single versioned baseline is what everything is compared against; the three checkpoints decide when you look; the table and rules decide what happens on a mismatch. OpenSCAP and the registry script below are the detectors for Linux and Windows hosts, and Ansible check mode and Group Policy are how you detect or enforce through the tool a fleet is already managed with.
Worked example: Linux check that fixes some settings and reports others
This script reads a table of settings. Rows marked fix are rewritten when they drift. Rows marked report are only flagged. It takes a root directory so it can be tested safely.
#!/usr/bin/env bash
# Report-or-fix drift check for key/value settings in config files.
# Usage: drift-check.sh ROOT (ROOT is / on a real host)
set -euo pipefail
root=${1:-/}
# file | key | expected value | mode (report|fix)
settings=(
"etc/ssh/sshd_config|PermitRootLogin|no|fix"
"etc/ssh/sshd_config|X11Forwarding|no|fix"
"etc/ssh/sshd_config|MaxAuthTries|4|report"
"etc/sysctl.d/99-hardening.conf|net.ipv4.conf.all.rp_filter|1|report"
)
drift=0
for row in "${settings[@]}"; do
IFS='|' read -r file key want mode <<<"$row"
path="$root/$file"
have=$({ sed -E 's/[[:space:]]*=[[:space:]]*/ /' "$path" 2>/dev/null || true; } |
awk -v k="$key" '$1 == k { v = $2 } END { print v }')
if [[ "$have" == "$want" ]]; then
printf 'OK %-40s %s\n' "$key" "$have"
continue
fi
drift=$((drift + 1))
printf 'DRIFT %-40s have=%s want=%s mode=%s\n' "$key" "${have:-<unset>}" "$want" "$mode"
if [[ "$mode" == fix ]]; then
if [[ ! -f "$path" ]]; then
printf 'SKIP %-40s not fixed, %s does not exist\n' "$key" "$file"
else
if grep -qE "^[[:space:]]*${key}[[:space:]]" "$path"; then
sed -i -E "s|^[[:space:]]*${key}[[:space:]].*|${key} ${want}|" "$path"
else
printf '%s %s\n' "$key" "$want" >>"$path"
fi
printf 'FIXED %-40s now=%s\n' "$key" "$want"
fi
fi
done
echo "drift_count=$drift"
exit $(( drift > 0 ))
Driver and real output (run inside a Debian container with GNU sed):
rm -rf /tmp/root && mkdir -p /tmp/root/etc/ssh /tmp/root/etc/sysctl.d
printf 'PermitRootLogin yes\nMaxAuthTries 6\n' > /tmp/root/etc/ssh/sshd_config
printf 'net.ipv4.conf.all.rp_filter = 0\n' > /tmp/root/etc/sysctl.d/99-hardening.conf
bash drift-check.sh /tmp/root; echo "exit=$?"
bash drift-check.sh /tmp/root; echo "exit=$?"
DRIFT PermitRootLogin have=yes want=no mode=fix
FIXED PermitRootLogin now=no
DRIFT X11Forwarding have=<unset> want=no mode=fix
FIXED X11Forwarding now=no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=4
exit=1
OK PermitRootLogin no
OK X11Forwarding no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=2
exit=1
Reading the script, line by line:
settings=( ... )is a table with one row per setting; the four fields in each row are separated by|.IFS='|' read -r file key want mode <<<"$row"splits one row on|into four variables (IFSis the field separatorreaduses,-rstops backslashes being treated as escapes, and<<<"$row"feeds the row toreadas its input). For the first row that givesfile=etc/ssh/sshd_config key=PermitRootLogin want=no mode=fix.- The
have=$(...)line finds the current value. The{ sed ... || true; }group matters: withset -euo pipefail, a missing config file would makesedfail and silently end the whole script with status 2 and no output. With the guard, a missing file simply reads as<unset>and is reported as drift (run on a root whosesshd_configalready meets the baseline and which has no sysctl file, it printsDRIFT net.ipv4.conf.all.rp_filter have=<unset> want=1 mode=reportanddrift_count=1). The same guard does not help afixrow, because the rewrite needs a file to edit: without the[[ ! -f "$path" ]]test in the fix branch, a root with nosshd_configmadegrepand then the>>append fail and the whole run stop after the first DRIFT line, with nodrift_countand exit status 1, which looks the same as ordinary drift. With the test, that row printsSKIP PermitRootLogin not fixed, etc/ssh/sshd_config does not exist, stays counted as drift, and the run finishes (on an empty root all four rows are drift, two of themSKIP,drift_count=4, exit 1; run in the same container).sed -E 's/[[:space:]]*=[[:space:]]*/ /'turns akey = valueline intokey valueby replacing the first=and the spaces around it with one space (lines already writtenkey valuepass through unchanged).awk -v k="$key" '$1 == k { v = $2 } END { print v }'then looks for lines whose first word equals the key, remembers the second word, and prints the last one remembered when the file ends. Run on a small file in a Debian container:
--- after sed:
PermitRootLogin yes
MaxAuthTries 6
MaxAuthTries 5
# X11Forwarding yes
--- awk for MaxAuthTries:
5
The commented-out X11Forwarding line is ignored because its first word is #, and when a key appears twice the last line wins. That is how sysctl files behave; sshd keeps the first value instead, one more reason the sketch's reading of a file can differ from the effective configuration.
[[ "$have" == "$want" ]]printsOKand moves to the next row; anything else counts as drift, incrementsdriftand printsDRIFT. Only rows with modefixreach the rewrite, and only when the file exists:sed -ireplaces the existing line, orprintf ... >>appends one if the key was missing.exit $(( drift > 0 ))ends the script with status 1 if any drift remains and 0 otherwise, because the arithmetic comparison evaluates to 1 or 0. Run in the same container,drift=0exits 0 anddrift=3exits 1.
The first run fixes two settings and reports two. The second run finds the fixed ones clean (idempotent) while the report-only ones stay open until a human acts. The non-zero exit code lets a scheduler or pipeline alarm on remaining drift.
Limits of this sketch, so you do not trust it blindly: it reads files, not effective configuration. The sshd manual says that for each keyword the first obtained value is used, and that keywords after a Match line apply only to matching connections, so an appended line can be shadowed or captured by a Match block (a section of the configuration whose settings apply only to connections matching a condition, such as a particular user or source address). In production, read the effective SSH configuration with sshd -T, which prints it, rather than grepping the file.
Worked example: Windows registry-backed setting
On Windows the same decision logic applies; prefer Group Policy or endpoint management to enforce, and use a script only for what they do not cover. This reports drift for SMB signing and fixes it only when run with -Fix:
param([switch]$Fix)
$path = 'HKLM:\SYSTEM\CurrentControlSet\Services\LanmanServer\Parameters'
$name = 'RequireSecuritySignature'
$want = 1
$have = Get-ItemPropertyValue -Path $path -Name $name -ErrorAction SilentlyContinue
if ($have -eq $want) { "OK $name = $have"; return }
"DRIFT $name have=$have want=$want"
if ($Fix) {
Set-ItemProperty -Path $path -Name $name -Value $want
"FIXED $name now=$want"
}
Run it on a Windows host, since the registry provider is Windows-only. RequireSecuritySignature is the registry setting Microsoft names for the "always sign" SMB policy. It is deliberately report-only by default.
Trade-offs and pitfalls
- Auto-fix everything and you will eventually take down an application at 3 a.m.; report everything and drift grows faster than people close tickets. The per-setting rule is the compromise.
- Detecting drift is half the job: find out why it happened (manual change, a package script, a stopped agent) or the same drift returns.
- Measure it: drift count per host per week, mean time to close report-only items, and number of exceptions past expiry.
- A scanner finding is only evidence if the scan could authenticate and covered the host.
Unlock Full Question Bank
Get access to all 47 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.