System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
A hardened baseline keeps drifting once servers are in production. Design how you would detect drift against it across Windows and Linux hosts, and how you decide per setting between reporting only and fixing it automatically.
Sample Answer
Direct answer
Make the baseline a single versioned source of truth, check it at three points (image build, provisioning, and continuously at runtime), and decide per setting between report-only and auto-fix with a rule: auto-fix only when the correction is deterministic (the same input always gives the same result, with no per-host judgement), idempotent, reversible and cannot take a service or an administrator out; report and ticket everything else. Keep the scan results as dated, tamper-evident evidence (any later edit would be visible) so an auditor can see the history, not only today's state.
Design: three checkpoints and evidence
- Image build: scan every image before it is published and fail the build on violations. On Linux, OpenSCAP (an open-source compliance scanner) can evaluate a profile and write results and an HTML report with
oscap xccdf eval --profile <profile> --results-arf results.xml --report report.html <datastream>. A profile is a named subset of rules, such as a CIS level; a datastream is the single file that bundles the rules and the checks that test them; XCCDF is the XML format those rules are written in and results are reported in; on Windows, build from a hardened image and apply the settings through Group Policy. - Provisioning: enforce at first boot with configuration management (Ansible, Group Policy or an endpoint management tool) so no host starts life drifted.
- Runtime: re-check continuously (for example daily) in detect mode. Ansible can run read-only with
--check(and--diff), where modules that support check mode report the changes they would make. Note the documented limits: modules without check-mode support report nothing, and tasks whose conditions depend on registered results from earlier tasks produce no output in check mode. - Evidence: store each scan result (the XML results file or the exported report) with the host, time, baseline version and operator, in write-once storage (storage that accepts new files but refuses edits and deletions, for example object storage with a retention lock). This is what an auditor asks for.
Deciding per setting: report only or fix automatically
| Setting | Mode | Why |
|---|---|---|
| Disable root SSH login | Fix | Deterministic, no application depends on it, reversible |
| Disable X11 forwarding | Fix | Same |
| Limit SSH authentication tries | Report | Rarely harmful but needs review against automation that retries |
| Reverse-path filtering (a kernel check that drops a packet whose reply would leave by a different network interface than it arrived on) | Report | Can break asymmetric routing (traffic that arrives by one path and returns by another); owner must approve per host |
| Require SMB signing on Windows file servers | Report | Legacy clients lose access; needs a rollout plan |
Rules behind the table:
- Fix automatically when the setting is a pure tightening that no workload legitimately needs and the fix is idempotent (running it twice changes nothing).
- Report only when it can interrupt service or lock people out, needs a reboot, varies by host role, or when the same setting is repeatedly reverted by a person or program. Repeated reverts mean something is fighting your baseline, so raise an alert instead of fixing again.
- Every report-only drift gets an owner and a due date, or a time-limited exception with an expiry.
- Examples of fixes that are not deterministic, so they stay report-only: "set
rp_filterto whatever suits this host's routing" (the right value differs per host), "remove local administrators who look unfamiliar" (who is legitimate is a human judgement), and any change that only takes effect after a reboot.
How the pieces fit: the single versioned baseline is what everything is compared against; the three checkpoints decide when you look; the table and rules decide what happens on a mismatch. OpenSCAP and the registry script below are the detectors for Linux and Windows hosts, and Ansible check mode and Group Policy are how you detect or enforce through the tool a fleet is already managed with.
Worked example: Linux check that fixes some settings and reports others
This script reads a table of settings. Rows marked fix are rewritten when they drift. Rows marked report are only flagged. It takes a root directory so it can be tested safely.
#!/usr/bin/env bash
# Report-or-fix drift check for key/value settings in config files.
# Usage: drift-check.sh ROOT (ROOT is / on a real host)
set -euo pipefail
root=${1:-/}
# file | key | expected value | mode (report|fix)
settings=(
"etc/ssh/sshd_config|PermitRootLogin|no|fix"
"etc/ssh/sshd_config|X11Forwarding|no|fix"
"etc/ssh/sshd_config|MaxAuthTries|4|report"
"etc/sysctl.d/99-hardening.conf|net.ipv4.conf.all.rp_filter|1|report"
)
drift=0
for row in "${settings[@]}"; do
IFS='|' read -r file key want mode <<<"$row"
path="$root/$file"
have=$({ sed -E 's/[[:space:]]*=[[:space:]]*/ /' "$path" 2>/dev/null || true; } |
awk -v k="$key" '$1 == k { v = $2 } END { print v }')
if [[ "$have" == "$want" ]]; then
printf 'OK %-40s %s\n' "$key" "$have"
continue
fi
drift=$((drift + 1))
printf 'DRIFT %-40s have=%s want=%s mode=%s\n' "$key" "${have:-<unset>}" "$want" "$mode"
if [[ "$mode" == fix ]]; then
if [[ ! -f "$path" ]]; then
printf 'SKIP %-40s not fixed, %s does not exist\n' "$key" "$file"
else
if grep -qE "^[[:space:]]*${key}[[:space:]]" "$path"; then
sed -i -E "s|^[[:space:]]*${key}[[:space:]].*|${key} ${want}|" "$path"
else
printf '%s %s\n' "$key" "$want" >>"$path"
fi
printf 'FIXED %-40s now=%s\n' "$key" "$want"
fi
fi
done
echo "drift_count=$drift"
exit $(( drift > 0 ))
Driver and real output (run inside a Debian container with GNU sed):
rm -rf /tmp/root && mkdir -p /tmp/root/etc/ssh /tmp/root/etc/sysctl.d
printf 'PermitRootLogin yes\nMaxAuthTries 6\n' > /tmp/root/etc/ssh/sshd_config
printf 'net.ipv4.conf.all.rp_filter = 0\n' > /tmp/root/etc/sysctl.d/99-hardening.conf
bash drift-check.sh /tmp/root; echo "exit=$?"
bash drift-check.sh /tmp/root; echo "exit=$?"
DRIFT PermitRootLogin have=yes want=no mode=fix
FIXED PermitRootLogin now=no
DRIFT X11Forwarding have=<unset> want=no mode=fix
FIXED X11Forwarding now=no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=4
exit=1
OK PermitRootLogin no
OK X11Forwarding no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=2
exit=1
Reading the script, line by line:
settings=( ... )is a table with one row per setting; the four fields in each row are separated by|.IFS='|' read -r file key want mode <<<"$row"splits one row on|into four variables (IFSis the field separatorreaduses,-rstops backslashes being treated as escapes, and<<<"$row"feeds the row toreadas its input). For the first row that givesfile=etc/ssh/sshd_config key=PermitRootLogin want=no mode=fix.- The
have=$(...)line finds the current value. The{ sed ... || true; }group matters: withset -euo pipefail, a missing config file would makesedfail and silently end the whole script with status 2 and no output. With the guard, a missing file simply reads as<unset>and is reported as drift (run on a root whosesshd_configalready meets the baseline and which has no sysctl file, it printsDRIFT net.ipv4.conf.all.rp_filter have=<unset> want=1 mode=reportanddrift_count=1). The same guard does not help afixrow, because the rewrite needs a file to edit: without the[[ ! -f "$path" ]]test in the fix branch, a root with nosshd_configmadegrepand then the>>append fail and the whole run stop after the first DRIFT line, with nodrift_countand exit status 1, which looks the same as ordinary drift. With the test, that row printsSKIP PermitRootLogin not fixed, etc/ssh/sshd_config does not exist, stays counted as drift, and the run finishes (on an empty root all four rows are drift, two of themSKIP,drift_count=4, exit 1; run in the same container).sed -E 's/[[:space:]]*=[[:space:]]*/ /'turns akey = valueline intokey valueby replacing the first=and the spaces around it with one space (lines already writtenkey valuepass through unchanged).awk -v k="$key" '$1 == k { v = $2 } END { print v }'then looks for lines whose first word equals the key, remembers the second word, and prints the last one remembered when the file ends. Run on a small file in a Debian container:
--- after sed:
PermitRootLogin yes
MaxAuthTries 6
MaxAuthTries 5
# X11Forwarding yes
--- awk for MaxAuthTries:
5
The commented-out X11Forwarding line is ignored because its first word is #, and when a key appears twice the last line wins. That is how sysctl files behave; sshd keeps the first value instead, one more reason the sketch's reading of a file can differ from the effective configuration.
[[ "$have" == "$want" ]]printsOKand moves to the next row; anything else counts as drift, incrementsdriftand printsDRIFT. Only rows with modefixreach the rewrite, and only when the file exists:sed -ireplaces the existing line, orprintf ... >>appends one if the key was missing.exit $(( drift > 0 ))ends the script with status 1 if any drift remains and 0 otherwise, because the arithmetic comparison evaluates to 1 or 0. Run in the same container,drift=0exits 0 anddrift=3exits 1.
The first run fixes two settings and reports two. The second run finds the fixed ones clean (idempotent) while the report-only ones stay open until a human acts. The non-zero exit code lets a scheduler or pipeline alarm on remaining drift.
Limits of this sketch, so you do not trust it blindly: it reads files, not effective configuration. The sshd manual says that for each keyword the first obtained value is used, and that keywords after a Match line apply only to matching connections, so an appended line can be shadowed or captured by a Match block (a section of the configuration whose settings apply only to connections matching a condition, such as a particular user or source address). In production, read the effective SSH configuration with sshd -T, which prints it, rather than grepping the file.
Worked example: Windows registry-backed setting
On Windows the same decision logic applies; prefer Group Policy or endpoint management to enforce, and use a script only for what they do not cover. This reports drift for SMB signing and fixes it only when run with -Fix:
param([switch]$Fix)
$path = 'HKLM:\SYSTEM\CurrentControlSet\Services\LanmanServer\Parameters'
$name = 'RequireSecuritySignature'
$want = 1
$have = Get-ItemPropertyValue -Path $path -Name $name -ErrorAction SilentlyContinue
if ($have -eq $want) { "OK $name = $have"; return }
"DRIFT $name have=$have want=$want"
if ($Fix) {
Set-ItemProperty -Path $path -Name $name -Value $want
"FIXED $name now=$want"
}
Run it on a Windows host, since the registry provider is Windows-only. RequireSecuritySignature is the registry setting Microsoft names for the "always sign" SMB policy. It is deliberately report-only by default.
Trade-offs and pitfalls
- Auto-fix everything and you will eventually take down an application at 3 a.m.; report everything and drift grows faster than people close tickets. The per-setting rule is the compromise.
- Detecting drift is half the job: find out why it happened (manual change, a package script, a stopped agent) or the same drift returns.
- Measure it: drift count per host per week, mean time to close report-only items, and number of exceptions past expiry.
- A scanner finding is only evidence if the scan could authenticate and covered the host.
A new Windows server has just been provisioned and should have your organisation's hardening baseline applied through Group Policy and local hardening scripts. How do you verify automatically that it really applied, which commands and logs do you inspect, and how do noncompliant hosts get surfaced to the team?
Sample Answer
Direct answer
Do not trust "the Group Policy Object (GPO) is linked" or "the script exited 0". Verify the effective state on the host, automatically, in four layers, and report centrally:
- Did policy process? Look at the Group Policy Operational event log and run
gpresultto see which GPOs applied or were denied. - Is the setting actually in effect? Query the live value with the tool that owns it (firewall profile, registry,
auditpol, SMB configuration) and compare with the expected value. - Did the local hardening script finish? Check a run record that the script writes at the end (version, exit code, time), not just the scheduler's "task succeeded".
- Surface the result. Each host emits a JSON result with a status per setting; a collector aggregates them and the team gets a list of non-compliant hosts, with "not yet applied" and "deliberately excepted" shown separately from "wrong".
Layer 1: did Group Policy process, and what applied
Run these on the host from an elevated prompt (or remotely through your management tooling):
gpupdate /force /target:computer
gpresult /scope computer /h C:\Temp\gp-computer.html /f
gpresultreports the Resultant Set of Policy (RSoP), the combined outcome of every GPO that applies to the computer. It needs an output option./hwrites an HTML report and/foverwrites an existing file;/rprints a summary to the console and/vor/zprint more detail. The report lists the GPOs that applied and those that were denied, with the reason. Two common reasons: a security filter (a setting on the GPO that limits it to named computers or groups, so a server outside them is denied) and a Windows Management Instrumentation (WMI) filter (a query about the machine, such as its OS version, that must return true for the GPO to apply). How to read it for a new server: find the hardening GPO by name. If it sits under the applied list, policy reached the host. If it sits under the denied list, the stated reason tells you whether the server is in the wrong group or OU, or failed a filter. If it appears in neither, the GPO is not linked above the server's OU. Archive the report as evidence for a failed host.- The event log
Microsoft-Windows-GroupPolicy/Operationalrecords the processing. A client-side extension (CSE) is the part of Windows that applies one family of Group Policy settings (security settings, firewall, advanced audit and so on); each runs separately, so one can fail while the others succeed. Event 4016 marks a client-side extension starting and 5016 marks it completing successfully, so a healthy processing run shows a 4016 and a 5016 pair per extension, and a 4016 with no later 5016 means that extension did not finish.Get-WinEventprints events with the columnsTimeCreated,Id,LevelDisplayNameandMessage; the extension's name is in theMessagetext. Warnings and errors in the System log (for example 1129, no network path to a domain controller, or 1058, the policy files could not be read from the domain controller) explain why policy did not process. - Clock skew matters: Microsoft's guidance notes that a difference of more than five minutes between a computer and its domain controller can stop the computer authenticating, so a new server with a wrong clock gets no policy.
w32tm /resyncfixes it. - A success event does not prove the setting is in effect. For the advanced audit extension, event 5016 can carry the return status
E_PENDING(a Windows status meaning "not finished yet") by design: an asynchronous thread was started to apply the audit settings, so the event says only that processing began. Check the Security-Audit-Configuration-Client Operational log (under Applications and Services Logs, Microsoft, Windows) and, more importantly, read the live audit policy withauditpolas in Layer 2.
Get-WinEvent -FilterHashtable @{ LogName = 'Microsoft-Windows-GroupPolicy/Operational'; Id = 5016 } -MaxEvents 5
Layer 2 and 3: check the live state and the local script
The script below runs on the host (from a scheduled task, a remoting call or your endpoint management tool). It checks a baseline of eight items covering six kinds of evidence, and writes a JSON result:
#requires -Version 5.1
[CmdletBinding()]
param(
[datetime]$ProvisionedAtUtc = [datetime]::MinValue,
[int]$GraceMinutes = 120,
[int]$MaxGpoAgeHours = 24,
[string]$RunRecordPath = 'C:\ProgramData\Hardening\last-run.json',
[string]$OutputPath = (Join-Path $env:TEMP 'baseline-result.json')
)
$Baseline = @{
Version = '2026.10'
Checks = @(
@{ Id = 'fw-domain-on'; Type = 'Firewall'; Profile = 'Domain'; Expected = 'True' }
@{ Id = 'fw-public-on'; Type = 'Firewall'; Profile = 'Public'; Expected = 'True' }
@{ Id = 'tls10-server'; Type = 'Registry'
Path = 'HKLM:\SYSTEM\CurrentControlSet\Control\SecurityProviders\SCHANNEL\Protocols\TLS 1.0\Server'
Name = 'Enabled'; Expected = '0' }
@{ Id = 'smb1-off'; Type = 'Smb'; Property = 'EnableSMB1Protocol'; Expected = 'False' }
@{ Id = 'smb-sign-req'; Type = 'Smb'; Property = 'RequireSecuritySignature'; Expected = 'True'
Exception = @{ Ticket = 'CHG-1042'; Approver = 'secops'; Expires = '2026-12-31' } }
@{ Id = 'audit-proc-create'; Type = 'Audit'; Category = 'Detailed Tracking'
Subcategory = 'Process Creation'; Expected = 'Success' }
@{ Id = 'local-script-ran'; Type = 'RunRecord'; Expected = '2026.10' }
@{ Id = 'gpo-fresh'; Type = 'GpoFresh'; Expected = 'True' }
)
}
function Get-ObservedValue {
param($Check)
switch ($Check.Type) {
'Firewall' { [string](Get-NetFirewallProfile -Name $Check.Profile -PolicyStore ActiveStore).Enabled }
'Registry' {
if (-not (Test-Path -LiteralPath $Check.Path)) { '<missing>' }
else {
$props = Get-ItemProperty -LiteralPath $Check.Path -ErrorAction Stop
if ($props.PSObject.Properties.Name -contains $Check.Name) { [string]$props.($Check.Name) } else { '<missing>' }
}
}
'Smb' { [string](Get-SmbServerConfiguration).($Check.Property) }
'Audit' {
$rows = auditpol /get /category:"$($Check.Category)" /r |
ConvertFrom-Csv -Header Machine, Target, Subcategory, Guid, Inclusion, Exclusion
$row = $rows | Where-Object Subcategory -match $Check.Subcategory | Select-Object -First 1
if (-not $row) { throw "subcategory $($Check.Subcategory) not in auditpol output" }
[string]$row.Inclusion
}
'RunRecord' {
if (-not (Test-Path -LiteralPath $RunRecordPath)) { '<missing>' }
else { [string](Get-Content -LiteralPath $RunRecordPath -Raw -ErrorAction Stop | ConvertFrom-Json).BaselineVersion }
}
'GpoFresh' {
$since = (Get-Date).AddHours(-$MaxGpoAgeHours)
$ev = Get-WinEvent -FilterHashtable @{ LogName = 'Microsoft-Windows-GroupPolicy/Operational'; Id = 5016; StartTime = $since } -MaxEvents 1 -ErrorAction SilentlyContinue
[string][bool]$ev
}
}
}
function Resolve-CheckStatus {
param($Check, [string]$Observed, [string]$ErrorText, [datetime]$Now, [datetime]$ProvisionedAtUtc, [int]$GraceMinutes)
if ($ErrorText) { return @{ Status = 'Unknown'; Note = $ErrorText } }
$ok = if ($Check.Type -eq 'Audit') { $Observed -match $Check.Expected } else { $Observed -eq $Check.Expected }
if ($ok) { return @{ Status = 'Compliant'; Note = '' } }
$ex = $Check.Exception
if ($ex -and ([datetime]$ex.Expires).Date -ge $Now.Date) {
return @{ Status = 'Excepted'; Note = "$($ex.Ticket) approved by $($ex.Approver) until $($ex.Expires)" }
}
if (($Now - $ProvisionedAtUtc).TotalMinutes -lt $GraceMinutes) {
return @{ Status = 'Pending'; Note = 'inside the post-provisioning grace window' }
}
$note = if ($ex) { "exception $($ex.Ticket) expired $($ex.Expires)" } else { '' }
return @{ Status = 'NonCompliant'; Note = $note }
}
function Invoke-Baseline {
$now = [datetime]::UtcNow
$results = foreach ($c in $Baseline.Checks) {
$obs = $null; $err = $null
try { $obs = Get-ObservedValue -Check $c } catch { $err = $_.Exception.Message }
$r = Resolve-CheckStatus -Check $c -Observed $obs -ErrorText $err -Now $now -ProvisionedAtUtc $ProvisionedAtUtc -GraceMinutes $GraceMinutes
[pscustomobject]@{ id = $c.Id; expected = $c.Expected; observed = $obs; status = $r.Status; note = $r.Note }
}
$doc = [pscustomobject]@{
host = $env:COMPUTERNAME; baselineVersion = $Baseline.Version
checkedAtUtc = $now.ToString('o'); results = @($results)
}
$doc | ConvertTo-Json -Depth 5 | Set-Content -LiteralPath $OutputPath -Encoding UTF8
if ($results | Where-Object status -eq 'NonCompliant') { exit 1 }
if ($results | Where-Object status -in 'Unknown', 'Pending') { exit 2 }
exit 0
}
Invoke-Baseline
What each kind of check does:
| Check | Source of truth | Expected value (example) |
|---|---|---|
Firewall | Get-NetFirewallProfile -PolicyStore ActiveStore, the effective policy summed over all GPOs and local rules | profile enabled |
Registry | Get-ItemProperty on the Schannel protocol key (Enabled of the Transport Layer Security 1.0 server role) | 0 |
Smb (Server Message Block, the Windows file-sharing protocol) | Get-SmbServerConfiguration properties EnableSMB1Protocol and RequireSecuritySignature | False and True |
Audit | auditpol /get /category:"Detailed Tracking" /r parsed as CSV; the Process Creation row must include Success | Success |
RunRecord | last-run.json written by the hardening script at the end of a successful run, holding BaselineVersion | 2026.10 |
GpoFresh | an event 5016 in Microsoft-Windows-GroupPolicy/Operational within the last 24 hours (an example threshold) | True |
The local hardening script must write the run record only after all its steps succeed, for example { "BaselineVersion": "2026.10", "ExitCode": 0, "FinishedUtc": "..." }. A missing or older record means the script did not finish or an old baseline version ran. A record, registry key or registry value that does not exist yet is not a failed check: the script reports the observed value <missing>, which is a mismatch like any other, so it goes through the exception and grace-window rules below. That is exactly what a server on which Group Policy has not yet processed looks like. Pass -ProvisionedAtUtc (from the asset inventory) so the grace window can work; its default of the minimum date value switches the grace window off.
What the collected evidence looks like
The auditpol report. With /r, auditpol prints CSV with six columns: machine name, policy target, subcategory, subcategory GUID, inclusion setting and exclusion setting (the Microsoft Learn description of the report format). The script supplies those six names itself with ConvertFrom-Csv -Header, so the printed header line becomes the first parsed row, which is harmless because the filter Where-Object Subcategory -match 'Process Creation' skips it. Illustrative rows for a host called WEB01, parsed on PowerShell 7:
Machine Subcategory Inclusion
------- ----------- ---------
Machine Name Subcategory Inclusion Setting
WEB01 Process Creation Success
WEB01 Process Termination No Auditing
The check reads the Inclusion value of the Process Creation row, here Success, which matches the expected Success. A row reading No Auditing would be NonCompliant.
An exception as a record. In the baseline, an exception is a small hash table attached to the check, with a ticket, an approver and an expiry date. This is the smb-sign-req check from the script:
@{ Id = 'smb-sign-req'; Type = 'Smb'; Property = 'RequireSecuritySignature'; Expected = 'True'
Exception = @{ Ticket = 'CHG-1042'; Approver = 'secops'; Expires = '2026-12-31' } }
The result file the collector receives. baseline-result.json has one entry per check. This illustrative example for WEB01 uses four of the eight checks and was produced by running the script's own Resolve-CheckStatus function on PowerShell 7 with a fixed clock; the observed values are made up to show each status:
{
"host": "WEB01",
"baselineVersion": "2026.10",
"checkedAtUtc": "2026-10-05T12:00:00.0000000Z",
"results": [
{ "id": "fw-domain-on", "expected": "True", "observed": "True", "status": "Compliant", "note": "" },
{ "id": "smb1-off", "expected": "False", "observed": "True", "status": "NonCompliant", "note": "" },
{ "id": "smb-sign-req", "expected": "True", "observed": "False", "status": "Excepted", "note": "CHG-1042 approved by secops until 2026-12-31" },
{ "id": "audit-proc-create", "expected": "Success", "observed": "Success", "status": "Compliant", "note": "" }
]
}
(The script writes it with ConvertTo-Json, which prints each field on its own line; it is compacted here.) The collector only needs the host and the status of each entry: this host would be listed as NonCompliant because of smb1-off, with the excepted SMB signing counted separately. The script's exit code for this host would be 1, since one item is non-compliant.
Distinguishing "not yet applied" from "deliberately excepted"
The function Resolve-CheckStatus assigns one of five statuses to each check. The rules, in order:
- Observed value equals expected:
Compliant. - Not equal, and the check carries an exception (ticket, approver, expiry date) that has not expired:
Excepted, with the ticket in the note. - Not equal, no valid exception, and the host was provisioned less than the grace period ago (120 minutes here, an example):
Pending. Group Policy and first-boot scripts need time, and flagging every new server red in its first minutes trains the team to ignore the report. - Not equal otherwise:
NonCompliant. If an exception existed but expired, the note says so. - The check itself failed (for example access denied, or a malformed run record):
Unknownwith the error text, never a pass. A key, value or record that simply does not exist is not a failure of the check; it is a mismatch with observed value<missing>and follows rules 2 to 4.
The script exits 1 for any non-compliant item and 2 when only Unknown or Pending items remain, so a pipeline can gate a server's promotion into production on exit code 0.
The decision function was executed on PowerShell 7 with fixed times. Save the script as Test-HardeningBaseline.ps1 and the harness below as t.ps1 in the same folder, then run pwsh -NoProfile -File t.ps1. The harness loads the function from the saved script:
$e=$null;$t=$null
$ast=[System.Management.Automation.Language.Parser]::ParseFile("$PWD/Test-HardeningBaseline.ps1",[ref]$t,[ref]$e)
"parse errors: $($e.Count)"
foreach($f in $ast.FindAll({param($n) $n -is [System.Management.Automation.Language.FunctionDefinitionAst] -and $n.Name -eq 'Resolve-CheckStatus'},$true)){ Invoke-Expression $f.Extent.Text }
$now=[datetime]'2026-10-05T12:00:00Z'
$prov=[datetime]'2026-10-05T10:30:00Z' # 90 min before
$old=[datetime]'2026-10-04T10:30:00Z'
$ex=@{Ticket='CHG-1042';Approver='secops';Expires='2026-12-31'}
$exOld=@{Ticket='CHG-0001';Approver='secops';Expires='2026-09-30'}
function show($label,$r){ '{0,-34} {1,-13} {2}' -f $label,$r.Status,$r.Note }
show 'match' (Resolve-CheckStatus -Check @{Type='Smb';Expected='True'} -Observed 'True' -ErrorText $null -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
show 'audit contains Success' (Resolve-CheckStatus -Check @{Type='Audit';Expected='Success'} -Observed 'Success and Failure' -ErrorText $null -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
show 'mismatch, valid exception' (Resolve-CheckStatus -Check @{Type='Smb';Expected='True';Exception=$ex} -Observed 'False' -ErrorText $null -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
show 'mismatch, 90 min old host' (Resolve-CheckStatus -Check @{Type='Smb';Expected='True'} -Observed 'False' -ErrorText $null -Now $now -ProvisionedAtUtc $prov -GraceMinutes 120)
show 'mismatch, expired exception' (Resolve-CheckStatus -Check @{Type='Smb';Expected='True';Exception=$exOld} -Observed 'False' -ErrorText $null -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
show 'mismatch, no exception' (Resolve-CheckStatus -Check @{Type='Smb';Expected='True'} -Observed 'False' -ErrorText $null -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
show 'check threw' (Resolve-CheckStatus -Check @{Type='Smb'} -Observed $null -ErrorText 'access denied' -Now $now -ProvisionedAtUtc $old -GraceMinutes 120)
Output:
parse errors: 0
match Compliant
audit contains Success Compliant
mismatch, valid exception Excepted CHG-1042 approved by secops until 2026-12-31
mismatch, 90 min old host Pending inside the post-provisioning grace window
mismatch, expired exception NonCompliant exception CHG-0001 expired 2026-09-30
mismatch, no exception NonCompliant
check threw Unknown access denied
The missing-record branch of the collection code was executed on PowerShell 7 with a run-record path that does not exist (save the harness as t3.ps1 next to the script and run pwsh -NoProfile -File t3.ps1):
$e=$null;$t=$null
$ast=[System.Management.Automation.Language.Parser]::ParseFile("$PWD/Test-HardeningBaseline.ps1",[ref]$t,[ref]$e)
"parse errors: $($e.Count)"
$RunRecordPath = Join-Path $PWD 'no-such-last-run.json'
foreach($f in $ast.FindAll({param($n) $n -is [System.Management.Automation.Language.FunctionDefinitionAst] -and $n.Name -in 'Get-ObservedValue','Resolve-CheckStatus'},$true)){ Invoke-Expression $f.Extent.Text }
$now = [datetime]'2026-10-05T12:00:00Z'
$check = @{ Id = 'local-script-ran'; Type = 'RunRecord'; Expected = '2026.10' }
$obs = Get-ObservedValue -Check $check
"observed with no run record: $obs"
foreach ($age in 10, 600) {
$r = Resolve-CheckStatus -Check $check -Observed $obs -ErrorText $null -Now $now -ProvisionedAtUtc $now.AddMinutes(-$age) -GraceMinutes 120
'{0,3}-minute-old host: {1}' -f $age, $r.Status
}
parse errors: 0
observed with no run record: <missing>
10-minute-old host: Pending
600-minute-old host: NonCompliant
The rest of the collection of live values (Get-NetFirewallProfile, auditpol, registry and SMB queries, including the registry branch, which uses the same Test-Path pattern) needs a Windows host, so that part was parse-checked and its cmdlets verified against Microsoft Learn, not executed.
Surfacing noncompliant hosts
- Each host writes
baseline-result.json; an agent or scheduled task ships it to a central store (a file share, a SIEM (security information and event management) index or a database), keyed by host name. - A daily report lists hosts by status:
NonCompliantfirst with the failing check ids, thenUnknown(the host could not be evaluated, which includes a host that stopped reporting), thenPendingolder than a few hours, and a count of active exceptions with their expiry dates. - Alerts go to the owning team's queue with the host, the failing check, the expected and observed values, and a link to the stored
gpresultreport. - Reconcile the list against the asset inventory: a server with no result at all is an alert in its own right.
- Trend the numbers: percentage compliant per baseline version, mean time from provisioning to compliant, and exceptions nearing expiry.
Pitfalls
- Reading GPO status from the Group Policy Management Console instead of the host: it shows what is linked, not what applied.
- Treating event 5016 as proof for audit settings (see the
E_PENDINGnote above). - A grace window that never closes: always compute
Pendingfrom the provisioning time and let it expire. - Exceptions without an expiry date become permanent.
- Checking only the registry value that a GPO sets: another source may override it later, so also check the effective state where a tool offers one (the firewall ActiveStore,
auditpol).
What does least privilege mean on a host, and how would you actually enforce it for users, files and directories on both Windows and Linux?
Sample Answer
Direct answer
Least privilege means every user, program and service gets only the permissions needed for its job, and only for as long as it needs them. On a host that means three layers: people (no everyday admin rights, separate admin identities, named sudo rules on Linux, managed local admin passwords on Windows), files and directories (owner and group set to the real service account, no access for "everyone", broken inheritance only where needed), and services (each running under its own low-privilege account). I enforce it with groups rather than individual grants, write the permissions as code, and then audit them with commands so I find drift rather than assume it.
Users: who can do what
| Layer | Linux | Windows |
|---|---|---|
| Daily account | Normal user, no root | Normal user, not a member of local Administrators |
| Elevation | sudo rules naming specific commands, per group | UAC (User Account Control) prompts for admin tasks, plus a separate admin account for administrators |
| Shared local admin credential | Do not have one; root login over SSH is disabled | Windows LAPS (Local Administrator Password Solution) gives every machine its own rotating local admin password, stored in Active Directory or Microsoft Entra ID (Microsoft's cloud identity directory, formerly called Azure AD) and readable only by authorised staff |
| Review | sudo -l -U <user> | Get-LocalGroupMember -Group 'Administrators' |
Windows LAPS is built into supported Windows versions, rotates the managed local administrator password regularly, and exists to stop one shared password enabling lateral movement from machine to machine (Microsoft Learn, Windows LAPS overview). Retrieval is a deliberate act: Get-LapsADPassword -Identity 'WEB01' returns the password as a SecureString unless you add -AsPlainText, which Microsoft says to use only in support or testing situations.
Linux sudo, tested. A sudoers file with a syntax error can lock out all admins, so it is validated before install. Executed in an ubuntu:24.04 container:
$ cat good.sudoers
deploy ALL=(root) NOPASSWD: /usr/bin/systemctl restart webapp
$ visudo -cf good.sudoers
good.sudoers: parsed OK
$ cat bad.sudoers
deploy ALL=(root) NOPASSWD /usr/bin/systemctl
$ visudo -cf bad.sudoers
bad.sudoers:1:46: syntax error
deploy ALL=(root) NOPASSWD /usr/bin/systemctl
^
$ install -m 0440 -o root -g root good.sudoers /etc/sudoers.d/deploy-webapp
$ sudo -l -U deploy
Matching Defaults entries for deploy on 75de9dd7b378:
env_reset, mail_badpass, secure_path=/usr/local/sbin\:/usr/local/bin\:/usr/sbin\:/usr/bin\:/sbin\:/bin\:/snap/bin, use_pty
User deploy may run the following commands on 75de9dd7b378:
(root) NOPASSWD: /usr/bin/systemctl restart webapp
deploy can restart one service as root and nothing else. The NOPASSWD tag removes the password prompt for that one command, which is acceptable for automation but should not be a blanket on ALL. Be careful with wildcards in command paths and with commands that can launch a shell (editors, pagers): the sudoers manual warns that wildcards can grant more than intended, and the NOEXEC tag stops dynamically linked programs from spawning further commands.
If you do three things first, make them these: remove everyday admin rights, give every machine its own LAPS password, and run each service under its own account. The finer points are details for when you write the rules: NOEXEC is a sudoers tag that stops a permitted program from launching further programs, UAC is the Windows prompt that asks for confirmation before an admin action, and a SecureString is a .NET string type that is kept encrypted in memory and does not print as plain text.
Files and directories on Linux
Set the owner and group to the account that really needs access, remove the "other" bits, and add ACLs (access control lists, per-user or per-group permissions beyond the owner, group and other triplet) rather than widening the group. Executed in the same container:
$ chmod 750 /srv/app; chmod 2770 /srv/app/shared
$ setfacl -m u:deploy:rx /srv/app
$ getfacl -p /srv/app
# file: /srv/app
# owner: webapp
# group: webapp
user::rwx
user:deploy:r-x
group::r-x
mask::r-x
other::---
$ stat -c '%A %U:%G %n' /srv/app /srv/app/shared
drwxr-x--- webapp:webapp /srv/app
drwxrws--- webapp:webapp /srv/app/shared
$ touch /srv/app/oops; chmod 666 /srv/app/oops
$ find /srv -xdev -type f -perm -0002 -print
/srv/app/oops
Reading the digits: each digit is read 4 + write 2 + execute 1, for owner, group and other in that order. 750 is 7 (4+2+1, rwx) for the owner, 5 (4+1, r-x) for the group and 0 for everyone else, which is the drwxr-x--- in the stat line. 2770 has a fourth, leading digit for special bits, where 2 means setgid: setgid, then rwx for owner, rwx for group, nothing for others, shown as drwxrws--- with the s standing in for the group x.
mask::r-x is the ceiling for every named-user entry and for the group entry: an entry's effective permission is its own bits ANDed with the mask. Here deploy has r-x and the mask allows r-x, so deploy gets r-x. Had the mask been r--, deploy would have lost execute (the right to enter the directory) and getfacl would have printed #effective:r-- beside the entry. setfacl -m recalculates the mask for you unless told not to.
other::--- means a user who is neither webapp nor deploy cannot even list /srv/app. deploy can read and traverse but not write. The setgid bit (the s in drwxrws---) makes new files in shared inherit its group. The find command is the audit for the classic mistake, a world-writable file; run it on a schedule. The same applies to setuid programs, which run with their owner's rights: list them with find / -xdev -type f -perm -4000 and review the list against what the server's role needs.
Files and directories on Windows
NTFS (the Windows file system) permissions are held in an ACL on each folder, and by default folders inherit entries from their parent. Least privilege usually means breaking inheritance on a sensitive folder and granting the service account, administrators and SYSTEM only what each needs. This script was parse-checked with PowerShell 7 and every cmdlet, method and parameter was checked against Microsoft Learn; it was not executed, because it needs a Windows host (Set-Acl is Windows-only, and Get-LapsADPassword needs the LAPS module and an AD):
# 1. Who is in the local Administrators group on this host?
Get-LocalGroupMember -Group 'Administrators' | Select-Object Name, ObjectClass, PrincipalSource
# 2. Restrict D:\AppData to the service account, Administrators and SYSTEM
$path = 'D:\AppData'
$acl = Get-Acl -Path $path
$acl.SetAccessRuleProtection($true, $false) # stop inheriting, drop the inherited entries
$inherit = [System.Security.AccessControl.InheritanceFlags]'ContainerInherit,ObjectInherit'
$propagate = [System.Security.AccessControl.PropagationFlags]::None
foreach ($rule in @(
@('CORP\svc-webapp', 'Modify'),
@('BUILTIN\Administrators', 'FullControl'),
@('NT AUTHORITY\SYSTEM', 'FullControl'))) {
$ace = New-Object System.Security.AccessControl.FileSystemAccessRule(
$rule[0], $rule[1], $inherit, $propagate, 'Allow')
$acl.AddAccessRule($ace)
}
Set-Acl -Path $path -AclObject $acl -WhatIf # remove -WhatIf after reviewing
icacls $path
# 3. Break-glass: read one machine's local admin password from AD
Get-LapsADPassword -Identity 'WEB01' -AsPlainText
ContainerInherit,ObjectInherit makes each entry apply to the subfolders (containers) and the files (objects) inside D:\AppData, so the grant flows down the tree; PropagationFlags::None adds no extra limit on that flow. In icacls notation the same ideas are (CI) and (OI), and RX means read and execute.
SetAccessRuleProtection($true, $false) means "protect from inheritance, and do not preserve the inherited entries", so the folder ends with only the three explicit entries. The three grants are added to the same in-memory ACL before Set-Acl writes it, so administrators are never locked out part-way. The -WhatIf switch makes Set-Acl report what it would do without doing it, so while -WhatIf is present nothing is written and icacls $path still prints the old ACL. After you remove -WhatIf and run the script for real, icacls $path lists the resulting ACL for review, one entry per line in the form name:(inheritance flags)(rights), where (M) is Modify and (F) is Full control. You should see exactly the three entries above (service account (M), Administrators (F), SYSTEM (F)); an entry marked (I) or any other name means inheritance was not removed or an extra grant exists. This describes the documented format and was not run, since it needs a Windows host. icacls $path /save <file> stores a copy of the ACL that /restore can re-apply if the change goes wrong.
Share permissions and NTFS permissions are separate layers: when a folder is reached over a network share, the access a user gets is the more restrictive of the two, so a locked-down NTFS ACL is not undone by a generous share.
Services
Each service runs under its own account with only the rights it needs, and never as root or as a domain administrator. On Linux that is a dedicated user such as webapp that owns only /srv/app. On Windows it is a dedicated service account, like CORP\svc-webapp above, granted rights only on its own folders and not given interactive logon.
Worked example
A web server hosts an app owned by webapp, deployed by a pipeline user deploy, administered by two people. Result: webapp owns /srv/app with mode 750 (full access for the owner, read and traverse for the group, nothing for others); deploy gets read via an ACL and one sudo rule to restart the service; the two administrators have personal logins with sudo for a short named list; there is no shared password anywhere; and the find audit above returns nothing. On the Windows file server the same pattern gives one service account modify rights on D:\AppData, administrators and SYSTEM full control, and everyone else no entry at all.
Trade-offs and pitfalls
- Least privilege costs setup time and generates "access denied" tickets. Pay that cost once in staging, by running the application as its real service account there.
- Group-based grants decay: people change roles. Review membership of the privileged groups quarterly, with
Get-LocalGroupMemberandsudo -l -U. - Standing admin rights are the thing to remove first. Time-limited elevation (just-in-time access) is the stronger end state but needs tooling; do the cheap steps first.
- A broad sudo rule on
ALL, or a sudo rule that lets someone run an editor as root, is effectively root; read rules for what they can be made to do.
Your organisation runs about 1,000 mixed Windows and Linux servers, on-prem and in cloud, with no consistent security baseline. Design the programme that gets them to one and keeps them there: what you enforce first, how you enforce it, how you prove it, and how you decide what to leave out.
Sample Answer
Direct answer
I would run this as a measured programme, not a one-off hardening project. Pick one published baseline per operating system (CIS Benchmarks, the Center for Internet Security's consensus configuration guides, at Level 1 as the starting profile), measure first without changing anything, enforce a short first tier of low-breakage, high-impact settings through the same tooling that builds and manages the servers, prove compliance with automated scans whose results feed a dashboard, and handle every deliberate omission through an exception record with an owner, a compensating control and an expiry date. The baseline is a living definition in version control, so "keeping them there" is a scheduled enforcement and scan loop, not a recurring audit.
For illustration assume 600 Linux and 400 Windows servers (1,000 in total) in a mix of on-premises and cloud. The split is an assumption; the structure does not depend on it.
What the baseline is, and why Level 1 first
CIS defines a Level 1 profile as a base recommendation that can be implemented fairly promptly without extensive performance impact, and Level 2 as defense in depth for environments where security is paramount, which can have an adverse effect if implemented without due care. CIS also advises testing either level in a test environment first. For Windows Server, Microsoft publishes security baselines as Group Policy Object (GPO) backups in its Security Compliance Toolkit, together with Policy Analyzer (compares sets of GPOs against each other or against local policy) and LGPO.exe (applies local policy to machines that are not domain-joined). So the plan is: CIS Level 1 server profile per Linux distribution, and either the Microsoft baseline or the CIS Windows Server benchmark per Windows version, choosing one per OS and not mixing, because overlapping baselines conflict.
Phase 0: discovery and assessment (measure before enforcing)
You cannot enforce what you have not inventoried. Build the server list from more than one source (the configuration management database, the hypervisor and cloud account inventories, the directory) and reconcile it, because the servers nobody remembers are the ones with no baseline. Record for each server: OS and version, owner, environment, internet exposure, whether it can be rebooted, and whether a vendor supports changes to it.
Then run the benchmark in assessment-only mode everywhere. For Linux, OpenSCAP (oscap, an open-source implementation of SCAP, the Security Content Automation Protocol) with the SCAP Security Guide content (the rule definitions) evaluates a profile and writes results without changing the host. A real run in a test AlmaLinux 9 container against the CIS server Level 1 profile:
dnf -y install openscap-scanner scap-security-guide
oscap xccdf eval \
--profile xccdf_org.ssgproject.content_profile_cis_server_l1 \
--results /tmp/results.xml --report /tmp/report.html \
/usr/share/xml/scap/ssg/content/ssg-almalinux9-ds.xml
echo "exit status: $?"
The run printed one block per rule: a Title, the Rule id and a Result. Three of them, copied from the run:
Title Ensure AlmaLinux GPG Key Installed
Rule xccdf_org.ssgproject.content_rule_ensure_almalinux_gpgkey_installed
Result pass
Title Implement Custom Crypto Policy Modules for CIS Benchmark
Rule xccdf_org.ssgproject.content_rule_configure_custom_crypto_policy_cis
Result fail
Title Install AIDE
Rule xccdf_org.ssgproject.content_rule_package_aide_installed
Result notapplicable
pass means the host matches the rule, fail means it does not and needs a fix or an exception, and notapplicable means the rule does not apply to this host (the scanner decides this from conditions in the content). The --profile value is the XCCDF (the XML format that SCAP uses for benchmarks) profile id, and here it selects the CIS server Level 1 rule set out of the content file. The tally was 70 pass, 5 fail, 220 notapplicable, 0 error, with exit status 2: according to the oscap manual, 0 means every rule passed, 1 means an error occurred during evaluation, and 2 means at least one rule failed or was unknown, so a pipeline can gate on the exit status. The honest compliance figure excludes the not-applicable rules from the denominator: 70 / (70 + 5) = 93.3% of the 75 applicable rules, not 70 of 295. Report both numbers, and say which one is the headline, or the dashboard will quietly flatter itself. (CIS-CAT Pro Assessor is CIS's own tool for assessing conformance, and is the alternative where a vendor-certified assessment is wanted.)
The assessment also tells you how bad it is by item: the fleet-wide failure count per rule is your prioritisation list.
What to enforce first (tier 1)
Rank candidate settings by three questions: does it close a path attackers actually use, how likely is it to break something, and can I verify it automatically? The first tier takes settings that score well on all three:
| Area | Linux | Windows | Why first |
|---|---|---|---|
| Remote administration | key-only SSH, no direct root login (sshd_config drop-in) | RDP (Remote Desktop Protocol, Windows' remote login) only from a jump host or gateway (a hardened server that administrators connect through) | most intrusions arrive over admin protocols |
| Local admin credentials | no shared local passwords | Windows LAPS (Windows Local Administrator Password Solution) rotates a unique local administrator password per machine and escrows it (stores a copy for authorised administrators to retrieve) in Active Directory or Microsoft Entra ID | one stolen local password otherwise opens every server |
| Patch posture | automatic security updates or a scheduled cadence with a reporting hook | WSUS (Windows Server Update Services, Microsoft's on-premises update server; Microsoft lists it as no longer actively developed in Windows Server 2025, with existing capabilities and content still available) or whichever update tooling the estate already runs, with a ring schedule | known-vulnerability exposure shrinks fastest here |
| Host firewall | default-deny inbound, named exceptions | host firewall enabled with inbound blocked by default | limits lateral movement (an attacker who has compromised one machine hopping to others) |
| Legacy and unused services | remove or disable what the server role does not need | disable obsolete protocols | smaller attack surface, rarely breaks anything |
| Logging and time | auditd (the Linux audit daemon) or journal forwarding, time sync | advanced audit policy, event forwarding | you need evidence for the "prove it" step and for incidents |
The Windows LAPS row of tier 1 has a concrete shape. Windows LAPS is configured through a Group Policy setting under Computer Configuration, Policies, Administrative Templates, System, LAPS. Illustrative policy for the fleet, using the setting names and defaults from Microsoft's Windows LAPS documentation: backup directory Active Directory, password age 30 days, password length 14, complexity 4 (upper case, lower case, numbers and special characters). An administrator with the right to decrypt can then read a server's current password:
Get-LapsADPassword -Identity SRV-APP-01 -AsPlainText
Output shaped like Microsoft's documentation example (host name illustrative, password omitted):
ComputerName : SRV-APP-01
Account : Administrator
Password : <the current 14-character password>
PasswordUpdateTime : 4/9/2023 9:39:38 AM
ExpirationTimestamp : 4/14/2023 9:39:38 AM
Source : EncryptedPassword
DecryptionStatus : Success
Read it as: the managed account is Administrator, the password was last rotated at PasswordUpdateTime and will be rotated again after ExpirationTimestamp, it is stored encrypted in Active Directory (Source), and the caller was allowed to decrypt it. A different password per server is the point: a stolen password for one machine opens nothing else.
Everything else in Level 1 follows in later tiers, and Level 2 items are opt-in per workload.
How to enforce: one definition, two delivery paths
The plan rests on three pieces: the published baseline as the definition, one enforcement tool per OS family (Group Policy for domain-joined Windows, Ansible for Linux), and one scanner per OS family (OpenSCAP for Linux, Policy Analyzer and CIS-CAT for Windows). Packer, LGPO.exe, Puppet, Chef and PowerShell DSC appear in the tables as the alternatives that fit particular situations (image builds, machines outside the domain, estates that already run an agent).
| Decision | Options | Recommendation |
|---|---|---|
| Image baking vs post-provision | Bake the baseline into the machine image (built with a tool such as Packer, which automates image creation, and scanned at build time) vs apply it after the server boots | Both, from the same definition. Bake for new cloud and template-built servers so they are born compliant; apply post-provision for the existing 1,000, since a new image does nothing for a server that already exists. A baked image goes stale until rebuilt, so the scheduled run still applies. |
| Agent vs agentless | Push over SSH or WinRM (Windows Remote Management) from a controller (Ansible) vs a resident agent that pulls policy (the Group Policy client, or configuration managers such as Puppet, Chef and PowerShell DSC, Desired State Configuration) | Use what already exists first. GPO for domain-joined Windows (the client is built in) and LGPO.exe for non-domain machines; Ansible, run on a schedule from a pipeline, for Linux. Agentless needs a network path and a credential vault and misses hosts that are offline during the run; an agent self-heals on its own interval and works behind NAT (network address translation, where a host sits behind a shared address and cannot be reached from outside, so it must call out to the controller instead), but adds a new always-running component to 1,000 servers. |
| Audit vs enforce | Report only vs change the setting | Audit mode in Phase 0 and for every new setting for one ring; enforce after the owner has seen the report. |
Roll out in rings, so a bad setting hurts few servers: ring 0 is 20 servers (2%), ring 1 is 100 (10%), ring 2 is 300 (30%), ring 3 is the remaining 580 (58%): 20 + 100 + 300 + 580 = 1,000. Each ring waits for a clean soak (no baseline-caused incidents) before the next starts. Pick ring 0 deliberately: non-production, then low-risk production, with owners who agreed in advance.
How to prove it, and keep proving it
- Scans on a schedule, results to one store: OpenSCAP or CIS-CAT results for Linux, Policy Analyzer comparisons and CIS-CAT for Windows. Keep the raw results file as evidence, because "dashboard says green" is not evidence.
- Metrics that cannot be gamed: percentage of servers passing tier 1 (applicable rules only), count of open exceptions and how many are expired, and median days from "new failing rule detected" to "fixed or excepted".
- Continuous assessment trade-offs: a daily scan catches drift quickly but costs CPU on the servers and generates volume; weekly scans are cheap and slower to notice. Stagger scan start times so 1,000 servers are not all scanning at once, run daily for tier 1 only and weekly for the full profile.
- Remediation automation: OpenSCAP can emit remediation content from a result (
oscap xccdf generate fix, with fix types including bash and ansible) and can apply it with--remediate. Treat generated fixes as a starting point to review and put in version control. Do not run--remediateacross production unreviewed. - Independent check: have someone outside the team sample 25 servers a quarter and compare the scan result with what is actually configured.
How to decide what to leave out
Leave a setting out (or defer it) only when one of these is true, and write it down:
| Reason | Example | Record |
|---|---|---|
| The server role needs the opposite | a legacy application needs a protocol the baseline disables | the specific hosts, the dependency, the plan to retire it |
| The control does not apply | a rule about a wireless interface on a server with none | marked not-applicable, no exception needed |
| A compensating control already covers the risk | a rule that duplicates a network-level restriction | name the control and how it is verified |
| Cost clearly exceeds benefit now | a Level 2 rule on a fleet where it needs a reboot per server | revisit date |
Each exception carries an owner, the exact hosts, the reason, the compensating control and an expiry date (I would default to six months) so it forces a decision again. An exception without an expiry is a permanent hole.
Patch windows and emergency changes
Align the monthly patch window with the rings, so a baseline change and a patch do not land on the same host in the same night, which makes failures impossible to attribute. For an actively exploited vulnerability, keep a pre-approved emergency path outside the monthly cycle; an out-of-band change that touches a hardened setting creates a time-limited exception (72 hours, for example) rather than a silent edit, and the scheduled enforcement run is paused for that host until it expires.
The first 90 days
| Days | Deliverable | Exit test |
|---|---|---|
| 1 to 30 | reconciled inventory, owners, baseline chosen per OS, assessment-only scans on every server, tier 1 list agreed | every server has an owner and a measured pass rate |
| 31 to 60 | golden images rebuilt with tier 1; enforcement live on rings 0 and 1 (120 servers); exception process running | zero baseline-caused outages in rings 0 and 1 |
| 61 to 90 | rings 2 and 3 (the other 880); scheduled scans and dashboard; drift alerts | proposed target: at least 95% of applicable tier 1 rules passing fleet-wide, every failure either ticketed or excepted |
The 95% figure is a target to agree, not a measurement.
Boundary
Container images and the cluster that schedules them are hardened separately. This programme covers the host operating systems that run them, since a container runtime inherits the host's weaknesses.
Pitfalls
- Enforcing everything in Level 1 on day one because "it is only Level 1".
- Counting not-applicable rules as passes, or excluding failing servers from the denominator.
- Two baselines applied to one host (a CIS GPO plus a Microsoft one) that silently overwrite each other.
- Exceptions with no expiry, which turn the baseline into a suggestion.
Write a script that compares a security baseline for a Linux host (SSH settings, sysctl values, file modes, enabled services) against the host's current state and reports drift as a short human-readable summary plus machine-readable JSON. The baseline can have missing keys and per-setting defaults. How do you structure the comparison so that adding a new kind of check is easy?
Sample Answer
Direct answer
Read a JSON baseline, run one small check function per kind of setting (SSH, sysctl, file mode, service), compare each observed value with the expected one, and give every check one of four statuses: ok, drift, not_applicable or error (could not be evaluated). Print a short summary of the problems and write the full result as JSON. Adding a new kind of check means writing one function and registering it with a decorator: the loader, comparison engine, summary and JSON writer do not change. The script only reads (it opens files, calls os.stat, and runs the query systemctl is-enabled), so it never repairs anything.
Plain meanings of the terms. sshd is the SSH server and sshd_config its settings file. A sysctl value is a kernel setting, readable as a file under /proc/sys where the dots in the name become slashes (net.ipv4.ip_forward is /proc/sys/net/ipv4/ip_forward). A file mode such as 0640 means owner read and write, group read, others nothing. systemctl is the command line of the systemd service manager. A Match line in sshd_config starts a conditional block whose settings apply only to connections that match it.
Design decisions
| Decision | Choice | Reason |
|---|---|---|
| Baseline format | JSON with a defaults block and a checks list; every check has an id, a type and an expected value (or max_mode for file modes) | Defaults for severity and on_missing are stated once; any check can override them |
| Missing setting | Three outcomes: the check supplies a default (the daemon's documented default is then compared), or on_missing is not_applicable, or the setting counts as drift | A missing key is only drift when the baseline says the absence matters. For example PermitEmptyPasswords defaults to no in sshd_config(5), so an absent line is compliant when the baseline expects no |
| Expecting absence | "expected": "absent" | For something that should not exist (a telnet socket), a Missing result is the compliant outcome and is reported ok; if the thing exists, its state (for example enabled) is compared with absent and reported as drift |
| Extensibility | @check("kind") registry: a check function returns the observed value or raises Missing or CannotEvaluate | The engine never needs to know what a kind does |
| File modes | max_mode means "no more permissive than this", so 0600 passes a 0640 limit; exact expected is also supported | A file that is stricter than the baseline is not drift in most baselines |
| JSON schema | schema_version, host, summary counts, results list with fixed keys and a fixed status vocabulary, sorted keys, no timestamp | Stable output diffs cleanly and downstream tooling can rely on it; the collector adds the timestamp |
| Exit codes | 0 clean, 1 drift found, 2 nothing drifted but some checks could not run | A scheduler can alert on non-zero and tell the cases apart |
| Read-only | No writes except the report file the caller names | Detect only; remediation is a separate, reviewed step |
The script
Two Python constructs carry the design. CHECKS is a dictionary from a kind name to a function, and @check("sshd") above a function is a decorator: it runs CHECKS["sshd"] = check_sshd when the file loads, so registering a new kind is one line above its function. class Missing(Exception) and CannotEvaluate are custom error types: a check function signals "the thing is not there" or "I could not look" by raising one, and evaluate catches them in one place and turns them into statuses, so no check function needs to know about statuses.
#!/usr/bin/env python3
"""Read-only drift check of a Linux host against a JSON security baseline."""
import argparse
import json
import os
import socket
import subprocess
import sys
CHECKS = {}
def check(kind):
"""Register a check type. Adding a new kind of check means adding one function."""
def register(fn):
CHECKS[kind] = fn
return fn
return register
class Missing(Exception):
"""The thing being checked does not exist on this host."""
class CannotEvaluate(Exception):
"""The check could not run (tool absent, permission denied)."""
def norm(value):
return " ".join(str(value).split()).lower()
@check("sshd")
def check_sshd(spec, ctx):
path = ctx["root"] + spec.get("file", "/etc/ssh/sshd_config")
try:
with open(path, encoding="utf-8") as fh:
lines = fh.read().splitlines()
except FileNotFoundError:
raise Missing(path)
except PermissionError as exc:
raise CannotEvaluate(str(exc))
wanted = spec["key"].lower()
for line in lines:
parts = line.split(None, 1)
if not parts or parts[0].startswith("#"):
continue
word = parts[0].lower()
if word == "match":
break # keywords after Match are conditional, not global
if word == wanted and len(parts) == 2:
return parts[1].strip() # first value wins, as in sshd_config(5)
raise Missing(spec["key"])
@check("sysctl")
def check_sysctl(spec, ctx):
path = ctx["root"] + "/proc/sys/" + spec["key"].replace(".", "/")
try:
with open(path, encoding="utf-8") as fh:
return fh.read().strip()
except FileNotFoundError:
raise Missing(spec["key"])
except PermissionError as exc:
raise CannotEvaluate(str(exc))
@check("file_mode")
def check_file_mode(spec, ctx):
path = ctx["root"] + spec["path"]
try:
return format(os.stat(path).st_mode & 0o7777, "04o")
except FileNotFoundError:
raise Missing(path)
except PermissionError as exc:
raise CannotEvaluate(str(exc))
@check("service")
def check_service(spec, ctx):
try:
proc = subprocess.run(["systemctl", "is-enabled", spec["name"]],
capture_output=True, text=True, timeout=10)
except FileNotFoundError:
raise CannotEvaluate("systemctl not found")
except subprocess.TimeoutExpired:
raise CannotEvaluate("systemctl timed out")
state = (proc.stdout.strip().splitlines() or [""])[0]
if state == "not-found" or "No such file or directory" in proc.stderr:
raise Missing(spec["name"]) # systemd 255+ prints not-found; systemd 252 prints the error text
if not state:
raise CannotEvaluate(proc.stderr.strip() or "no output from systemctl")
return state
def compare(spec, actual):
if spec["type"] == "file_mode" and "max_mode" in spec:
extra = int(actual, 8) & ~int(spec["max_mode"], 8)
return extra == 0
return norm(actual) == norm(spec["expected"])
def evaluate(spec, defaults, ctx):
rec = {"id": spec["id"], "type": spec["type"],
"severity": spec.get("severity", defaults.get("severity", "medium")),
"expected": spec.get("expected", spec.get("max_mode")),
"actual": None, "status": None, "note": ""}
on_missing = spec.get("on_missing", defaults.get("on_missing", "drift"))
try:
fn = CHECKS[spec["type"]]
except KeyError:
rec["status"], rec["note"] = "error", "unknown check type"
return rec
try:
actual = fn(spec, ctx)
except Missing:
if spec.get("expected") == "absent":
rec["actual"], rec["status"], rec["note"] = "absent", "ok", ""
return rec
if "default" in spec: # setting absent: the program's documented default applies
actual = spec["default"]
rec["note"] = "absent, using stated default"
elif on_missing == "not_applicable":
rec["status"], rec["note"] = "not_applicable", "absent on this host"
return rec
else:
rec["status"], rec["note"] = "drift", "absent"
return rec
except CannotEvaluate as exc:
rec["status"], rec["note"] = "error", str(exc)
return rec
rec["actual"] = actual
rec["status"] = "ok" if compare(spec, actual) else "drift"
return rec
def main(argv=None):
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("baseline")
ap.add_argument("--root", default="", help="prefix for file paths (testing)")
ap.add_argument("--host", default=socket.gethostname())
ap.add_argument("--json-out", required=True)
args = ap.parse_args(argv)
with open(args.baseline, encoding="utf-8") as fh:
baseline = json.load(fh)
defaults = baseline.get("defaults", {})
ctx = {"root": args.root}
results = [evaluate(s, defaults, ctx) for s in baseline["checks"]]
counts = {k: 0 for k in ("ok", "drift", "not_applicable", "error")}
for r in results:
counts[r["status"]] += 1
report = {"schema_version": 1, "host": args.host, "summary": counts,
"results": results}
with open(args.json_out, "w", encoding="utf-8") as fh:
json.dump(report, fh, indent=2, sort_keys=True)
fh.write("\n")
print(f"{args.host}: {counts['drift']} drift, {counts['ok']} ok, "
f"{counts['not_applicable']} n/a, {counts['error']} could not check")
for r in results:
if r["status"] in ("drift", "error"):
print(f" [{r['severity']}] {r['id']}: {r['status']} "
f"(expected {r['expected']}, actual {r['actual']}) {r['note']}".rstrip())
if counts["drift"]:
return 1
return 2 if counts["error"] else 0
if __name__ == "__main__":
sys.exit(main())
The baseline
{
"defaults": {"severity": "medium", "on_missing": "drift"},
"checks": [
{"id": "ssh-permit-root-login", "type": "sshd", "key": "PermitRootLogin", "expected": "no", "severity": "high"},
{"id": "ssh-password-auth", "type": "sshd", "key": "PasswordAuthentication", "expected": "no", "severity": "high"},
{"id": "ssh-empty-passwords", "type": "sshd", "key": "PermitEmptyPasswords", "expected": "no", "default": "no", "severity": "high"},
{"id": "ssh-x11-forwarding", "type": "sshd", "key": "X11Forwarding", "expected": "no", "default": "no", "severity": "low"},
{"id": "sysctl-ip-forward", "type": "sysctl", "key": "net.ipv4.ip_forward", "expected": "0"},
{"id": "sysctl-rp-filter", "type": "sysctl", "key": "net.ipv4.conf.all.rp_filter", "expected": "1"},
{"id": "sysctl-ipv6-ra", "type": "sysctl", "key": "net.ipv6.conf.all.accept_ra", "expected": "0", "on_missing": "not_applicable"},
{"id": "mode-shadow", "type": "file_mode", "path": "/etc/shadow", "max_mode": "0640", "severity": "high"},
{"id": "mode-sshd-config", "type": "file_mode", "path": "/etc/ssh/sshd_config", "max_mode": "0644"},
{"id": "mode-gshadow", "type": "file_mode", "path": "/etc/gshadow", "max_mode": "0640", "on_missing": "not_applicable"},
{"id": "svc-telnet", "type": "service", "name": "telnet.socket", "expected": "absent", "severity": "high"}
]
}
Each sshd check reads the global part of sshd_config. The OpenSSH manual states "Unless noted otherwise, for each keyword, the first obtained value will be used", so the parser returns the first occurrence of a keyword and ignores later duplicates (a few keywords accumulate instead, which is one more reason to prefer sshd -T for the effective answer). It stops at the first Match line because keywords after it only apply to matching connections. Include files are not followed. When you need the effective configuration including includes and Match evaluation, run sshd -T on the host and parse that instead. Executed in a debian:12-slim container with openssh-server installed (after mkdir -p /run/sshd and ssh-keygen -A), with an sshd_config holding PermitRootLogin yes, PasswordAuthentication=no, PermitRootLogin no and a Match User x block that sets PasswordAuthentication yes, the command sshd -T | grep -Ei '^(permitrootlogin|passwordauthentication|permitemptypasswords)' printed:
permitrootlogin yes
passwordauthentication no
permitemptypasswords no
The output is lowercase keyword value, the first PermitRootLogin wins, the Match block is not applied to the global value, and an unset keyword shows its built-in default. sshd also accepts = as the separator (PasswordAuthentication=no above), but the parser shown here splits on whitespace only, so a line written with = is read as missing; normalize the separator before comparing or use sshd -T.
Running it
The demo builds a throwaway directory tree that imitates a host, so the result is reproducible on any machine with Docker. Save drift_check.py, baseline.json and the script below in one directory (the demo script is named demo.sh):
set -u
R=/tmp/demo-root
mkdir -p $R/etc/ssh $R/proc/sys/net/ipv4/conf/all
cat > $R/etc/ssh/sshd_config <<'CFG'
# demo config
PermitRootLogin yes
PasswordAuthentication no
PermitRootLogin no
Match User backup
PasswordAuthentication yes
CFG
chmod 0644 $R/etc/ssh/sshd_config
echo 1 > $R/proc/sys/net/ipv4/ip_forward
echo 1 > $R/proc/sys/net/ipv4/conf/all/rp_filter
echo "root:*:19000:0:99999:7:::" > $R/etc/shadow; chmod 0644 $R/etc/shadow
python3 drift_check.py baseline.json --root $R --host demo-host --json-out report.json
echo "exit=$?"
cat report.json | head -40
docker run --rm -v "$PWD":/w -w /w python:3.12-slim sh demo.sh
Printed output (the script then prints the first lines of report.json, which are omitted here):
demo-host: 3 drift, 5 ok, 2 n/a, 1 could not check
[high] ssh-permit-root-login: drift (expected no, actual yes)
[medium] sysctl-ip-forward: drift (expected 0, actual 1)
[high] mode-shadow: drift (expected 0640, actual 0644)
[high] svc-telnet: error (expected absent, actual None) systemctl not found
exit=1
The summary counts (3 drift, 5 ok, 2 n/a, 1 could not check) add up to the 11 checks in the baseline, and the JSON summary object holds the same numbers: {'drift': 3, 'error': 1, 'not_applicable': 2, 'ok': 5}.
Reading the result:
ssh-permit-root-loginisdrifteven though a later line saysPermitRootLogin no: the first line wins.ssh-password-authisok: theMatch User backupblock that setsPasswordAuthentication yesis ignored by design because it only applies to that user.ssh-empty-passwordsandssh-x11-forwardingareokthrough the stated defaults (no), and the JSON records the noteabsent, using stated default.mode-shadowisdrift(0644is more permissive than the0640limit);mode-sshd-configpasses at0644.svc-telnetiserror: the container has nosystemctl, so the check reports that it could not run. It is not counted as compliant.sysctl-ipv6-raandmode-gshadowarenot_applicablebecause those paths are absent here and the baseline says absence is fine.
How the file-mode comparison works
compare checks that a file has no permission bit beyond the limit. ~limit flips every bit of the limit, and & keeps only bits that are set in both numbers, so actual & ~limit is the set of permission bits the file has that the limit does not allow. Zero means the file is within the limit. Executed in a python:3.12-slim container with the limit 0640 (the nine bits are owner rwx, group rwx, other rwx):
actual 0600 = 110000000 limit 0640 = 110100000 bits beyond limit = 000000000 -> ok
actual 0640 = 110100000 limit 0640 = 110100000 bits beyond limit = 000000000 -> ok
actual 0644 = 110100100 limit 0640 = 110100000 bits beyond limit = 000000100 -> drift
actual 0660 = 110110000 limit 0640 = 110100000 bits beyond limit = 000010000 -> drift
0644 fails because it adds the "other read" bit (the last 1 in the third line's beyond-limit column), and 0660 fails because it adds group write. 0600 passes because it only removes bits.
Adding a new kind of check
A check to confirm a line exists in a file is one function. The program below registers it and reuses main:
set -u
R=/tmp/demo-root
mkdir -p $R/etc
printf 'blacklist usb-storage\n' > $R/etc/modprobe-hardening.conf
cat > extra_check.py <<'PY'
import sys
import drift_check
from drift_check import check, Missing
@check("file_line")
def check_file_line(spec, ctx):
path = ctx["root"] + spec["path"]
try:
with open(path, encoding="utf-8") as fh:
lines = fh.read().splitlines()
except FileNotFoundError:
raise Missing(path)
return "present" if spec["line"] in lines else "absent"
sys.exit(drift_check.main())
PY
cat > extra.json <<'JSON'
{"checks": [
{"id": "usb-storage-blocked", "type": "file_line", "path": "/etc/modprobe-hardening.conf", "line": "blacklist usb-storage", "expected": "present"},
{"id": "cramfs-blocked", "type": "file_line", "path": "/etc/modprobe-hardening.conf", "line": "blacklist cramfs", "expected": "present", "severity": "low"}
]}
JSON
python3 extra_check.py extra.json --root $R --host demo-host --json-out extra-report.json
echo "exit=$?"
Output of that run:
demo-host: 1 drift, 1 ok, 0 n/a, 0 could not check
[low] cramfs-blocked: drift (expected present, actual absent)
exit=1
The first check passes because the file contains the line; the second reports drift. Nothing outside the new function and the baseline entries changed.
Complexity and edge cases
- Complexity: each check does constant work apart from I/O; an
sshdcheck readssshd_configonce per check, so N sshd checks over an L-line file cost O(N x L). Caching the parsed file per path makes it O(L + N) if the baseline grows large. - Edge cases handled: missing keys, missing files, unreadable files (reported as
error, not as pass), repeated keywords, comments, multi-value sysctl output (whitespace is normalized), a service that does not exist (newer systemd, 255 and later, printsnot-foundon stdout for an unknown unit; systemd 252, which Debian 12 ships, prints nothing on stdout andFailed to get unit file state for telnet.socket: No such file or directoryon stderr; both are treated as missing, which matters because without the second test an absent unit on a 252 host would be reported aserrorinstead of compliant), and an absent tool. - Not handled, and worth stating:
IncludeandMatchevaluation (usesshd -T), runtime sysctl values that differ from the persisted files in/etc/sysctl.d(add asysctl_persistedcheck kind to catch a setting that will revert at reboot), symlinks (os.statfollows them), and extended attributes or ownership (addownerchecks as another registered kind).
Unlock Full Question Bank
Get access to all 22 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.