System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
What are CIS Benchmarks, and how would you use them to build and maintain secure baselines for a mixed Windows and Linux estate? What do you do when a recommendation breaks something you depend on?
Sample Answer
Direct answer
The CIS Benchmarks are prescriptive, consensus-developed configuration guides published by the Center for Internet Security (CIS) for operating systems, cloud platforms, databases, network devices and applications. CIS describes them as consensus-based work by cybersecurity experts worldwide, and its catalogue lists well over a hundred of them. Each one lists individual recommendations: a setting, why it matters, how to audit it and how to remediate it. For a mixed Windows and Linux estate I would pick the benchmark that matches each operating system and exact version, choose a profile per server role, test it, deploy it as code, scan against it continuously, and run a formal exception process for the recommendations that break something we depend on. I would not apply a benchmark blindly, and I would never silently skip an item.
What one recommendation looks like
Each benchmark item has the same parts: a number and title, a description, a rationale, an audit procedure and a remediation. Here is one, a Linux item from the CIS Ubuntu Linux 22.04 LTS Benchmark as quoted in the public Wazuh policy file, which implements benchmark v1.0.0 and numbers the item 1.1.2.3 (the Ansible Lockdown automation, which at the time of checking targets v3.0.0 of the same benchmark, numbers it 1.1.2.1.4; item numbers move between benchmark versions, which is why a finding must always cite the benchmark version as well as the number):
| Part | Content |
|---|---|
| Title | Ensure noexec option set on /tmp partition |
| Description | The noexec mount option specifies that the filesystem cannot contain executable binaries |
| Rationale | /tmp is only meant for temporary file storage, so users should not be able to run executable binaries from it |
| Audit | findmnt --kernel /tmp and check that noexec appears in the options |
| Remediation | add noexec to the options field of the /tmp line in /etc/fstab, for example tmpfs /tmp tmpfs defaults,rw,nosuid,nodev,noexec,relatime 0 0, then mount -o remount /tmp |
The audit step reads the mount's options. I ran the equivalent on a throwaway tmpfs mount carrying those options:
$ findmnt --kernel /mnt/tmpdemo
TARGET SOURCE FSTYPE OPTIONS
/mnt/tmpdemo tmpfs tmpfs rw,nosuid,nodev,noexec,relatime
$ /mnt/tmpdemo/setup.sh
bash: line 6: /mnt/tmpdemo/setup.sh: Permission denied
$ sh /mnt/tmpdemo/setup.sh
installer ran
The noexec in the OPTIONS column is the pass condition. The last two commands show the effect, and its limit: the kernel refuses to execute the file directly, but handing the script to an interpreter (sh script) still works, because the interpreter only reads the file. That is why a benchmark treats one setting as a layer, not a seal.
The same benchmark family also shows the profile split. In the Ansible Lockdown role, which tags each rule with the benchmark's level, the noexec item above is tagged Level 1 (server and workstation), while "Ensure squashfs kernel module is not available" (rule 1.1.1.7 in that role's v3.0.0 numbering) is tagged Level 2 (server and workstation). The first is a low-risk mount flag; the second removes support for a whole filesystem type, so it needs a check that nothing on the host mounts that format first.
Profiles: Level 1 and Level 2
CIS defines two profile levels. Level 1 is described by CIS as a base recommendation that can be implemented fairly promptly and is designed not to have an extensive performance impact. Level 2 is described as "defense in depth" and is intended for environments where security is paramount; applying it can negatively affect operations if not carefully deployed. Benchmarks for some platforms also split profiles by role (for example server versus workstation), so the profile name on the document you download tells you which applies.
| Estate segment | Starting profile | Reason |
|---|---|---|
| General application servers, Windows and Linux | Level 1 | Low operational risk, fast to roll out, closes the common gaps |
| Domain controllers, bastion hosts, systems holding regulated data | Level 1, then selected Level 2 items after testing | Higher impact if compromised, so the extra items are worth the testing cost |
| Developer workstations | Level 1, with documented exceptions | Developer tooling is the commonest source of broken recommendations |
How I build and maintain the baselines
-
Inventory and scope. List operating systems and versions in use, because a benchmark is written for one product version. A host on a version with no benchmark yet gets the nearest one, flagged as such.
-
Pick the profile per role (table above) and record the decision.
-
Build the baseline as code, not as a document. Windows: the benchmark settings become Group Policy Objects (GPOs, the Active Directory mechanism for pushing Windows settings). The Windows and Linux tools each do a different job:
Need Tool Start from Microsoft's own recommended Windows settings as importable GPO backups Microsoft's Security Compliance Toolkit See where your existing GPOs differ from a baseline Policy Analyzer (part of the toolkit) Apply a policy file to a single machine, for testing or for a host not joined to the domain LGPO.exe(part of the toolkit)Measure a host against a CIS benchmark and get a pass or fail per recommendation CIS-CAT Pro Assessor, CIS's own assessment tool, or another scanner that imports the benchmark Linux: apply settings as code Ansible, Puppet or similar, or a hardened base image Buy ready-made implementations of the benchmark CIS Build Kits (GPOs and scripts) and Hardened Images A first trial on one Windows test machine needs two things from this list: a way to apply settings (
LGPO.exe) and a way to measure them (the assessor); the toolkit's baselines are Microsoft's own, not CIS's, so record which of the two a host group follows. Policy Analyzer and the Build Kits matter once the settings are being managed across many hosts. Bought kits and images still need the testing step below. -
Test in a pilot ring. A small group of representative hosts (including one running each critical application) gets the baseline first, followed by the application's smoke tests.
-
Deploy in rings (pilot, then a slice of production, then everything), as for patches.
-
Scan and report. A conformance scanner such as CIS-CAT Pro Assessor (CIS's own tool for assessing a system against a benchmark) or another tool that imports the benchmark produces a pass rate per host and per recommendation. I track the pass rate by host group, and the list of failing recommendations that are not approved exceptions.
-
Re-baseline on change. New OS versions, new benchmark releases and newly deployed applications all trigger a review. A baseline that is two benchmark versions old is a finding.
When a recommendation breaks something
The sequence matters: do not turn the setting back off first.
- Reproduce and pin down. Confirm in the pilot ring that this exact setting, and not something else deployed that week, causes the failure. Capture the error.
- Look for a fix that keeps the control. Often the dependency is what is wrong: a legacy service account, a vendor installer that unpacks in a temporary folder, an old protocol. Upgrading or reconfiguring the dependency satisfies the recommendation permanently.
- Narrow the scope. If it only breaks one application, apply the setting everywhere except that host group (a separate organisational unit for Windows, a host variable for Linux), rather than dropping it estate-wide.
- Add a compensating control (a different safeguard that covers the same risk): for example, if a recommendation to restrict an old authentication protocol cannot be applied to one legacy application server, put that server in its own network segment, restrict who can reach it, and add monitoring for use of the old protocol.
- Record a time-boxed exception, owned by a named person and approved by the person who owns the risk (not by the engineer who wants the exception). The scanner must show it as an accepted deviation so it does not look like an unknown failure.
An exception record has these fields:
| Field | Example content |
|---|---|
| Recommendation | The benchmark item title and version, copied from the document |
| Hosts affected | The exact host group, not "some servers" |
| What breaks | The application and the observed error |
| Compensating control | The segment restriction and the monitoring rule |
| Risk owner and approver | Named people |
| Review date | A calendar date, normally within the year or sooner |
| Exit plan | The work that would let the exception be removed |
Worked example: the noexec item above. Suppose a vendor installer unpacks a helper and runs it directly from /tmp. After the baseline lands, the pilot ring catches it: running the helper fails with Permission denied, as in the output above, and the installer aborts. Options, in the order of the process above:
- Make the dependency comply (best): the vendor ships an installer that honours
TMPDIR, so it unpacks into a directory you choose. - Scope an exception to the few hosts that run the installer: give them a dedicated mount, for example
/var/tmp/installerwithoutnoexec, readable and writable only by the install account, while/tmpstaysnoexeceverywhere. - Move installation into the image build, so production hosts never run the installer at all.
The exception record names the second or third option, the owner, and the review date.
Pitfalls
- Treating the pass percentage as the goal. A 98 percent score with five unreviewed failures on domain controllers is worse than 92 percent with every gap understood.
- Applying a benchmark written for one OS version to another.
- Benchmarks are a configuration baseline, not a complete security programme: they say nothing about patch currency, identity hygiene or detection.
- Level 2 applied everywhere without testing; it is explicitly the level that can hurt operations.
Write PowerShell that audits the local Administrators group across domain-joined Windows servers and reports hosts with unexpected members. How do you handle credentials and unreachable hosts?
Sample Answer
Direct answer
Fan out one read-only PowerShell remoting call (PowerShell's way of running a command on another computer over WinRM, Windows Remote Management) to every member server, have each server return exactly one object (either its member list or the error it hit), compare the members to an allow-list on the auditing side, and report four statuses: OK, UNEXPECTED, ERROR and UNREACHABLE. The two design rules are that the script never handles a password (it runs under the operator's or a service account's Kerberos identity: Kerberos is the Windows domain sign-in protocol, and the script simply borrows the ticket of whoever runs it), and that a host which did not answer is reported as UNREACHABLE rather than silently counted as clean. Silence must never read as a pass.
The script
#Requires -Version 5.1
<#
Audit the local Administrators group on domain-joined member servers.
Runs under the caller's Kerberos identity (no passwords handled).
#>
[CmdletBinding()]
param(
[string[]]$ComputerName,
# Principals that are expected in local Administrators on every server.
[string[]]$Allowed = @('LOCAL\Administrator', 'CORP\Domain Admins', 'CORP\SRV-Local-Admins'),
[int]$ThrottleLimit = 32,
[int]$ConnectTimeoutSec = 20
)
# Pure comparison logic: no remoting, so it can be tested anywhere.
function Compare-AdminMember {
param(
[Parameter(Mandatory)][string[]]$Actual,
[Parameter(Mandatory)][string[]]$Allowed,
# Host being audited: HOST\name is rewritten to LOCAL\name so one allow-list fits every server.
[string]$HostName
)
$allowedSet = [System.Collections.Generic.HashSet[string]]::new(
[string[]]$Allowed, [System.StringComparer]::OrdinalIgnoreCase)
# Local accounts are prefixed with the short (NetBIOS) computer name, even when the host is addressed by FQDN.
$shortName = $HostName.Split('.')[0]
foreach ($member in $Actual) {
$normalised = $member -replace ('^' + [regex]::Escape($shortName) + '\\'), 'LOCAL\'
if (-not $allowedSet.Contains($normalised)) { $member }
}
}
if (-not $ComputerName) {
$dcNames = (Get-ADDomainController -Filter *).Name
$ComputerName = (Get-ADComputer -Filter 'OperatingSystem -like "*Server*" -and Enabled -eq $true').Name |
Where-Object { $_ -notin $dcNames }
}
$probe = {
# One object per host, success or failure, so silence can never mean "clean".
try {
$members = Get-LocalGroupMember -SID 'S-1-5-32-544' -ErrorAction Stop
[pscustomobject]@{
Error = $null
Members = @($members | ForEach-Object { $_.Name })
}
}
catch {
[pscustomobject]@{ Error = $_.Exception.Message; Members = @() }
}
}
$so = New-PSSessionOption -OpenTimeout ($ConnectTimeoutSec * 1000) -NoMachineProfile
$raw = Invoke-Command -ComputerName $ComputerName -ScriptBlock $probe `
-SessionOption $so -ThrottleLimit $ThrottleLimit `
-ErrorAction SilentlyContinue -ErrorVariable remotingErrors
$remotingErrors | ForEach-Object { Write-Verbose $_.Exception.Message }
$reached = @($raw | Select-Object -ExpandProperty PSComputerName -Unique)
$report = foreach ($r in $raw) {
$extra = if ($r.Error) { @() } else {
@(Compare-AdminMember -Actual $r.Members -Allowed $Allowed -HostName $r.PSComputerName)
}
[pscustomobject]@{
Host = $r.PSComputerName
Status = if ($r.Error) { 'ERROR' } elseif ($extra.Count) { 'UNEXPECTED' } else { 'OK' }
Unexpected = $extra -join '; '
Detail = $r.Error
}
}
$unreachable = $ComputerName | Where-Object { $_ -notin $reached } | ForEach-Object {
[pscustomobject]@{ Host = $_; Status = 'UNREACHABLE'; Unexpected = ''; Detail = 'no result returned (see -Verbose)' }
}
@($report) + @($unreachable) | Sort-Object Status, Host
Run it as .\Audit-LocalAdmins.ps1 -Verbose | Export-Csv admins.csv -NoTypeInformation from a machine with the Active Directory PowerShell module (installed with RSAT, the Remote Server Administration Tools) and an account that is allowed to open remoting sessions to the servers. Pass -ComputerName for a subset; with no parameter it discovers every enabled server computer account in the Active Directory (AD) domain except domain controllers.
What the report looks like. The four rows below are illustrative: the host names are made up, and I produced them by running the script's own comparison and report-building lines (unchanged, in PowerShell 7) on sample probe results, because the remoting itself needs Windows servers. The CSV is sorted by Status then Host:
"Host","Status","Unexpected","Detail"
"sql01","ERROR","","Access is denied."
"app01","OK","",
"file01","UNEXPECTED","CORP\jsmith; FILE01\svc_backup",
"web02","UNREACHABLE","","no result returned (see -Verbose)"
Reading it: OK means the host answered and every member is on the allow-list; UNEXPECTED lists the members that are not (Unexpected column), which is what you investigate; ERROR means the host answered but the group read failed, with the reason in Detail (the error text in the sample row is made up); UNREACHABLE means no result came back from that host: it did not answer, or it answered but refused the connection (for example access denied at logon), so nothing is known about its group. The reason, where Windows gave one, is in $remotingErrors and is printed with -Verbose. The empty cells are normal.
Why it is built this way
Reading the group. Get-LocalGroupMember -SID 'S-1-5-32-544' addresses the local Administrators group by its well-known SID (security identifier), so it still works on a server whose OS language renames the group. Note on scope: the Microsoft.PowerShell.LocalAccounts module is not available in 32-bit PowerShell on a 64-bit system. Domain controllers are excluded because they have no local account database of their own: the Administrators group there is a domain-wide built-in group, which needs a separate AD audit.
Credentials.
- No password appears in the script, a parameter, or a file.
Invoke-Commanduses the caller's current Kerberos credentials, and the script block touches nothing beyond the local server, so no credential has to be forwarded to a third machine, and CredSSP is not needed. The "second hop" picture: you run the script on workstation A, and it connects to server B (the first hop). If code running on B then had to reach a file share or server C (the second hop), B would need your credentials to authenticate there, and Kerberos does not hand B a reusable copy of them by default, so that call fails. CredSSP is the setting that works around this by sending your actual credentials to B, which is why it is best avoided. Here B only reads its own local group, so there is no second hop. - For a one-off with a different account,
Invoke-Command -Credential (Get-Credential)prompts and keeps the secret in aPSCredentialobject only in memory. - For a scheduled run, run the task as a group managed service account (gMSA): Active Directory rotates its password, so no secret is stored anywhere. Give that account the least privilege that allows opening a remoting session and reading local groups, and test it. If local administrator rights turn out to be required, treat the account as a privileged identity: it should not log on interactively anywhere, and its use should be alerted on, because an account that can read every server's admins can usually do more.
Unreachable hosts. A connection failure in Invoke-Command is a non-terminating error: the hosts that answer still return their results and the command continues. That is convenient, but it also means a dead host contributes nothing to the output. The script therefore computes UNREACHABLE as the requested names minus the names that returned an object (PSComputerName is added to every remote result). Reading the Invoke-Command call: -ComputerName and -ScriptBlock say where to run the probe and what to run; -SessionOption $so applies the timeout and no-profile settings; -ErrorAction SilentlyContinue stops each connection failure from printing a red error and halting the loop, so reachable hosts still return results; -ErrorVariable remotingErrors (written without a $) captures those same errors into the variable $remotingErrors so they are not lost, and -Verbose prints them. Errors are silenced, but hosts are never lost, because the UNREACHABLE list is computed afterwards from the names that did not answer. New-PSSessionOption -OpenTimeout is in milliseconds and defaults to 180000 (three minutes); lowering it to 20 seconds keeps dead hosts from stalling the run, and -ThrottleLimit caps concurrent connections (default 32). Error text from failed connections is available with -Verbose.
Errors on reachable hosts. The probe wraps the cmdlet in try/catch with -ErrorAction Stop and returns the message, so a host where the call failed shows as ERROR with its reason instead of an empty, apparently clean list.
The comparison is separate and testable. Compare-AdminMember takes the actual members, the allow-list and the host name, rewrites HOST\name to LOCAL\name (the local prefix is the short computer name, even when the host is addressed by its FQDN, fully qualified domain name) and returns whatever is not on the list. Testing it needs no Windows host. The script cannot simply be loaded into a test session, because loading it would run the remoting. So the first line uses Parser.ParseFile to read the script file and build a syntax tree (an AST, abstract syntax tree) without executing anything; the second line asks the tree for the first function definition in it, which is Compare-AdminMember; and the third wraps that function's text in a script block and runs it with a leading dot, which defines just that function in the current session:
$ast=[System.Management.Automation.Language.Parser]::ParseFile("$PSScriptRoot/Audit-LocalAdmins.ps1",[ref]$null,[ref]$null)
$fn=$ast.Find({param($n) $n -is [System.Management.Automation.Language.FunctionDefinitionAst]},$true)
. ([scriptblock]::Create($fn.Extent.Text))
$allowed='LOCAL\Administrator','CORP\Domain Admins','CORP\SRV-Local-Admins'
$actual='FILE01\Administrator','CORP\Domain Admins','CORP\jsmith','FILE01\svc_backup','corp\srv-local-admins'
Compare-AdminMember -Actual $actual -Allowed $allowed -HostName file01.corp.example.com
CORP\jsmith
FILE01\svc_backup
FILE01\Administrator matched LOCAL\Administrator, CORP\Domain Admins matched exactly, and corp\srv-local-admins matched despite the case difference. Only the two genuinely unexpected members came back. The whole script was parse-checked and its cmdlets and parameters were checked against Microsoft Learn; the remoting path itself needs Windows servers and was not executed here.
Complexity and edge cases
- Work is one remote call per server, run up to
-ThrottleLimitat a time, so the cost grows linearly with the server count; the comparison is O(members) per host with a hash set lookup. - Both lists are compared as strings, so the requested names and the reported
PSComputerNamemust use the same form. The AD discovery path uses theNameproperty throughout; if you supply FQDNs yourself, supply them consistently. - The allow-list matches by name. A renamed built-in Administrator account would be flagged, which is a safe failure but noisy. A hardened version compares SIDs.
- Domain groups in the allow-list are trusted as a unit: the cmdlet reports direct members, so who sits inside
SRV-Local-Adminsmust be audited in AD. - Servers that are powered off, in a different forest, or with WinRM (Windows Remote Management, the transport that PowerShell remoting uses) blocked all land in UNREACHABLE. Follow up on that list; a host that stays unreachable for several runs is a finding in itself.
PrincipalSource(Local, Active Directory, Microsoft Entra group, Microsoft Account) is only populated on Windows 10, Windows Server 2016 and later, so do not rely on it for older servers.
Trade-offs
A scheduled audit like this detects drift but does not prevent it. The preventive control is a Group Policy "Restricted Groups" or Local Users and Groups preference that sets the membership, with this report as the independent check that the policy is applying. Reporting only is deliberate here: an auditing script that also removes members can lock out an admin during an incident.
Design a patch baseline and schedule for a hybrid estate of Windows and Linux servers managed through a cloud update service. Cover update classifications, pre and post scripts, maintenance windows, and machines that are offline when the window opens.
Sample Answer
Direct answer
Use Azure Update Manager (the Azure service that assesses and installs updates on Windows and Linux machines) with one maintenance configuration per patch ring, attach machines to rings by tag through dynamic scoping, and install only the Critical and Security classifications on a monthly cadence for Windows (anchored to "Patch Tuesday", the second Tuesday of the month) and a weekly security-only cadence for Linux. Use "Reboot if required" for stateless tiers and a controlled reboot for databases. Use pre and post events (small automations triggered before and after each run) to power on machines that are normally off, to snapshot or stop services beforehand, and to run health checks afterwards. For machines that are offline when the window opens, rely on three things: a power-on pre-event for machines you control, a catch-up schedule a few days later, and the 24-hour periodic assessment to show you who was missed.
Azure terms in plain words. A maintenance configuration is a saved schedule object (when, for how long, which update classifications, which reboot option) that you attach to machines. Dynamic scoping is a saved filter (subscription, resource group, location, operating system, tags) that decides which machines a maintenance configuration covers, re-evaluated each time it runs. Customer Managed Schedules is a patch-orchestration setting meaning only your schedule patches the machine. An availability set is an Azure grouping that spreads virtual machines across separate hardware, and an update domain is a slice of it that the platform restarts at one time. Azure Resource Graph is Azure's query service over resource metadata and results.
The estate is hybrid: Azure virtual machines are patched directly, and on-premises servers are patched after they are connected through Azure Arc (the service that represents non-Azure servers as Azure resources). Update Manager honors the update source already configured on each machine, so Windows machines keep using Windows Update, Microsoft Update or WSUS (Windows Server Update Services, where you approve updates yourself; Microsoft documents WSUS as deprecated, meaning no new features but still supported with security and quality updates), and Linux machines keep using their configured repositories. It installs what that source offers; it does not publish updates.
1. Baseline: what gets installed
| Decision | Choice | Reason |
|---|---|---|
| Classifications | Critical and Security on both operating systems | Smallest set that closes known exploited flaws; broader categories (feature updates and the rest) go through a separate, slower review |
| Exclusions | Exclude packages or KB (Knowledge Base article) IDs you have found to break an application; Linux accepts package names with wildcards such as kernel* | Lets the database tier hold kernel updates until a controlled window |
| Drivers | Not handled | Update Manager does not support driver updates, so keep a separate process for firmware and drivers |
| Update source | WSUS approvals for Windows if you need to hold a patch; repository snapshots for Linux if you need reproducibility | The service installs what the source offers, so the source is where you control "which patch" |
| Patch orchestration on Azure VMs | Set to Customer Managed Schedules (required for scheduled patching on Azure VMs; not required for Azure Arc-enabled machines) | Otherwise the platform or the guest operating system patches on its own timetable |
What one ring looks like as a configuration (illustrative values)
| Field | Pilot ring, Windows |
|---|---|
| Name | patch-ring-0-windows |
| Scope (dynamic scoping filter) | tag patch-ring = 0, operating system = Windows |
| Schedule | monthly, second Tuesday, offset +1 day, 22:00 |
| Maintenance window length | 3 hours |
| Classifications | Critical, Security |
| Exclusions | KB IDs known to break the application |
| Reboot | Reboot if required |
A pre-event handler is a small piece of code (for example an Azure Function, code Azure runs on demand) subscribed through Event Grid to the pre-maintenance event. In this design it does three things: start machines tagged as normally off, take the data-tier snapshot, and call the cancellation API if either step fails.
Which numbers to design around: the 3-hour window with its 25-minute cut-off (decides when the last install can start), the "machine on at least 15 minutes before the start" rule together with the pre-event period (decide when power-on must begin), and the 24-hour assessment (drives the catch-up schedule). The 40-minute edit rule, the 20-machine one-time update limit and the 7 and 30 day retention periods are operating details to look up when you meet them.
2. Rings, schedules and maintenance windows
A maintenance configuration starts updates on all of its attached machines at the same time, so rings must be separate configurations. Dynamic scoping (selecting machines by subscription, resource group, location, operating system and tags) or the built-in Azure Policy named "Schedule recurring updates using Azure Update Manager" attaches new machines automatically, so a server built next month is patched without anyone remembering it.
| Ring | Tag | Windows schedule | Linux schedule | Reboot |
|---|---|---|---|---|
| Pilot | patch-ring=0 | Monthly, second Tuesday with +1 day offset (Wednesday), 22:00 | Weekly, Wednesday 22:00 | If required |
| Production A | patch-ring=1 | Monthly, second Tuesday with +4 day offset (Saturday), 22:00 | Weekly, Saturday 22:00 | If required |
| Production B (databases, clusters) | patch-ring=2 | Monthly, second Tuesday with +6 day offset (Monday), 02:00 | Weekly, Sunday 02:00 | Never, followed by a planned reboot |
The offsets use the product's own recurrence rule ("nth weekday of the month, with an offset of up to six days either way"); Tuesday plus 1 is Wednesday, plus 4 is Saturday and plus 6 is Monday.
Window length: the portal accepts 1 hour 30 minutes to 3 hours 55 minutes; the schedule needs to repeat no more often than every 6 hours. Choose 3 hours (180 minutes) for Windows. Update Manager checks the remaining time before each step: per the troubleshooting guidance, with fewer than 25 minutes left (15 for the update plus 10 reserved for the reboot) it does not scan or install, and a service pack needs 30. So 180 - 25 = 155 minutes is the latest point at which a new ordinary update can start. Updates already running are never killed, and anything not attempted is reported as "Not attempted", so a window that is too short shows up as a recurring pattern in the history instead of a silent failure. On Linux a reboot needs 15 minutes left in the window on Azure VMs.
Reboot options are Reboot if required, Never reboot and Always reboot. Two details matter. "Never reboot" leaves a machine whose updates need a restart waiting for one that nothing in the run will perform, so the planned reboot for Production B must be a deliberate step (a post-event or an operations task) and not an assumption. And some Windows registry settings (the Windows Update and restart policies) can cause a reboot even when you chose Never reboot, so set those to match.
Machines in a common availability set are not all updated at once: they are updated within update domain boundaries, and machines across multiple update domains are not updated concurrently. If you patch members of one availability set under different schedules and one window overruns, a member can fail or be skipped. Split such machines across schedules at different times and widen the window.
3. Pre and post events
Update Manager uses Azure Event Grid (a service that routes events to handlers such as a webhook or an Azure Function) for this. The documented order is: pre-event, optional cancellation, update installation, post-event.
| Phase | Typical actions here |
|---|---|
| Pre-event | Power on machines that are normally off; take a disk snapshot of the data tier; send a notification; stop an application service that must not be interrupted |
| Cancellation | Your pre-event logic must call the cancellation API itself when a pre-step fails. The service does not cancel automatically, and if you do nothing the run proceeds |
| Post-event | Start services again; run a smoke test; power machines back off; send the patch summary |
Timing facts to design around (from the Microsoft documentation, using its own example of a 3:00 PM start): pre-events run outside the maintenance window, between about 2:30 PM and 2:50 PM, and the cancellation call must be made by 2:50 PM; post-events start as soon as installation ends and can run outside the window; and the status of the run covers update installation only, not the pre and post events, so monitor the event handler separately. If you create or edit a schedule that has a pre-event, do it at least 40 minutes before the start or that run is cancelled automatically.
4. Machines that are offline when the window opens
- Shut-down machines cannot be patched. The machine must be on at least 15 minutes before the scheduled start. Machines may also appear disassociated from their configuration while off; that is a display issue, and the association is still there.
- For machines you control that are off by design (test and development virtual machines), power them on from a pre-event, and start the power-on at the beginning of the pre-event period. For a 3:00 PM start the machine must be running by 2:45 PM (3:00 minus 15 minutes), while the documented pre-event period runs from about 2:30 to 2:50, so a power-on that begins at the end of the period is already too late. Power them off from the post-event.
- For servers that are simply unreachable (a branch server with a network outage), the run misses them. The periodic assessment (every 24 hours when enabled) keeps reporting what is missing, so build two things on it: a dashboard or report of machines with pending Critical and Security updates, and a catch-up maintenance configuration a few days after the main window, scoped to the same tag. The catch-up is safe to run broadly, because every run starts with a fresh assessment, so machines already patched have nothing left to install. For individual stragglers, a one-time update can be started from the portal (up to 20 machines at a time).
- Keep evidence outside the service. Pending-update data is retained for 7 days and installation results for 30 days in Azure Resource Graph (the tables
patchassessmentresourcesandpatchinstallationresources), so export compliance results monthly if audit evidence is needed. - Agent health is the other cause of silent misses: Update Manager installs its own extension on each machine the first time it runs, and Azure Arc machines must be connected. Alert on machines that have not been assessed.
Mapping the design to other tools
The same design uses vendor-neutral parts: a baseline (which classifications and exclusions are approved), rings selected by tag, a schedule per ring, hooks before and after, and an assessment that reports who was missed.
| Part | Azure Update Manager | AWS Systems Manager Patch Manager | Ansible-driven estate |
|---|---|---|---|
| Baseline | classifications and exclusions in the maintenance configuration | a patch baseline (approval rules by classification and severity, plus approved and rejected patch lists) | a package list or security-only update task per play |
| Ring membership | dynamic scoping by tag | targets chosen by tag | inventory groups per ring |
| Schedule | maintenance configuration recurrence | maintenance window or patch policy | the scheduler that runs the play |
| Before and after steps | pre and post events through Event Grid | lifecycle hooks (Systems Manager documents run before and after patching) | tasks before and after the update task |
| Who was missed | periodic assessment every 24 hours | a Scan operation and compliance reports | a read-only check play on a schedule |
Trade-offs and what would change the choices
- Monthly Windows plus weekly Linux security splits the estate by how fast each ecosystem ships fixes. If the audit requires one cadence for both, use weekly Critical and Security for both and accept more reboots.
- Pre-events add moving parts that can fail. Keep them to jobs that really need to happen before patching, and make the failure path explicit (cancel the run, notify).
- If a patch must be validated before it reaches production, hold it at the source (WSUS approval, repository snapshot) and not by delaying a schedule, so the pilot ring still installs what production will later install.
Tell me about a time you introduced a hardening or secure-configuration change in production. How did you validate it, and what was your rollback plan?
Sample Answer
Direct answer
Tell a story with a clear arc: the risk that motivated the change, how you tested it on a small set before production, how you watched it afterwards, and a rollback that was written and rehearsed before the change. The sample below uses a real setting so you can see the shape. Use your own incident in its place.
Situation
Our file servers allowed unsigned SMB (Server Message Block, the Windows file-sharing protocol) traffic. Microsoft documents that unsigned SMB packets can be intercepted and modified in transit, and that on member servers the policy "Microsoft network server: Digitally sign communications (always)" is disabled by default, while it is enabled on domain controllers. Those are the long-standing defaults. Microsoft's SMB documentation says Windows 11 24H2 and Windows Server 2025 changed them (Windows 11 24H2 requires signing both outbound and inbound; Windows Server 2025 requires outbound signing, and Microsoft's pages differ on whether it also requires inbound), so on newer builds test what is actually enforced instead of assuming unsigned traffic is allowed. This story assumes older server builds that still accepted unsigned inbound SMB. Security asked us to require signing on the file servers.
What I did
- Wrote the change: set the policy to Enabled on a pilot OU (organizational unit, a folder of computers in Active Directory) containing two low-risk file servers.
- Identified who could break: older clients and appliances that cannot sign SMB. Microsoft's SMB signing page describes a connection to a server that does not support signing failing (for example with STATUS_INVALID_SIGNATURE), so by the same logic unsigned-only devices would lose access once signing is required. I inventoried clients before the change by turning on SMB signing auditing on the pilot servers for two weeks and listing each client that could not sign.
How the audit works, from Microsoft's SMB signing documentation: on a server build that has the setting (Microsoft documents these audit settings from Windows 11 24H2 and Windows Server 2025), run this in an elevated PowerShell session (or set the Group PolicyComputer Configuration\Administrative Templates\Network\Lanman Server\Audit client does not support signing):
Set-SmbServerConfiguration -AuditClientDoesNotSupportSigning $true
The server then writes an event for each client that connects without signing support to Event Viewer under Applications and Services Logs, Microsoft, Windows, SMBServer, Audit (the documentation lists event IDs 3021 and 3022 for that log). Reading the log after two weeks gives a client list; in this story it named two devices (illustrative names: nas-old-01 and copier-scan-02). If a server build lacks the setting, the fallback is the inventory of client operating systems and appliances plus a test connection from each type.
3. Validated: after enabling on the pilot, I tested from one current Windows 11 laptop, one older Windows 10 laptop, one Linux client mounting the share, and the two flagged devices: from each I opened the share by name, saved a file into it and read it back, so 'shares still opened' meant those three actions succeeded. I also watched file-copy performance. Microsoft's documentation says signing has limited impact on 1 Gb Ethernet but can matter more on faster networks, so I measured large copies on the pilot rather than assuming.
4. Wrote the rollback in advance: set the policy back to Disabled in the pilot GPO (Group Policy Object). Microsoft documents no restart requirement for this policy, so rollback is a policy change, not a reboot, and I confirmed that on the pilot before relying on it.
5. Rolled out in waves, with a communications note and a service-desk script for "cannot open share".
Result and lesson
The pilot surfaced two legacy devices that could not sign. We upgraded one and isolated the other on a separate share with a documented exception and an end date. Looking back, I would have run the audit first on every server group instead of just the pilot, and would have scheduled the first production wave away from month-end close.
Pitfalls
Do not say "no issues." A hardening story without a surprise sounds untested. Keep claims about numbers to what you measured. And say what you would do differently.
Your organisation runs about 1,000 mixed Windows and Linux servers, on-prem and in cloud, with no consistent security baseline. Design the programme that gets them to one and keeps them there: what you enforce first, how you enforce it, how you prove it, and how you decide what to leave out.
Sample Answer
Direct answer
I would run this as a measured programme, not a one-off hardening project. Pick one published baseline per operating system (CIS Benchmarks, the Center for Internet Security's consensus configuration guides, at Level 1 as the starting profile), measure first without changing anything, enforce a short first tier of low-breakage, high-impact settings through the same tooling that builds and manages the servers, prove compliance with automated scans whose results feed a dashboard, and handle every deliberate omission through an exception record with an owner, a compensating control and an expiry date. The baseline is a living definition in version control, so "keeping them there" is a scheduled enforcement and scan loop, not a recurring audit.
For illustration assume 600 Linux and 400 Windows servers (1,000 in total) in a mix of on-premises and cloud. The split is an assumption; the structure does not depend on it.
What the baseline is, and why Level 1 first
CIS defines a Level 1 profile as a base recommendation that can be implemented fairly promptly without extensive performance impact, and Level 2 as defense in depth for environments where security is paramount, which can have an adverse effect if implemented without due care. CIS also advises testing either level in a test environment first. For Windows Server, Microsoft publishes security baselines as Group Policy Object (GPO) backups in its Security Compliance Toolkit, together with Policy Analyzer (compares sets of GPOs against each other or against local policy) and LGPO.exe (applies local policy to machines that are not domain-joined). So the plan is: CIS Level 1 server profile per Linux distribution, and either the Microsoft baseline or the CIS Windows Server benchmark per Windows version, choosing one per OS and not mixing, because overlapping baselines conflict.
Phase 0: discovery and assessment (measure before enforcing)
You cannot enforce what you have not inventoried. Build the server list from more than one source (the configuration management database, the hypervisor and cloud account inventories, the directory) and reconcile it, because the servers nobody remembers are the ones with no baseline. Record for each server: OS and version, owner, environment, internet exposure, whether it can be rebooted, and whether a vendor supports changes to it.
Then run the benchmark in assessment-only mode everywhere. For Linux, OpenSCAP (oscap, an open-source implementation of SCAP, the Security Content Automation Protocol) with the SCAP Security Guide content (the rule definitions) evaluates a profile and writes results without changing the host. A real run in a test AlmaLinux 9 container against the CIS server Level 1 profile:
dnf -y install openscap-scanner scap-security-guide
oscap xccdf eval \
--profile xccdf_org.ssgproject.content_profile_cis_server_l1 \
--results /tmp/results.xml --report /tmp/report.html \
/usr/share/xml/scap/ssg/content/ssg-almalinux9-ds.xml
echo "exit status: $?"
The run printed one block per rule: a Title, the Rule id and a Result. Three of them, copied from the run:
Title Ensure AlmaLinux GPG Key Installed
Rule xccdf_org.ssgproject.content_rule_ensure_almalinux_gpgkey_installed
Result pass
Title Implement Custom Crypto Policy Modules for CIS Benchmark
Rule xccdf_org.ssgproject.content_rule_configure_custom_crypto_policy_cis
Result fail
Title Install AIDE
Rule xccdf_org.ssgproject.content_rule_package_aide_installed
Result notapplicable
pass means the host matches the rule, fail means it does not and needs a fix or an exception, and notapplicable means the rule does not apply to this host (the scanner decides this from conditions in the content). The --profile value is the XCCDF (the XML format that SCAP uses for benchmarks) profile id, and here it selects the CIS server Level 1 rule set out of the content file. The tally was 70 pass, 5 fail, 220 notapplicable, 0 error, with exit status 2: according to the oscap manual, 0 means every rule passed, 1 means an error occurred during evaluation, and 2 means at least one rule failed or was unknown, so a pipeline can gate on the exit status. The honest compliance figure excludes the not-applicable rules from the denominator: 70 / (70 + 5) = 93.3% of the 75 applicable rules, not 70 of 295. Report both numbers, and say which one is the headline, or the dashboard will quietly flatter itself. (CIS-CAT Pro Assessor is CIS's own tool for assessing conformance, and is the alternative where a vendor-certified assessment is wanted.)
The assessment also tells you how bad it is by item: the fleet-wide failure count per rule is your prioritisation list.
What to enforce first (tier 1)
Rank candidate settings by three questions: does it close a path attackers actually use, how likely is it to break something, and can I verify it automatically? The first tier takes settings that score well on all three:
| Area | Linux | Windows | Why first |
|---|---|---|---|
| Remote administration | key-only SSH, no direct root login (sshd_config drop-in) | RDP (Remote Desktop Protocol, Windows' remote login) only from a jump host or gateway (a hardened server that administrators connect through) | most intrusions arrive over admin protocols |
| Local admin credentials | no shared local passwords | Windows LAPS (Windows Local Administrator Password Solution) rotates a unique local administrator password per machine and escrows it (stores a copy for authorised administrators to retrieve) in Active Directory or Microsoft Entra ID | one stolen local password otherwise opens every server |
| Patch posture | automatic security updates or a scheduled cadence with a reporting hook | WSUS (Windows Server Update Services, Microsoft's on-premises update server; Microsoft lists it as no longer actively developed in Windows Server 2025, with existing capabilities and content still available) or whichever update tooling the estate already runs, with a ring schedule | known-vulnerability exposure shrinks fastest here |
| Host firewall | default-deny inbound, named exceptions | host firewall enabled with inbound blocked by default | limits lateral movement (an attacker who has compromised one machine hopping to others) |
| Legacy and unused services | remove or disable what the server role does not need | disable obsolete protocols | smaller attack surface, rarely breaks anything |
| Logging and time | auditd (the Linux audit daemon) or journal forwarding, time sync | advanced audit policy, event forwarding | you need evidence for the "prove it" step and for incidents |
The Windows LAPS row of tier 1 has a concrete shape. Windows LAPS is configured through a Group Policy setting under Computer Configuration, Policies, Administrative Templates, System, LAPS. Illustrative policy for the fleet, using the setting names and defaults from Microsoft's Windows LAPS documentation: backup directory Active Directory, password age 30 days, password length 14, complexity 4 (upper case, lower case, numbers and special characters). An administrator with the right to decrypt can then read a server's current password:
Get-LapsADPassword -Identity SRV-APP-01 -AsPlainText
Output shaped like Microsoft's documentation example (host name illustrative, password omitted):
ComputerName : SRV-APP-01
Account : Administrator
Password : <the current 14-character password>
PasswordUpdateTime : 4/9/2023 9:39:38 AM
ExpirationTimestamp : 4/14/2023 9:39:38 AM
Source : EncryptedPassword
DecryptionStatus : Success
Read it as: the managed account is Administrator, the password was last rotated at PasswordUpdateTime and will be rotated again after ExpirationTimestamp, it is stored encrypted in Active Directory (Source), and the caller was allowed to decrypt it. A different password per server is the point: a stolen password for one machine opens nothing else.
Everything else in Level 1 follows in later tiers, and Level 2 items are opt-in per workload.
How to enforce: one definition, two delivery paths
The plan rests on three pieces: the published baseline as the definition, one enforcement tool per OS family (Group Policy for domain-joined Windows, Ansible for Linux), and one scanner per OS family (OpenSCAP for Linux, Policy Analyzer and CIS-CAT for Windows). Packer, LGPO.exe, Puppet, Chef and PowerShell DSC appear in the tables as the alternatives that fit particular situations (image builds, machines outside the domain, estates that already run an agent).
| Decision | Options | Recommendation |
|---|---|---|
| Image baking vs post-provision | Bake the baseline into the machine image (built with a tool such as Packer, which automates image creation, and scanned at build time) vs apply it after the server boots | Both, from the same definition. Bake for new cloud and template-built servers so they are born compliant; apply post-provision for the existing 1,000, since a new image does nothing for a server that already exists. A baked image goes stale until rebuilt, so the scheduled run still applies. |
| Agent vs agentless | Push over SSH or WinRM (Windows Remote Management) from a controller (Ansible) vs a resident agent that pulls policy (the Group Policy client, or configuration managers such as Puppet, Chef and PowerShell DSC, Desired State Configuration) | Use what already exists first. GPO for domain-joined Windows (the client is built in) and LGPO.exe for non-domain machines; Ansible, run on a schedule from a pipeline, for Linux. Agentless needs a network path and a credential vault and misses hosts that are offline during the run; an agent self-heals on its own interval and works behind NAT (network address translation, where a host sits behind a shared address and cannot be reached from outside, so it must call out to the controller instead), but adds a new always-running component to 1,000 servers. |
| Audit vs enforce | Report only vs change the setting | Audit mode in Phase 0 and for every new setting for one ring; enforce after the owner has seen the report. |
Roll out in rings, so a bad setting hurts few servers: ring 0 is 20 servers (2%), ring 1 is 100 (10%), ring 2 is 300 (30%), ring 3 is the remaining 580 (58%): 20 + 100 + 300 + 580 = 1,000. Each ring waits for a clean soak (no baseline-caused incidents) before the next starts. Pick ring 0 deliberately: non-production, then low-risk production, with owners who agreed in advance.
How to prove it, and keep proving it
- Scans on a schedule, results to one store: OpenSCAP or CIS-CAT results for Linux, Policy Analyzer comparisons and CIS-CAT for Windows. Keep the raw results file as evidence, because "dashboard says green" is not evidence.
- Metrics that cannot be gamed: percentage of servers passing tier 1 (applicable rules only), count of open exceptions and how many are expired, and median days from "new failing rule detected" to "fixed or excepted".
- Continuous assessment trade-offs: a daily scan catches drift quickly but costs CPU on the servers and generates volume; weekly scans are cheap and slower to notice. Stagger scan start times so 1,000 servers are not all scanning at once, run daily for tier 1 only and weekly for the full profile.
- Remediation automation: OpenSCAP can emit remediation content from a result (
oscap xccdf generate fix, with fix types including bash and ansible) and can apply it with--remediate. Treat generated fixes as a starting point to review and put in version control. Do not run--remediateacross production unreviewed. - Independent check: have someone outside the team sample 25 servers a quarter and compare the scan result with what is actually configured.
How to decide what to leave out
Leave a setting out (or defer it) only when one of these is true, and write it down:
| Reason | Example | Record |
|---|---|---|
| The server role needs the opposite | a legacy application needs a protocol the baseline disables | the specific hosts, the dependency, the plan to retire it |
| The control does not apply | a rule about a wireless interface on a server with none | marked not-applicable, no exception needed |
| A compensating control already covers the risk | a rule that duplicates a network-level restriction | name the control and how it is verified |
| Cost clearly exceeds benefit now | a Level 2 rule on a fleet where it needs a reboot per server | revisit date |
Each exception carries an owner, the exact hosts, the reason, the compensating control and an expiry date (I would default to six months) so it forces a decision again. An exception without an expiry is a permanent hole.
Patch windows and emergency changes
Align the monthly patch window with the rings, so a baseline change and a patch do not land on the same host in the same night, which makes failures impossible to attribute. For an actively exploited vulnerability, keep a pre-approved emergency path outside the monthly cycle; an out-of-band change that touches a hardened setting creates a time-limited exception (72 hours, for example) rather than a silent edit, and the scheduled enforcement run is paused for that host until it expires.
The first 90 days
| Days | Deliverable | Exit test |
|---|---|---|
| 1 to 30 | reconciled inventory, owners, baseline chosen per OS, assessment-only scans on every server, tier 1 list agreed | every server has an owner and a measured pass rate |
| 31 to 60 | golden images rebuilt with tier 1; enforcement live on rings 0 and 1 (120 servers); exception process running | zero baseline-caused outages in rings 0 and 1 |
| 61 to 90 | rings 2 and 3 (the other 880); scheduled scans and dashboard; drift alerts | proposed target: at least 95% of applicable tier 1 rules passing fleet-wide, every failure either ticketed or excepted |
The 95% figure is a target to agree, not a measurement.
Boundary
Container images and the cluster that schedules them are hardened separately. This programme covers the host operating systems that run them, since a container runtime inherits the host's weaknesses.
Pitfalls
- Enforcing everything in Level 1 on day one because "it is only Level 1".
- Counting not-applicable rules as passes, or excluding failing servers from the denominator.
- Two baselines applied to one host (a CIS GPO plus a Microsoft one) that silently overwrite each other.
- Exceptions with no expiry, which turn the baseline into a suggestion.
Unlock Full Question Bank
Get access to all 47 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.