System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
Your organisation runs about 1,000 mixed Windows and Linux servers, on-prem and in cloud, with no consistent security baseline. Design the programme that gets them to one and keeps them there: what you enforce first, how you enforce it, how you prove it, and how you decide what to leave out.
Sample Answer
Direct answer
I would run this as a measured programme, not a one-off hardening project. Pick one published baseline per operating system (CIS Benchmarks, the Center for Internet Security's consensus configuration guides, at Level 1 as the starting profile), measure first without changing anything, enforce a short first tier of low-breakage, high-impact settings through the same tooling that builds and manages the servers, prove compliance with automated scans whose results feed a dashboard, and handle every deliberate omission through an exception record with an owner, a compensating control and an expiry date. The baseline is a living definition in version control, so "keeping them there" is a scheduled enforcement and scan loop, not a recurring audit.
For illustration assume 600 Linux and 400 Windows servers (1,000 in total) in a mix of on-premises and cloud. The split is an assumption; the structure does not depend on it.
What the baseline is, and why Level 1 first
CIS defines a Level 1 profile as a base recommendation that can be implemented fairly promptly without extensive performance impact, and Level 2 as defense in depth for environments where security is paramount, which can have an adverse effect if implemented without due care. CIS also advises testing either level in a test environment first. For Windows Server, Microsoft publishes security baselines as Group Policy Object (GPO) backups in its Security Compliance Toolkit, together with Policy Analyzer (compares sets of GPOs against each other or against local policy) and LGPO.exe (applies local policy to machines that are not domain-joined). So the plan is: CIS Level 1 server profile per Linux distribution, and either the Microsoft baseline or the CIS Windows Server benchmark per Windows version, choosing one per OS and not mixing, because overlapping baselines conflict.
Phase 0: discovery and assessment (measure before enforcing)
You cannot enforce what you have not inventoried. Build the server list from more than one source (the configuration management database, the hypervisor and cloud account inventories, the directory) and reconcile it, because the servers nobody remembers are the ones with no baseline. Record for each server: OS and version, owner, environment, internet exposure, whether it can be rebooted, and whether a vendor supports changes to it.
Then run the benchmark in assessment-only mode everywhere. For Linux, OpenSCAP (oscap, an open-source implementation of SCAP, the Security Content Automation Protocol) with the SCAP Security Guide content (the rule definitions) evaluates a profile and writes results without changing the host. A real run in a test AlmaLinux 9 container against the CIS server Level 1 profile:
dnf -y install openscap-scanner scap-security-guide
oscap xccdf eval \
--profile xccdf_org.ssgproject.content_profile_cis_server_l1 \
--results /tmp/results.xml --report /tmp/report.html \
/usr/share/xml/scap/ssg/content/ssg-almalinux9-ds.xml
echo "exit status: $?"
The run printed one block per rule: a Title, the Rule id and a Result. Three of them, copied from the run:
Title Ensure AlmaLinux GPG Key Installed
Rule xccdf_org.ssgproject.content_rule_ensure_almalinux_gpgkey_installed
Result pass
Title Implement Custom Crypto Policy Modules for CIS Benchmark
Rule xccdf_org.ssgproject.content_rule_configure_custom_crypto_policy_cis
Result fail
Title Install AIDE
Rule xccdf_org.ssgproject.content_rule_package_aide_installed
Result notapplicable
pass means the host matches the rule, fail means it does not and needs a fix or an exception, and notapplicable means the rule does not apply to this host (the scanner decides this from conditions in the content). The --profile value is the XCCDF (the XML format that SCAP uses for benchmarks) profile id, and here it selects the CIS server Level 1 rule set out of the content file. The tally was 70 pass, 5 fail, 220 notapplicable, 0 error, with exit status 2: according to the oscap manual, 0 means every rule passed, 1 means an error occurred during evaluation, and 2 means at least one rule failed or was unknown, so a pipeline can gate on the exit status. The honest compliance figure excludes the not-applicable rules from the denominator: 70 / (70 + 5) = 93.3% of the 75 applicable rules, not 70 of 295. Report both numbers, and say which one is the headline, or the dashboard will quietly flatter itself. (CIS-CAT Pro Assessor is CIS's own tool for assessing conformance, and is the alternative where a vendor-certified assessment is wanted.)
The assessment also tells you how bad it is by item: the fleet-wide failure count per rule is your prioritisation list.
What to enforce first (tier 1)
Rank candidate settings by three questions: does it close a path attackers actually use, how likely is it to break something, and can I verify it automatically? The first tier takes settings that score well on all three:
| Area | Linux | Windows | Why first |
|---|---|---|---|
| Remote administration | key-only SSH, no direct root login (sshd_config drop-in) | RDP (Remote Desktop Protocol, Windows' remote login) only from a jump host or gateway (a hardened server that administrators connect through) | most intrusions arrive over admin protocols |
| Local admin credentials | no shared local passwords | Windows LAPS (Windows Local Administrator Password Solution) rotates a unique local administrator password per machine and escrows it (stores a copy for authorised administrators to retrieve) in Active Directory or Microsoft Entra ID | one stolen local password otherwise opens every server |
| Patch posture | automatic security updates or a scheduled cadence with a reporting hook | WSUS (Windows Server Update Services, Microsoft's on-premises update server; Microsoft lists it as no longer actively developed in Windows Server 2025, with existing capabilities and content still available) or whichever update tooling the estate already runs, with a ring schedule | known-vulnerability exposure shrinks fastest here |
| Host firewall | default-deny inbound, named exceptions | host firewall enabled with inbound blocked by default | limits lateral movement (an attacker who has compromised one machine hopping to others) |
| Legacy and unused services | remove or disable what the server role does not need | disable obsolete protocols | smaller attack surface, rarely breaks anything |
| Logging and time | auditd (the Linux audit daemon) or journal forwarding, time sync | advanced audit policy, event forwarding | you need evidence for the "prove it" step and for incidents |
The Windows LAPS row of tier 1 has a concrete shape. Windows LAPS is configured through a Group Policy setting under Computer Configuration, Policies, Administrative Templates, System, LAPS. Illustrative policy for the fleet, using the setting names and defaults from Microsoft's Windows LAPS documentation: backup directory Active Directory, password age 30 days, password length 14, complexity 4 (upper case, lower case, numbers and special characters). An administrator with the right to decrypt can then read a server's current password:
Get-LapsADPassword -Identity SRV-APP-01 -AsPlainText
Output shaped like Microsoft's documentation example (host name illustrative, password omitted):
ComputerName : SRV-APP-01
Account : Administrator
Password : <the current 14-character password>
PasswordUpdateTime : 4/9/2023 9:39:38 AM
ExpirationTimestamp : 4/14/2023 9:39:38 AM
Source : EncryptedPassword
DecryptionStatus : Success
Read it as: the managed account is Administrator, the password was last rotated at PasswordUpdateTime and will be rotated again after ExpirationTimestamp, it is stored encrypted in Active Directory (Source), and the caller was allowed to decrypt it. A different password per server is the point: a stolen password for one machine opens nothing else.
Everything else in Level 1 follows in later tiers, and Level 2 items are opt-in per workload.
How to enforce: one definition, two delivery paths
The plan rests on three pieces: the published baseline as the definition, one enforcement tool per OS family (Group Policy for domain-joined Windows, Ansible for Linux), and one scanner per OS family (OpenSCAP for Linux, Policy Analyzer and CIS-CAT for Windows). Packer, LGPO.exe, Puppet, Chef and PowerShell DSC appear in the tables as the alternatives that fit particular situations (image builds, machines outside the domain, estates that already run an agent).
| Decision | Options | Recommendation |
|---|---|---|
| Image baking vs post-provision | Bake the baseline into the machine image (built with a tool such as Packer, which automates image creation, and scanned at build time) vs apply it after the server boots | Both, from the same definition. Bake for new cloud and template-built servers so they are born compliant; apply post-provision for the existing 1,000, since a new image does nothing for a server that already exists. A baked image goes stale until rebuilt, so the scheduled run still applies. |
| Agent vs agentless | Push over SSH or WinRM (Windows Remote Management) from a controller (Ansible) vs a resident agent that pulls policy (the Group Policy client, or configuration managers such as Puppet, Chef and PowerShell DSC, Desired State Configuration) | Use what already exists first. GPO for domain-joined Windows (the client is built in) and LGPO.exe for non-domain machines; Ansible, run on a schedule from a pipeline, for Linux. Agentless needs a network path and a credential vault and misses hosts that are offline during the run; an agent self-heals on its own interval and works behind NAT (network address translation, where a host sits behind a shared address and cannot be reached from outside, so it must call out to the controller instead), but adds a new always-running component to 1,000 servers. |
| Audit vs enforce | Report only vs change the setting | Audit mode in Phase 0 and for every new setting for one ring; enforce after the owner has seen the report. |
Roll out in rings, so a bad setting hurts few servers: ring 0 is 20 servers (2%), ring 1 is 100 (10%), ring 2 is 300 (30%), ring 3 is the remaining 580 (58%): 20 + 100 + 300 + 580 = 1,000. Each ring waits for a clean soak (no baseline-caused incidents) before the next starts. Pick ring 0 deliberately: non-production, then low-risk production, with owners who agreed in advance.
How to prove it, and keep proving it
- Scans on a schedule, results to one store: OpenSCAP or CIS-CAT results for Linux, Policy Analyzer comparisons and CIS-CAT for Windows. Keep the raw results file as evidence, because "dashboard says green" is not evidence.
- Metrics that cannot be gamed: percentage of servers passing tier 1 (applicable rules only), count of open exceptions and how many are expired, and median days from "new failing rule detected" to "fixed or excepted".
- Continuous assessment trade-offs: a daily scan catches drift quickly but costs CPU on the servers and generates volume; weekly scans are cheap and slower to notice. Stagger scan start times so 1,000 servers are not all scanning at once, run daily for tier 1 only and weekly for the full profile.
- Remediation automation: OpenSCAP can emit remediation content from a result (
oscap xccdf generate fix, with fix types including bash and ansible) and can apply it with--remediate. Treat generated fixes as a starting point to review and put in version control. Do not run--remediateacross production unreviewed. - Independent check: have someone outside the team sample 25 servers a quarter and compare the scan result with what is actually configured.
How to decide what to leave out
Leave a setting out (or defer it) only when one of these is true, and write it down:
| Reason | Example | Record |
|---|---|---|
| The server role needs the opposite | a legacy application needs a protocol the baseline disables | the specific hosts, the dependency, the plan to retire it |
| The control does not apply | a rule about a wireless interface on a server with none | marked not-applicable, no exception needed |
| A compensating control already covers the risk | a rule that duplicates a network-level restriction | name the control and how it is verified |
| Cost clearly exceeds benefit now | a Level 2 rule on a fleet where it needs a reboot per server | revisit date |
Each exception carries an owner, the exact hosts, the reason, the compensating control and an expiry date (I would default to six months) so it forces a decision again. An exception without an expiry is a permanent hole.
Patch windows and emergency changes
Align the monthly patch window with the rings, so a baseline change and a patch do not land on the same host in the same night, which makes failures impossible to attribute. For an actively exploited vulnerability, keep a pre-approved emergency path outside the monthly cycle; an out-of-band change that touches a hardened setting creates a time-limited exception (72 hours, for example) rather than a silent edit, and the scheduled enforcement run is paused for that host until it expires.
The first 90 days
| Days | Deliverable | Exit test |
|---|---|---|
| 1 to 30 | reconciled inventory, owners, baseline chosen per OS, assessment-only scans on every server, tier 1 list agreed | every server has an owner and a measured pass rate |
| 31 to 60 | golden images rebuilt with tier 1; enforcement live on rings 0 and 1 (120 servers); exception process running | zero baseline-caused outages in rings 0 and 1 |
| 61 to 90 | rings 2 and 3 (the other 880); scheduled scans and dashboard; drift alerts | proposed target: at least 95% of applicable tier 1 rules passing fleet-wide, every failure either ticketed or excepted |
The 95% figure is a target to agree, not a measurement.
Boundary
Container images and the cluster that schedules them are hardened separately. This programme covers the host operating systems that run them, since a container runtime inherits the host's weaknesses.
Pitfalls
- Enforcing everything in Level 1 on day one because "it is only Level 1".
- Counting not-applicable rules as passes, or excluding failing servers from the denominator.
- Two baselines applied to one host (a CIS GPO plus a Microsoft one) that silently overwrite each other.
- Exceptions with no expiry, which turn the baseline into a suggestion.
Design a patch baseline and schedule for a hybrid estate of Windows and Linux servers managed through a cloud update service. Cover update classifications, pre and post scripts, maintenance windows, and machines that are offline when the window opens.
Sample Answer
Direct answer
Use Azure Update Manager (the Azure service that assesses and installs updates on Windows and Linux machines) with one maintenance configuration per patch ring, attach machines to rings by tag through dynamic scoping, and install only the Critical and Security classifications on a monthly cadence for Windows (anchored to "Patch Tuesday", the second Tuesday of the month) and a weekly security-only cadence for Linux. Use "Reboot if required" for stateless tiers and a controlled reboot for databases. Use pre and post events (small automations triggered before and after each run) to power on machines that are normally off, to snapshot or stop services beforehand, and to run health checks afterwards. For machines that are offline when the window opens, rely on three things: a power-on pre-event for machines you control, a catch-up schedule a few days later, and the 24-hour periodic assessment to show you who was missed.
Azure terms in plain words. A maintenance configuration is a saved schedule object (when, for how long, which update classifications, which reboot option) that you attach to machines. Dynamic scoping is a saved filter (subscription, resource group, location, operating system, tags) that decides which machines a maintenance configuration covers, re-evaluated each time it runs. Customer Managed Schedules is a patch-orchestration setting meaning only your schedule patches the machine. An availability set is an Azure grouping that spreads virtual machines across separate hardware, and an update domain is a slice of it that the platform restarts at one time. Azure Resource Graph is Azure's query service over resource metadata and results.
The estate is hybrid: Azure virtual machines are patched directly, and on-premises servers are patched after they are connected through Azure Arc (the service that represents non-Azure servers as Azure resources). Update Manager honors the update source already configured on each machine, so Windows machines keep using Windows Update, Microsoft Update or WSUS (Windows Server Update Services, where you approve updates yourself; Microsoft documents WSUS as deprecated, meaning no new features but still supported with security and quality updates), and Linux machines keep using their configured repositories. It installs what that source offers; it does not publish updates.
1. Baseline: what gets installed
| Decision | Choice | Reason |
|---|---|---|
| Classifications | Critical and Security on both operating systems | Smallest set that closes known exploited flaws; broader categories (feature updates and the rest) go through a separate, slower review |
| Exclusions | Exclude packages or KB (Knowledge Base article) IDs you have found to break an application; Linux accepts package names with wildcards such as kernel* | Lets the database tier hold kernel updates until a controlled window |
| Drivers | Not handled | Update Manager does not support driver updates, so keep a separate process for firmware and drivers |
| Update source | WSUS approvals for Windows if you need to hold a patch; repository snapshots for Linux if you need reproducibility | The service installs what the source offers, so the source is where you control "which patch" |
| Patch orchestration on Azure VMs | Set to Customer Managed Schedules (required for scheduled patching on Azure VMs; not required for Azure Arc-enabled machines) | Otherwise the platform or the guest operating system patches on its own timetable |
What one ring looks like as a configuration (illustrative values)
| Field | Pilot ring, Windows |
|---|---|
| Name | patch-ring-0-windows |
| Scope (dynamic scoping filter) | tag patch-ring = 0, operating system = Windows |
| Schedule | monthly, second Tuesday, offset +1 day, 22:00 |
| Maintenance window length | 3 hours |
| Classifications | Critical, Security |
| Exclusions | KB IDs known to break the application |
| Reboot | Reboot if required |
A pre-event handler is a small piece of code (for example an Azure Function, code Azure runs on demand) subscribed through Event Grid to the pre-maintenance event. In this design it does three things: start machines tagged as normally off, take the data-tier snapshot, and call the cancellation API if either step fails.
Which numbers to design around: the 3-hour window with its 25-minute cut-off (decides when the last install can start), the "machine on at least 15 minutes before the start" rule together with the pre-event period (decide when power-on must begin), and the 24-hour assessment (drives the catch-up schedule). The 40-minute edit rule, the 20-machine one-time update limit and the 7 and 30 day retention periods are operating details to look up when you meet them.
2. Rings, schedules and maintenance windows
A maintenance configuration starts updates on all of its attached machines at the same time, so rings must be separate configurations. Dynamic scoping (selecting machines by subscription, resource group, location, operating system and tags) or the built-in Azure Policy named "Schedule recurring updates using Azure Update Manager" attaches new machines automatically, so a server built next month is patched without anyone remembering it.
| Ring | Tag | Windows schedule | Linux schedule | Reboot |
|---|---|---|---|---|
| Pilot | patch-ring=0 | Monthly, second Tuesday with +1 day offset (Wednesday), 22:00 | Weekly, Wednesday 22:00 | If required |
| Production A | patch-ring=1 | Monthly, second Tuesday with +4 day offset (Saturday), 22:00 | Weekly, Saturday 22:00 | If required |
| Production B (databases, clusters) | patch-ring=2 | Monthly, second Tuesday with +6 day offset (Monday), 02:00 | Weekly, Sunday 02:00 | Never, followed by a planned reboot |
The offsets use the product's own recurrence rule ("nth weekday of the month, with an offset of up to six days either way"); Tuesday plus 1 is Wednesday, plus 4 is Saturday and plus 6 is Monday.
Window length: the portal accepts 1 hour 30 minutes to 3 hours 55 minutes; the schedule needs to repeat no more often than every 6 hours. Choose 3 hours (180 minutes) for Windows. Update Manager checks the remaining time before each step: per the troubleshooting guidance, with fewer than 25 minutes left (15 for the update plus 10 reserved for the reboot) it does not scan or install, and a service pack needs 30. So 180 - 25 = 155 minutes is the latest point at which a new ordinary update can start. Updates already running are never killed, and anything not attempted is reported as "Not attempted", so a window that is too short shows up as a recurring pattern in the history instead of a silent failure. On Linux a reboot needs 15 minutes left in the window on Azure VMs.
Reboot options are Reboot if required, Never reboot and Always reboot. Two details matter. "Never reboot" leaves a machine whose updates need a restart waiting for one that nothing in the run will perform, so the planned reboot for Production B must be a deliberate step (a post-event or an operations task) and not an assumption. And some Windows registry settings (the Windows Update and restart policies) can cause a reboot even when you chose Never reboot, so set those to match.
Machines in a common availability set are not all updated at once: they are updated within update domain boundaries, and machines across multiple update domains are not updated concurrently. If you patch members of one availability set under different schedules and one window overruns, a member can fail or be skipped. Split such machines across schedules at different times and widen the window.
3. Pre and post events
Update Manager uses Azure Event Grid (a service that routes events to handlers such as a webhook or an Azure Function) for this. The documented order is: pre-event, optional cancellation, update installation, post-event.
| Phase | Typical actions here |
|---|---|
| Pre-event | Power on machines that are normally off; take a disk snapshot of the data tier; send a notification; stop an application service that must not be interrupted |
| Cancellation | Your pre-event logic must call the cancellation API itself when a pre-step fails. The service does not cancel automatically, and if you do nothing the run proceeds |
| Post-event | Start services again; run a smoke test; power machines back off; send the patch summary |
Timing facts to design around (from the Microsoft documentation, using its own example of a 3:00 PM start): pre-events run outside the maintenance window, between about 2:30 PM and 2:50 PM, and the cancellation call must be made by 2:50 PM; post-events start as soon as installation ends and can run outside the window; and the status of the run covers update installation only, not the pre and post events, so monitor the event handler separately. If you create or edit a schedule that has a pre-event, do it at least 40 minutes before the start or that run is cancelled automatically.
4. Machines that are offline when the window opens
- Shut-down machines cannot be patched. The machine must be on at least 15 minutes before the scheduled start. Machines may also appear disassociated from their configuration while off; that is a display issue, and the association is still there.
- For machines you control that are off by design (test and development virtual machines), power them on from a pre-event, and start the power-on at the beginning of the pre-event period. For a 3:00 PM start the machine must be running by 2:45 PM (3:00 minus 15 minutes), while the documented pre-event period runs from about 2:30 to 2:50, so a power-on that begins at the end of the period is already too late. Power them off from the post-event.
- For servers that are simply unreachable (a branch server with a network outage), the run misses them. The periodic assessment (every 24 hours when enabled) keeps reporting what is missing, so build two things on it: a dashboard or report of machines with pending Critical and Security updates, and a catch-up maintenance configuration a few days after the main window, scoped to the same tag. The catch-up is safe to run broadly, because every run starts with a fresh assessment, so machines already patched have nothing left to install. For individual stragglers, a one-time update can be started from the portal (up to 20 machines at a time).
- Keep evidence outside the service. Pending-update data is retained for 7 days and installation results for 30 days in Azure Resource Graph (the tables
patchassessmentresourcesandpatchinstallationresources), so export compliance results monthly if audit evidence is needed. - Agent health is the other cause of silent misses: Update Manager installs its own extension on each machine the first time it runs, and Azure Arc machines must be connected. Alert on machines that have not been assessed.
Mapping the design to other tools
The same design uses vendor-neutral parts: a baseline (which classifications and exclusions are approved), rings selected by tag, a schedule per ring, hooks before and after, and an assessment that reports who was missed.
| Part | Azure Update Manager | AWS Systems Manager Patch Manager | Ansible-driven estate |
|---|---|---|---|
| Baseline | classifications and exclusions in the maintenance configuration | a patch baseline (approval rules by classification and severity, plus approved and rejected patch lists) | a package list or security-only update task per play |
| Ring membership | dynamic scoping by tag | targets chosen by tag | inventory groups per ring |
| Schedule | maintenance configuration recurrence | maintenance window or patch policy | the scheduler that runs the play |
| Before and after steps | pre and post events through Event Grid | lifecycle hooks (Systems Manager documents run before and after patching) | tasks before and after the update task |
| Who was missed | periodic assessment every 24 hours | a Scan operation and compliance reports | a read-only check play on a schedule |
Trade-offs and what would change the choices
- Monthly Windows plus weekly Linux security splits the estate by how fast each ecosystem ships fixes. If the audit requires one cadence for both, use weekly Critical and Security for both and accept more reboots.
- Pre-events add moving parts that can fail. Keep them to jobs that really need to happen before patching, and make the failure path explicit (cancel the run, notify).
- If a patch must be validated before it reaches production, hold it at the source (WSUS approval, repository snapshot) and not by delaying a schedule, so the pilot ring still installs what production will later install.
That is every published System and Endpoint Hardening question for Cloud Architect so far. Browse the other topics in this category, or practice this one interactively.