System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
What are SELinux and AppArmor, and how does mandatory access control differ from the ordinary file permissions most people rely on? When would you insist on it and when would you hesitate?
Sample Answer
Direct answer
SELinux (Security-Enhanced Linux) and AppArmor are two implementations of mandatory access control (MAC) for the Linux kernel. Ordinary file permissions are discretionary access control (DAC): the owner of a file decides who may read or write it, and any program running as that user inherits all of that user's rights. MAC adds a second, system-wide policy written by the administrator or the distribution, which the kernel enforces on top of the permission bits and which a user or a compromised program cannot change at will. That lets you confine a program to what it legitimately needs, so a hijacked web server running as an ordinary account cannot, for example, read every file that account happens to own. I would insist on it for anything internet-facing or handling sensitive data, and hesitate only about how fast to enforce it on software nobody understands, never about whether to use it.
DAC versus MAC, concretely
| DAC (owner / group / other bits, ACLs) | MAC (SELinux or AppArmor) | |
|---|---|---|
| Who sets the rule | the file's owner (or root) | the system policy, loaded into the kernel |
| What the rule is about | which users may touch a file | which programs may touch which resources |
| Can a process widen its own access | yes, if it owns the object | no, the policy is outside its control |
| Order of checks | first | after DAC: the Red Hat documentation states that SELinux rules "are checked after DAC rules", so if DAC denies, no SELinux denial is even logged |
Both checks must pass. A MAC policy in effect narrows what DAC allows, which is why a permissions fix (chmod) alone does not cure a MAC denial.
SELinux and AppArmor: how they differ
| SELinux | AppArmor | |
|---|---|---|
| Identifies things by | labels: every process and file carries a context of user, role, type and level (user:role:type:level); the type (ending in _t, such as httpd_t for Apache) is what policy rules mostly use; on a process the type is also called its domain | paths: a profile per program lists the files and capabilities (Linux splits root's powers into separate privileges, such as binding a port below 1024 or changing file ownership) it may use |
| Default on | Red Hat family (Red Hat's documentation describes enforcing as the default and recommended mode) | Ubuntu (pre-installed and active) and, from Debian 10, Debian (enabled by default) |
| Modes | enforcing, permissive (logs but does not block), disabled | enforce, complain (logs but allows) |
| Strength | very fine-grained: covers processes, files, network ports and more | profiles are plain-text files under /etc/apparmor.d/, named after the program's path |
| Cost | steep learning curve; a wrong label on a file is a common cause of denials | a profile must name every path the program uses, and anything not listed is denied, so a program whose files move needs its profile updated |
An actual AppArmor profile fragment shipped on Ubuntu 24.04 (/etc/apparmor.d/lsb_release, trimmed) shows the idea: the program may read a few files, execute a few helpers, and nothing else:
profile lsb_release {
include <abstractions/base>
/dev/tty rw,
/usr/bin/lsb_release r,
/etc/lsb-release r,
/{usr/,}bin/bash ixr,
/usr/bin/cat ixr,
}
r is read, w is write, x execute (ix means the helper runs under this same profile), and anything not granted is denied. Letters combine, so /usr/bin/cat ixr, grants two things at once: execute cat under this same profile (ix) and read the file (r), which the kernel needs in order to load the program. A SELinux rule says the same thing in terms of types: "processes in domain httpd_t may read files of type httpd_sys_content_t".
Working example: a web server returns 403 on a correct-looking directory
Apache on a RHEL-family host serves /srv/myweb, owned by the right user with mode 755, yet returns 403. DAC is fine; the files carry the wrong type for a web server to read. The sequence:
-
getenforceto see the mode, andls -Zd /srv/mywebto see the directory's context (ls -Zprints the security context;-dlists the directory itself). Under/srvthe default label in the RHEL policy isvar_t(read from the policy's file-context list), a type Apache's domain is not allowed to read. Illustrative output before and after the fix in step 3 (formatuser:role:type:level; the type is the third field):textbefore: unconfined_u:object_r:var_t:s0 /srv/myweb after: unconfined_u:object_r:httpd_sys_content_t:s0 /srv/myweb -
ausearch -m AVC,USER_AVC,SELINUX_ERR,USER_SELINUX_ERR -ts recentsearches the audit log for denial records (an AVC, access vector cache, message is how SELinux logs a denial; the name comes from the kernel cache that holds access decisions), orsealert -l "*"for a readable explanation (-mselects the message types to list,-ts recentlimits the search to the last ten minutes, and-l "*"askssealertto describe every alert it has). A denial record for this story looks like this (illustrative, in the format SELinux writes):texttype=AVC msg=audit(1728200000.123:456): avc: denied { read } for pid=1840 comm="httpd" name="index.html" dev="dm-0" ino=412 scontext=system_u:system_r:httpd_t:s0 tcontext=unconfined_u:object_r:var_t:s0 tclass=file permissive=0Read it as:
{ read }is the action refused;comm="httpd"andscontextsay who asked (domainhttpd_t);nameandtcontextsay what was touched (a file of typevar_t);tclass=fileis the object kind;permissive=0means the host was enforcing, so the request really was blocked. -
Fix the label first:
semanage fcontext -a -t httpd_sys_content_t "/srv/myweb(/.*)?"records the rule permanently, thenrestorecon -R -v /srv/mywebapplies it. Thesemanage fcontextrule makes the labelling part of the policy, andrestoreconapplies it to the existing files. The pattern at the end of the path is a regular expression:(/.*)?means an optional slash followed by anything, so the rule covers/srv/mywebitself and every file and folder below it (checked withgrep -E:/srv/myweband/srv/myweb/a/b.htmlmatch^/srv/myweb(/.*)?$,/srv/mywebsitedoes not). -
If labelling is correct and the denial is a legitimate feature (the app must connect to a database), use a boolean, a switch built into the policy:
setsebool -P httpd_can_network_connect_db on. A non-standard port is likewise a label:semanage port -a -t http_port_t -p tcp 9876. -
Only if neither fits, write a small custom module. Red Hat explicitly warns against using
audit2allowto generate a local module as the first response to a denial, because it turns whatever was blocked, including a real attack, into permitted behaviour.
The temptation is setenforce 0, which switches the whole host to permissive. The better tool is to make one domain permissive while the rest stays enforcing: semanage permissive -a httpd_t, then remove it when done. Never set SELinux to disabled: with it disabled objects are not even labelled, which makes turning it back on painful.
The AppArmor equivalent: aa-status lists profiles and modes; aa-complain /path/to/bin moves one profile to complain mode while you collect what the program really does; edit /etc/apparmor.d/<profile>; apparmor_parser -r /etc/apparmor.d/<profile> reloads it; aa-enforce /path/to/bin returns it to enforcement.
When I would insist, and when I would hesitate
Insist when:
- the host is internet-facing (web servers, mail, VPN endpoints) or processes untrusted input, since confinement limits what a successful exploit can reach;
- the host is multi-tenant or runs code from several teams;
- a compliance baseline requires it (the SCAP Security Guide's CIS Server Level 1 profile for AlmaLinux 9, for example, includes the rules "Ensure SELinux is Not Disabled" and "Configure SELinux Policy", so a scan will flag a host that has it off);
- the host runs containers, where MAC is an additional layer between a container and the host (installing
apparmorandapparmor-utilson Ubuntu 24.04 places profiles forrunc,crun,podmanandbuildahin/etc/apparmor.d).
Hesitate (that is, stage it, do not skip it) when:
- a vendor application with unknown behaviour has files in unusual places and nobody can say what it legitimately touches: run it in permissive or complain mode with logging for a defined period, review the denials, then enforce with a date attached;
- the team has no one who can read a denial, in which case train someone before enforcing on a revenue system, not after the first outage;
- the application is in the middle of a migration and its file layout is still changing.
The pattern "hesitate" never means "disable". A permissive domain with logging still gives you the audit trail and a path to enforcement, while a disabled module gives you neither.
Pitfalls
- Fixing a denial by loosening DAC (
chmod 777), which weakens the layer that was working and does not address the MAC denial. - Generating a policy module from every denial without reading them.
- Restoring or copying data into place and assuming the labels are right: run
restoreconon it and check withls -Z. - Assuming AppArmor and SELinux are interchangeable in an operating procedure: the commands, defaults and failure modes differ, so write the runbook per distribution.
What does least privilege mean on a host, and how would you actually enforce it for users, files and directories on both Windows and Linux?
Sample Answer
Direct answer
Least privilege means every user, program and service gets only the permissions needed for its job, and only for as long as it needs them. On a host that means three layers: people (no everyday admin rights, separate admin identities, named sudo rules on Linux, managed local admin passwords on Windows), files and directories (owner and group set to the real service account, no access for "everyone", broken inheritance only where needed), and services (each running under its own low-privilege account). I enforce it with groups rather than individual grants, write the permissions as code, and then audit them with commands so I find drift rather than assume it.
Users: who can do what
| Layer | Linux | Windows |
|---|---|---|
| Daily account | Normal user, no root | Normal user, not a member of local Administrators |
| Elevation | sudo rules naming specific commands, per group | UAC (User Account Control) prompts for admin tasks, plus a separate admin account for administrators |
| Shared local admin credential | Do not have one; root login over SSH is disabled | Windows LAPS (Local Administrator Password Solution) gives every machine its own rotating local admin password, stored in Active Directory or Microsoft Entra ID (Microsoft's cloud identity directory, formerly called Azure AD) and readable only by authorised staff |
| Review | sudo -l -U <user> | Get-LocalGroupMember -Group 'Administrators' |
Windows LAPS is built into supported Windows versions, rotates the managed local administrator password regularly, and exists to stop one shared password enabling lateral movement from machine to machine (Microsoft Learn, Windows LAPS overview). Retrieval is a deliberate act: Get-LapsADPassword -Identity 'WEB01' returns the password as a SecureString unless you add -AsPlainText, which Microsoft says to use only in support or testing situations.
Linux sudo, tested. A sudoers file with a syntax error can lock out all admins, so it is validated before install. Executed in an ubuntu:24.04 container:
$ cat good.sudoers
deploy ALL=(root) NOPASSWD: /usr/bin/systemctl restart webapp
$ visudo -cf good.sudoers
good.sudoers: parsed OK
$ cat bad.sudoers
deploy ALL=(root) NOPASSWD /usr/bin/systemctl
$ visudo -cf bad.sudoers
bad.sudoers:1:46: syntax error
deploy ALL=(root) NOPASSWD /usr/bin/systemctl
^
$ install -m 0440 -o root -g root good.sudoers /etc/sudoers.d/deploy-webapp
$ sudo -l -U deploy
Matching Defaults entries for deploy on 75de9dd7b378:
env_reset, mail_badpass, secure_path=/usr/local/sbin\:/usr/local/bin\:/usr/sbin\:/usr/bin\:/sbin\:/bin\:/snap/bin, use_pty
User deploy may run the following commands on 75de9dd7b378:
(root) NOPASSWD: /usr/bin/systemctl restart webapp
deploy can restart one service as root and nothing else. The NOPASSWD tag removes the password prompt for that one command, which is acceptable for automation but should not be a blanket on ALL. Be careful with wildcards in command paths and with commands that can launch a shell (editors, pagers): the sudoers manual warns that wildcards can grant more than intended, and the NOEXEC tag stops dynamically linked programs from spawning further commands.
If you do three things first, make them these: remove everyday admin rights, give every machine its own LAPS password, and run each service under its own account. The finer points are details for when you write the rules: NOEXEC is a sudoers tag that stops a permitted program from launching further programs, UAC is the Windows prompt that asks for confirmation before an admin action, and a SecureString is a .NET string type that is kept encrypted in memory and does not print as plain text.
Files and directories on Linux
Set the owner and group to the account that really needs access, remove the "other" bits, and add ACLs (access control lists, per-user or per-group permissions beyond the owner, group and other triplet) rather than widening the group. Executed in the same container:
$ chmod 750 /srv/app; chmod 2770 /srv/app/shared
$ setfacl -m u:deploy:rx /srv/app
$ getfacl -p /srv/app
# file: /srv/app
# owner: webapp
# group: webapp
user::rwx
user:deploy:r-x
group::r-x
mask::r-x
other::---
$ stat -c '%A %U:%G %n' /srv/app /srv/app/shared
drwxr-x--- webapp:webapp /srv/app
drwxrws--- webapp:webapp /srv/app/shared
$ touch /srv/app/oops; chmod 666 /srv/app/oops
$ find /srv -xdev -type f -perm -0002 -print
/srv/app/oops
Reading the digits: each digit is read 4 + write 2 + execute 1, for owner, group and other in that order. 750 is 7 (4+2+1, rwx) for the owner, 5 (4+1, r-x) for the group and 0 for everyone else, which is the drwxr-x--- in the stat line. 2770 has a fourth, leading digit for special bits, where 2 means setgid: setgid, then rwx for owner, rwx for group, nothing for others, shown as drwxrws--- with the s standing in for the group x.
mask::r-x is the ceiling for every named-user entry and for the group entry: an entry's effective permission is its own bits ANDed with the mask. Here deploy has r-x and the mask allows r-x, so deploy gets r-x. Had the mask been r--, deploy would have lost execute (the right to enter the directory) and getfacl would have printed #effective:r-- beside the entry. setfacl -m recalculates the mask for you unless told not to.
other::--- means a user who is neither webapp nor deploy cannot even list /srv/app. deploy can read and traverse but not write. The setgid bit (the s in drwxrws---) makes new files in shared inherit its group. The find command is the audit for the classic mistake, a world-writable file; run it on a schedule. The same applies to setuid programs, which run with their owner's rights: list them with find / -xdev -type f -perm -4000 and review the list against what the server's role needs.
Files and directories on Windows
NTFS (the Windows file system) permissions are held in an ACL on each folder, and by default folders inherit entries from their parent. Least privilege usually means breaking inheritance on a sensitive folder and granting the service account, administrators and SYSTEM only what each needs. This script was parse-checked with PowerShell 7 and every cmdlet, method and parameter was checked against Microsoft Learn; it was not executed, because it needs a Windows host (Set-Acl is Windows-only, and Get-LapsADPassword needs the LAPS module and an AD):
# 1. Who is in the local Administrators group on this host?
Get-LocalGroupMember -Group 'Administrators' | Select-Object Name, ObjectClass, PrincipalSource
# 2. Restrict D:\AppData to the service account, Administrators and SYSTEM
$path = 'D:\AppData'
$acl = Get-Acl -Path $path
$acl.SetAccessRuleProtection($true, $false) # stop inheriting, drop the inherited entries
$inherit = [System.Security.AccessControl.InheritanceFlags]'ContainerInherit,ObjectInherit'
$propagate = [System.Security.AccessControl.PropagationFlags]::None
foreach ($rule in @(
@('CORP\svc-webapp', 'Modify'),
@('BUILTIN\Administrators', 'FullControl'),
@('NT AUTHORITY\SYSTEM', 'FullControl'))) {
$ace = New-Object System.Security.AccessControl.FileSystemAccessRule(
$rule[0], $rule[1], $inherit, $propagate, 'Allow')
$acl.AddAccessRule($ace)
}
Set-Acl -Path $path -AclObject $acl -WhatIf # remove -WhatIf after reviewing
icacls $path
# 3. Break-glass: read one machine's local admin password from AD
Get-LapsADPassword -Identity 'WEB01' -AsPlainText
ContainerInherit,ObjectInherit makes each entry apply to the subfolders (containers) and the files (objects) inside D:\AppData, so the grant flows down the tree; PropagationFlags::None adds no extra limit on that flow. In icacls notation the same ideas are (CI) and (OI), and RX means read and execute.
SetAccessRuleProtection($true, $false) means "protect from inheritance, and do not preserve the inherited entries", so the folder ends with only the three explicit entries. The three grants are added to the same in-memory ACL before Set-Acl writes it, so administrators are never locked out part-way. The -WhatIf switch makes Set-Acl report what it would do without doing it, so while -WhatIf is present nothing is written and icacls $path still prints the old ACL. After you remove -WhatIf and run the script for real, icacls $path lists the resulting ACL for review, one entry per line in the form name:(inheritance flags)(rights), where (M) is Modify and (F) is Full control. You should see exactly the three entries above (service account (M), Administrators (F), SYSTEM (F)); an entry marked (I) or any other name means inheritance was not removed or an extra grant exists. This describes the documented format and was not run, since it needs a Windows host. icacls $path /save <file> stores a copy of the ACL that /restore can re-apply if the change goes wrong.
Share permissions and NTFS permissions are separate layers: when a folder is reached over a network share, the access a user gets is the more restrictive of the two, so a locked-down NTFS ACL is not undone by a generous share.
Services
Each service runs under its own account with only the rights it needs, and never as root or as a domain administrator. On Linux that is a dedicated user such as webapp that owns only /srv/app. On Windows it is a dedicated service account, like CORP\svc-webapp above, granted rights only on its own folders and not given interactive logon.
Worked example
A web server hosts an app owned by webapp, deployed by a pipeline user deploy, administered by two people. Result: webapp owns /srv/app with mode 750 (full access for the owner, read and traverse for the group, nothing for others); deploy gets read via an ACL and one sudo rule to restart the service; the two administrators have personal logins with sudo for a short named list; there is no shared password anywhere; and the find audit above returns nothing. On the Windows file server the same pattern gives one service account modify rights on D:\AppData, administrators and SYSTEM full control, and everyone else no entry at all.
Trade-offs and pitfalls
- Least privilege costs setup time and generates "access denied" tickets. Pay that cost once in staging, by running the application as its real service account there.
- Group-based grants decay: people change roles. Review membership of the privileged groups quarterly, with
Get-LocalGroupMemberandsudo -l -U. - Standing admin rights are the thing to remove first. Time-limited elevation (just-in-time access) is the stronger end state but needs tooling; do the cheap steps first.
- A broad sudo rule on
ALL, or a sudo rule that lets someone run an editor as root, is effectively root; read rules for what they can be made to do.
A hardened baseline keeps drifting once servers are in production. Design how you would detect drift against it across Windows and Linux hosts, and how you decide per setting between reporting only and fixing it automatically.
Sample Answer
Direct answer
Make the baseline a single versioned source of truth, check it at three points (image build, provisioning, and continuously at runtime), and decide per setting between report-only and auto-fix with a rule: auto-fix only when the correction is deterministic (the same input always gives the same result, with no per-host judgement), idempotent, reversible and cannot take a service or an administrator out; report and ticket everything else. Keep the scan results as dated, tamper-evident evidence (any later edit would be visible) so an auditor can see the history, not only today's state.
Design: three checkpoints and evidence
- Image build: scan every image before it is published and fail the build on violations. On Linux, OpenSCAP (an open-source compliance scanner) can evaluate a profile and write results and an HTML report with
oscap xccdf eval --profile <profile> --results-arf results.xml --report report.html <datastream>. A profile is a named subset of rules, such as a CIS level; a datastream is the single file that bundles the rules and the checks that test them; XCCDF is the XML format those rules are written in and results are reported in; on Windows, build from a hardened image and apply the settings through Group Policy. - Provisioning: enforce at first boot with configuration management (Ansible, Group Policy or an endpoint management tool) so no host starts life drifted.
- Runtime: re-check continuously (for example daily) in detect mode. Ansible can run read-only with
--check(and--diff), where modules that support check mode report the changes they would make. Note the documented limits: modules without check-mode support report nothing, and tasks whose conditions depend on registered results from earlier tasks produce no output in check mode. - Evidence: store each scan result (the XML results file or the exported report) with the host, time, baseline version and operator, in write-once storage (storage that accepts new files but refuses edits and deletions, for example object storage with a retention lock). This is what an auditor asks for.
Deciding per setting: report only or fix automatically
| Setting | Mode | Why |
|---|---|---|
| Disable root SSH login | Fix | Deterministic, no application depends on it, reversible |
| Disable X11 forwarding | Fix | Same |
| Limit SSH authentication tries | Report | Rarely harmful but needs review against automation that retries |
| Reverse-path filtering (a kernel check that drops a packet whose reply would leave by a different network interface than it arrived on) | Report | Can break asymmetric routing (traffic that arrives by one path and returns by another); owner must approve per host |
| Require SMB signing on Windows file servers | Report | Legacy clients lose access; needs a rollout plan |
Rules behind the table:
- Fix automatically when the setting is a pure tightening that no workload legitimately needs and the fix is idempotent (running it twice changes nothing).
- Report only when it can interrupt service or lock people out, needs a reboot, varies by host role, or when the same setting is repeatedly reverted by a person or program. Repeated reverts mean something is fighting your baseline, so raise an alert instead of fixing again.
- Every report-only drift gets an owner and a due date, or a time-limited exception with an expiry.
- Examples of fixes that are not deterministic, so they stay report-only: "set
rp_filterto whatever suits this host's routing" (the right value differs per host), "remove local administrators who look unfamiliar" (who is legitimate is a human judgement), and any change that only takes effect after a reboot.
How the pieces fit: the single versioned baseline is what everything is compared against; the three checkpoints decide when you look; the table and rules decide what happens on a mismatch. OpenSCAP and the registry script below are the detectors for Linux and Windows hosts, and Ansible check mode and Group Policy are how you detect or enforce through the tool a fleet is already managed with.
Worked example: Linux check that fixes some settings and reports others
This script reads a table of settings. Rows marked fix are rewritten when they drift. Rows marked report are only flagged. It takes a root directory so it can be tested safely.
#!/usr/bin/env bash
# Report-or-fix drift check for key/value settings in config files.
# Usage: drift-check.sh ROOT (ROOT is / on a real host)
set -euo pipefail
root=${1:-/}
# file | key | expected value | mode (report|fix)
settings=(
"etc/ssh/sshd_config|PermitRootLogin|no|fix"
"etc/ssh/sshd_config|X11Forwarding|no|fix"
"etc/ssh/sshd_config|MaxAuthTries|4|report"
"etc/sysctl.d/99-hardening.conf|net.ipv4.conf.all.rp_filter|1|report"
)
drift=0
for row in "${settings[@]}"; do
IFS='|' read -r file key want mode <<<"$row"
path="$root/$file"
have=$({ sed -E 's/[[:space:]]*=[[:space:]]*/ /' "$path" 2>/dev/null || true; } |
awk -v k="$key" '$1 == k { v = $2 } END { print v }')
if [[ "$have" == "$want" ]]; then
printf 'OK %-40s %s\n' "$key" "$have"
continue
fi
drift=$((drift + 1))
printf 'DRIFT %-40s have=%s want=%s mode=%s\n' "$key" "${have:-<unset>}" "$want" "$mode"
if [[ "$mode" == fix ]]; then
if [[ ! -f "$path" ]]; then
printf 'SKIP %-40s not fixed, %s does not exist\n' "$key" "$file"
else
if grep -qE "^[[:space:]]*${key}[[:space:]]" "$path"; then
sed -i -E "s|^[[:space:]]*${key}[[:space:]].*|${key} ${want}|" "$path"
else
printf '%s %s\n' "$key" "$want" >>"$path"
fi
printf 'FIXED %-40s now=%s\n' "$key" "$want"
fi
fi
done
echo "drift_count=$drift"
exit $(( drift > 0 ))
Driver and real output (run inside a Debian container with GNU sed):
rm -rf /tmp/root && mkdir -p /tmp/root/etc/ssh /tmp/root/etc/sysctl.d
printf 'PermitRootLogin yes\nMaxAuthTries 6\n' > /tmp/root/etc/ssh/sshd_config
printf 'net.ipv4.conf.all.rp_filter = 0\n' > /tmp/root/etc/sysctl.d/99-hardening.conf
bash drift-check.sh /tmp/root; echo "exit=$?"
bash drift-check.sh /tmp/root; echo "exit=$?"
DRIFT PermitRootLogin have=yes want=no mode=fix
FIXED PermitRootLogin now=no
DRIFT X11Forwarding have=<unset> want=no mode=fix
FIXED X11Forwarding now=no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=4
exit=1
OK PermitRootLogin no
OK X11Forwarding no
DRIFT MaxAuthTries have=6 want=4 mode=report
DRIFT net.ipv4.conf.all.rp_filter have=0 want=1 mode=report
drift_count=2
exit=1
Reading the script, line by line:
settings=( ... )is a table with one row per setting; the four fields in each row are separated by|.IFS='|' read -r file key want mode <<<"$row"splits one row on|into four variables (IFSis the field separatorreaduses,-rstops backslashes being treated as escapes, and<<<"$row"feeds the row toreadas its input). For the first row that givesfile=etc/ssh/sshd_config key=PermitRootLogin want=no mode=fix.- The
have=$(...)line finds the current value. The{ sed ... || true; }group matters: withset -euo pipefail, a missing config file would makesedfail and silently end the whole script with status 2 and no output. With the guard, a missing file simply reads as<unset>and is reported as drift (run on a root whosesshd_configalready meets the baseline and which has no sysctl file, it printsDRIFT net.ipv4.conf.all.rp_filter have=<unset> want=1 mode=reportanddrift_count=1). The same guard does not help afixrow, because the rewrite needs a file to edit: without the[[ ! -f "$path" ]]test in the fix branch, a root with nosshd_configmadegrepand then the>>append fail and the whole run stop after the first DRIFT line, with nodrift_countand exit status 1, which looks the same as ordinary drift. With the test, that row printsSKIP PermitRootLogin not fixed, etc/ssh/sshd_config does not exist, stays counted as drift, and the run finishes (on an empty root all four rows are drift, two of themSKIP,drift_count=4, exit 1; run in the same container).sed -E 's/[[:space:]]*=[[:space:]]*/ /'turns akey = valueline intokey valueby replacing the first=and the spaces around it with one space (lines already writtenkey valuepass through unchanged).awk -v k="$key" '$1 == k { v = $2 } END { print v }'then looks for lines whose first word equals the key, remembers the second word, and prints the last one remembered when the file ends. Run on a small file in a Debian container:
--- after sed:
PermitRootLogin yes
MaxAuthTries 6
MaxAuthTries 5
# X11Forwarding yes
--- awk for MaxAuthTries:
5
The commented-out X11Forwarding line is ignored because its first word is #, and when a key appears twice the last line wins. That is how sysctl files behave; sshd keeps the first value instead, one more reason the sketch's reading of a file can differ from the effective configuration.
[[ "$have" == "$want" ]]printsOKand moves to the next row; anything else counts as drift, incrementsdriftand printsDRIFT. Only rows with modefixreach the rewrite, and only when the file exists:sed -ireplaces the existing line, orprintf ... >>appends one if the key was missing.exit $(( drift > 0 ))ends the script with status 1 if any drift remains and 0 otherwise, because the arithmetic comparison evaluates to 1 or 0. Run in the same container,drift=0exits 0 anddrift=3exits 1.
The first run fixes two settings and reports two. The second run finds the fixed ones clean (idempotent) while the report-only ones stay open until a human acts. The non-zero exit code lets a scheduler or pipeline alarm on remaining drift.
Limits of this sketch, so you do not trust it blindly: it reads files, not effective configuration. The sshd manual says that for each keyword the first obtained value is used, and that keywords after a Match line apply only to matching connections, so an appended line can be shadowed or captured by a Match block (a section of the configuration whose settings apply only to connections matching a condition, such as a particular user or source address). In production, read the effective SSH configuration with sshd -T, which prints it, rather than grepping the file.
Worked example: Windows registry-backed setting
On Windows the same decision logic applies; prefer Group Policy or endpoint management to enforce, and use a script only for what they do not cover. This reports drift for SMB signing and fixes it only when run with -Fix:
param([switch]$Fix)
$path = 'HKLM:\SYSTEM\CurrentControlSet\Services\LanmanServer\Parameters'
$name = 'RequireSecuritySignature'
$want = 1
$have = Get-ItemPropertyValue -Path $path -Name $name -ErrorAction SilentlyContinue
if ($have -eq $want) { "OK $name = $have"; return }
"DRIFT $name have=$have want=$want"
if ($Fix) {
Set-ItemProperty -Path $path -Name $name -Value $want
"FIXED $name now=$want"
}
Run it on a Windows host, since the registry provider is Windows-only. RequireSecuritySignature is the registry setting Microsoft names for the "always sign" SMB policy. It is deliberately report-only by default.
Trade-offs and pitfalls
- Auto-fix everything and you will eventually take down an application at 3 a.m.; report everything and drift grows faster than people close tickets. The per-setting rule is the compromise.
- Detecting drift is half the job: find out why it happened (manual change, a package script, a stopped agent) or the same drift returns.
- Measure it: drift count per host per week, mean time to close report-only items, and number of exceptions past expiry.
- A scanner finding is only evidence if the scan could authenticate and covered the host.
Your team still hardens servers by hand and chases findings after each audit. Make the case to leadership for investing in automated baseline enforcement: what you would claim, how you would size the benefit honestly, and how you would phase the work.
Sample Answer
Direct answer
Make the case in three parts. Claim what is true: automating baseline enforcement removes repeated manual labour and cuts the window between a change and its detection, and it makes audits evidence-driven instead of a scramble. Size the benefit in hours you can count from your own tickets, show the payback including the ramp, and say plainly what you cannot size (risk reduction). Then propose phases so leadership funds the first one and sees evidence before funding the next.
What to claim, and what not to
Claim: fewer hands-on hours per build, fewer audit findings that need chasing, faster detection of drift (a setting changing after the server was hardened), and a record of compliance on every host. Do not claim: "no more breaches", a dollar figure for risk avoided that you cannot derive, or that audits disappear.
Sizing the benefit honestly
Use inputs you pull from ticket history and time records. The numbers below are assumed inputs for the arithmetic; replace them with your own measured counts.
- Manual hardening: 80 new builds per year x 6 hours each = 480 hours.
- Audit follow-up: 120 findings per year x 3 hours each = 360 hours.
- Current total: 480 + 360 = 840 hours per year.
With automation:
- Running the automated build: 80 x 0.5 hour = 40 hours.
- Findings drop to 30 percent of today: 120 x 0.30 = 36 findings x 3 hours = 108 hours.
- New total: 40 + 108 = 148 hours per year.
- Saving: 840 - 148 = 692 hours per year, which is about 57.7 hours per month.
Cost: assume 400 hours of engineering to build and test the baseline and pipeline.
The naive payback is 400 / 692 x 12 = 6.9 months. Payback means the point where the cumulative saving has repaid the build cost. The naive figure is wrong because it assumes the full saving from month 1, but the benefit does not start on day one. A ramp is the period before the saving reaches its full size. Assume the ramp is: months 1 to 3 deliver nothing (the build phase), months 4 to 6 deliver half the monthly saving, and month 7 onward delivers the full saving. Treat the 400 hours as spent during months 1 to 3, so the whole cost is already counted before the first hour is saved.
The steps, from the sizing above:
- Full monthly saving: 692 / 12 = 57.67 hours. Half rate: 57.67 / 2 = 28.83 hours (692 / 24, unrounded).
- End of month 3: 0 saved, so the position is -400 hours.
- End of month 6: -400 + 3 months x 28.83 = -400 + 86.5 = -313.5 hours.
- End of month 12: -313.5 + 6 months (months 7 to 12) x 57.67 = -313.5 + 346.0 = +32.5 hours.
Month by month, in hours: month 3 -400.0, month 4 -371.2, month 5 -342.3, month 6 -313.5, month 7 -255.8, month 8 -198.2, month 9 -140.5, month 10 -82.8, month 11 -25.2, month 12 +32.5. The position crosses zero during month 12, so payback arrives in month 12, not month 7. Say the later number. It survives scrutiny.
All the figures in this sizing are in hours on one basis; do not convert some to money and leave others in hours.
Phasing
| Phase | Months | Scope | Exit evidence |
|---|---|---|---|
| 1 | 1 to 3 | Pick one baseline (a CIS benchmark profile or your own list), automate scan and report-only on a pilot group; build the exception process | Pilot hosts reported, owners agreed on what may auto-fix |
| 2 | 4 to 6 | Enforce at provisioning for new builds; auto-fix only the safe settings | New builds pass at first boot |
| 3 | 7 onward | Runtime drift detection and evidence storage for the whole fleet | Audit evidence pulled from the system, not assembled by hand |
CIS here means the Center for Internet Security, which publishes hardening benchmarks.
Risks to name and the honest limit
- Automation applied wrongly can break production, so phase 1 is detect-only and the auto-fix list is short.
- The benefit assumes findings truly fall by 70 percent; check this at the end of phase 2 and report the real number, even if it is lower.
- Risk reduction is real but not sizeable in hours. Present it as a qualitative benefit, and tie it to the audit findings you can count.
What would change the recommendation
If you build only 10 servers a year, the manual cost is small and the case rests on audit evidence alone. If the fleet is mostly managed images already, enforcement is mostly built and the investment is smaller.
Corporate access from personal laptops and phones must depend on the device being in a safe state. Which host-side conditions would you check, what happens to a device that fails, and how do you avoid pushing users toward workarounds?
Sample Answer
Direct answer
Check a short list of host conditions that each map to a real attack path (disk encryption, operating system patch level, a protection agent that is installed and healthy, a firewall and screen lock, and not being rooted or jailbroken), plus whether the device is managed or not. Evaluate them at sign-in and again while the session lives. A device that fails is not simply blocked: it moves down a ladder (grace period, then restricted access, with an outright block reserved for a device that cannot be trusted at all, such as a rooted or jailbroken one), and an unmanaged personal device gets a lower access tier (a level of access, such as full, browser-only or none) by design rather than being asked to meet managed-device rules. Users go looking for workarounds when failure is opaque and the fix is slow, so the design effort goes into clear messages, self-service fixes and a pilot that measures the failure rate before anything is enforced. The identity rules that assign each tier are set by access policy; the host signals and the handling of a failing device are what the endpoint side owns.
Which host conditions to check, and how the signal is read
| Condition | What it protects against | Example of how the signal is read | Signal quality |
|---|---|---|---|
| Disk encryption on | Data exposure from a lost or stolen device | Windows: Get-BitLockerVolume reports VolumeStatus (FullyEncrypted) and ProtectionStatus (On) for the OS volume. macOS FileVault, Linux LUKS and mobile device encryption are reported by the management agent | Good on a managed device |
| OS patch level | Known exploited vulnerabilities | Build or version at or above a minimum that you set (for example "within the last two monthly cumulative updates", which is a policy choice you tune, not a standard) | Good, hard to fake on a managed device |
| Protection agent present, running, current | Commodity malware, missing detection | Microsoft Defender: Get-MpComputerStatus exposes AMServiceEnabled, RealTimeProtectionEnabled and AntivirusSignatureAge. For a third-party endpoint detection and response (EDR) agent, use the agent's own health report | Moderate: an attacker with admin rights can stop an agent, so combine with a server-side "last seen" time |
| Firewall on, screen lock or passcode set | Opportunistic network and physical access | MDM (mobile device management) compliance report | Good on a managed device |
| Not rooted or jailbroken | Local controls bypassed by the user or malware | Platform integrity or attestation result via MDM (attestation is a signed statement, produced by the device's own platform and checked by a vendor or management service, that the operating system started unmodified; a rooted or jailbroken device cannot give a clean one) | Moderate to good; self-reported checks on a compromised device are not trustworthy |
| Managed or unmanaged | Decides the access tier rather than pass or fail | Device is enrolled in MDM or presents a device certificate (a certificate the management system installs on an enrolled device, so sign-in can tell a known device from an unknown one) | Strong when certificate-based |
Keep the list short. Every extra condition adds false failures, and each one should be defensible as "this blocks a specific attack". The strength of the signal matters more than the number of signals: a check that the device reports about itself proves less on an unmanaged phone than a hardware-backed result on a managed laptop, so unmanaged devices get less access instead of the same checks with less trust.
What happens to a device that fails
| State | Access | Why |
|---|---|---|
| Managed and healthy | Full access to the apps the user's role allows | The normal case |
| Managed, out of compliance (late patch, signatures stale, agent stopped or missing) | Grace period with a banner, then restricted access until fixed. For example 7 days for patch age, none for a stopped agent | Gives people time to apply a patch without punishing a vacation, but does not wait on an agent that is off |
| Unmanaged personal device | Web applications in the browser only, no downloads, no local sync, no client apps | You cannot verify or fix the device, so you limit what can leave through it |
| Rooted or jailbroken | Blocked for corporate resources | The device can no longer be trusted to enforce anything |
Written as access rules, in plain words (illustrative, not any product's syntax):
- If the device presents a valid device certificate and its compliance report says compliant, allow the apps the user's role needs.
- If it is enrolled but out of compliance, allow with a banner for the grace period, then restrict access until it is fixed.
- If it presents no device certificate (unmanaged), allow only browser sessions with downloads off.
- If the attestation result says rooted or jailbroken, block corporate resources.
How to avoid pushing users to workarounds
- Say exactly what failed and how to fix it. The block page should name the condition ("disk encryption is off"), give the fix in one step, and show a self-service button or link where the platform allows. A generic "device not compliant" message sends people to the help desk or to personal email.
- Make the fix quick. Push patches and enable encryption through management where you can, and keep a short path for people who cannot self-fix (a help desk queue with a target time).
- Grace periods tied to the cause, not a single number. A newly released patch needs days. A disabled protection agent does not.
- Leave an audited emergency path for the people who repair devices (help desk, on-call), so enforcement never blocks the fix.
- Respect privacy on personal devices. Collect only the signals in the table, say what is collected, and on an unmanaged device remove only corporate app data (selective wipe), never the whole device.
- Watch for workaround signals after launch: corporate mail forwarded to personal accounts, files moved through unsanctioned sharing tools, spikes in help desk tickets about access.
- Pilot in report-only mode first (the policy is evaluated for every sign-in and the result is logged, but nothing is blocked). Log who would have failed and why without enforcing, read the numbers, fix the noisy conditions, then enforce.
Worked example
A sales manager signs in from a personal Windows laptop whose disk is not encrypted and which is four months behind on updates. The device is unmanaged, so the rule that applies is the unmanaged tier, not a compliance check: the CRM works in the browser, downloads and desktop sync are off, and the page tells them that enrolling the laptop would earn full access. A different user's managed laptop first failed the patch check yesterday: the user sees a banner each day and keeps full access, and on the eighth day after that first failure access drops to the restricted tier until the update installs. The grace clock starts at the first failed check, not at the patch release date.
Planning the help desk load for the pilot (illustrative arithmetic, not a measurement): if the report-only run shows 6% of 1,000 managed devices would fail on patch age, that is 60 devices; at 15 minutes of help desk time each, enforcement day costs 900 minutes, or 15 hours. If that is too much, extend the grace period or push the patch first, then enforce.
Trade-offs and what would change the answer
- A stricter model (managed devices only, personal devices through a virtual desktop) is simpler to reason about and right for regulated data, but it costs more and hurts contractors. The tiered model is the better default when most unmanaged users only need web apps.
- If the workforce is mostly contractors on their own devices, protect the applications and data (app-level controls) rather than depending on host posture you cannot verify.
- Continuous evaluation (re-running these rules while the session is open, not only at sign-in) catches a device that drifts after sign-in, at the price of more disruptive mid-session restrictions. Use it for the conditions that change quickly (agent stopped) and not for slow ones (patch age).
Unlock Full Question Bank
Get access to all 10 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.