System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
What are SELinux and AppArmor, and how does mandatory access control differ from the ordinary file permissions most people rely on? When would you insist on it and when would you hesitate?
Sample Answer
Direct answer
SELinux (Security-Enhanced Linux) and AppArmor are two implementations of mandatory access control (MAC) for the Linux kernel. Ordinary file permissions are discretionary access control (DAC): the owner of a file decides who may read or write it, and any program running as that user inherits all of that user's rights. MAC adds a second, system-wide policy written by the administrator or the distribution, which the kernel enforces on top of the permission bits and which a user or a compromised program cannot change at will. That lets you confine a program to what it legitimately needs, so a hijacked web server running as an ordinary account cannot, for example, read every file that account happens to own. I would insist on it for anything internet-facing or handling sensitive data, and hesitate only about how fast to enforce it on software nobody understands, never about whether to use it.
DAC versus MAC, concretely
| DAC (owner / group / other bits, ACLs) | MAC (SELinux or AppArmor) | |
|---|---|---|
| Who sets the rule | the file's owner (or root) | the system policy, loaded into the kernel |
| What the rule is about | which users may touch a file | which programs may touch which resources |
| Can a process widen its own access | yes, if it owns the object | no, the policy is outside its control |
| Order of checks | first | after DAC: the Red Hat documentation states that SELinux rules "are checked after DAC rules", so if DAC denies, no SELinux denial is even logged |
Both checks must pass. A MAC policy in effect narrows what DAC allows, which is why a permissions fix (chmod) alone does not cure a MAC denial.
SELinux and AppArmor: how they differ
| SELinux | AppArmor | |
|---|---|---|
| Identifies things by | labels: every process and file carries a context of user, role, type and level (user:role:type:level); the type (ending in _t, such as httpd_t for Apache) is what policy rules mostly use; on a process the type is also called its domain | paths: a profile per program lists the files and capabilities (Linux splits root's powers into separate privileges, such as binding a port below 1024 or changing file ownership) it may use |
| Default on | Red Hat family (Red Hat's documentation describes enforcing as the default and recommended mode) | Ubuntu (pre-installed and active) and, from Debian 10, Debian (enabled by default) |
| Modes | enforcing, permissive (logs but does not block), disabled | enforce, complain (logs but allows) |
| Strength | very fine-grained: covers processes, files, network ports and more | profiles are plain-text files under /etc/apparmor.d/, named after the program's path |
| Cost | steep learning curve; a wrong label on a file is a common cause of denials | a profile must name every path the program uses, and anything not listed is denied, so a program whose files move needs its profile updated |
An actual AppArmor profile fragment shipped on Ubuntu 24.04 (/etc/apparmor.d/lsb_release, trimmed) shows the idea: the program may read a few files, execute a few helpers, and nothing else:
profile lsb_release {
include <abstractions/base>
/dev/tty rw,
/usr/bin/lsb_release r,
/etc/lsb-release r,
/{usr/,}bin/bash ixr,
/usr/bin/cat ixr,
}
r is read, w is write, x execute (ix means the helper runs under this same profile), and anything not granted is denied. Letters combine, so /usr/bin/cat ixr, grants two things at once: execute cat under this same profile (ix) and read the file (r), which the kernel needs in order to load the program. A SELinux rule says the same thing in terms of types: "processes in domain httpd_t may read files of type httpd_sys_content_t".
Working example: a web server returns 403 on a correct-looking directory
Apache on a RHEL-family host serves /srv/myweb, owned by the right user with mode 755, yet returns 403. DAC is fine; the files carry the wrong type for a web server to read. The sequence:
-
getenforceto see the mode, andls -Zd /srv/mywebto see the directory's context (ls -Zprints the security context;-dlists the directory itself). Under/srvthe default label in the RHEL policy isvar_t(read from the policy's file-context list), a type Apache's domain is not allowed to read. Illustrative output before and after the fix in step 3 (formatuser:role:type:level; the type is the third field):textbefore: unconfined_u:object_r:var_t:s0 /srv/myweb after: unconfined_u:object_r:httpd_sys_content_t:s0 /srv/myweb -
ausearch -m AVC,USER_AVC,SELINUX_ERR,USER_SELINUX_ERR -ts recentsearches the audit log for denial records (an AVC, access vector cache, message is how SELinux logs a denial; the name comes from the kernel cache that holds access decisions), orsealert -l "*"for a readable explanation (-mselects the message types to list,-ts recentlimits the search to the last ten minutes, and-l "*"askssealertto describe every alert it has). A denial record for this story looks like this (illustrative, in the format SELinux writes):texttype=AVC msg=audit(1728200000.123:456): avc: denied { read } for pid=1840 comm="httpd" name="index.html" dev="dm-0" ino=412 scontext=system_u:system_r:httpd_t:s0 tcontext=unconfined_u:object_r:var_t:s0 tclass=file permissive=0Read it as:
{ read }is the action refused;comm="httpd"andscontextsay who asked (domainhttpd_t);nameandtcontextsay what was touched (a file of typevar_t);tclass=fileis the object kind;permissive=0means the host was enforcing, so the request really was blocked. -
Fix the label first:
semanage fcontext -a -t httpd_sys_content_t "/srv/myweb(/.*)?"records the rule permanently, thenrestorecon -R -v /srv/mywebapplies it. Thesemanage fcontextrule makes the labelling part of the policy, andrestoreconapplies it to the existing files. The pattern at the end of the path is a regular expression:(/.*)?means an optional slash followed by anything, so the rule covers/srv/mywebitself and every file and folder below it (checked withgrep -E:/srv/myweband/srv/myweb/a/b.htmlmatch^/srv/myweb(/.*)?$,/srv/mywebsitedoes not). -
If labelling is correct and the denial is a legitimate feature (the app must connect to a database), use a boolean, a switch built into the policy:
setsebool -P httpd_can_network_connect_db on. A non-standard port is likewise a label:semanage port -a -t http_port_t -p tcp 9876. -
Only if neither fits, write a small custom module. Red Hat explicitly warns against using
audit2allowto generate a local module as the first response to a denial, because it turns whatever was blocked, including a real attack, into permitted behaviour.
The temptation is setenforce 0, which switches the whole host to permissive. The better tool is to make one domain permissive while the rest stays enforcing: semanage permissive -a httpd_t, then remove it when done. Never set SELinux to disabled: with it disabled objects are not even labelled, which makes turning it back on painful.
The AppArmor equivalent: aa-status lists profiles and modes; aa-complain /path/to/bin moves one profile to complain mode while you collect what the program really does; edit /etc/apparmor.d/<profile>; apparmor_parser -r /etc/apparmor.d/<profile> reloads it; aa-enforce /path/to/bin returns it to enforcement.
When I would insist, and when I would hesitate
Insist when:
- the host is internet-facing (web servers, mail, VPN endpoints) or processes untrusted input, since confinement limits what a successful exploit can reach;
- the host is multi-tenant or runs code from several teams;
- a compliance baseline requires it (the SCAP Security Guide's CIS Server Level 1 profile for AlmaLinux 9, for example, includes the rules "Ensure SELinux is Not Disabled" and "Configure SELinux Policy", so a scan will flag a host that has it off);
- the host runs containers, where MAC is an additional layer between a container and the host (installing
apparmorandapparmor-utilson Ubuntu 24.04 places profiles forrunc,crun,podmanandbuildahin/etc/apparmor.d).
Hesitate (that is, stage it, do not skip it) when:
- a vendor application with unknown behaviour has files in unusual places and nobody can say what it legitimately touches: run it in permissive or complain mode with logging for a defined period, review the denials, then enforce with a date attached;
- the team has no one who can read a denial, in which case train someone before enforcing on a revenue system, not after the first outage;
- the application is in the middle of a migration and its file layout is still changing.
The pattern "hesitate" never means "disable". A permissive domain with logging still gives you the audit trail and a path to enforcement, while a disabled module gives you neither.
Pitfalls
- Fixing a denial by loosening DAC (
chmod 777), which weakens the layer that was working and does not address the MAC denial. - Generating a policy module from every denial without reading them.
- Restoring or copying data into place and assuming the labels are right: run
restoreconon it and check withls -Z. - Assuming AppArmor and SELinux are interchangeable in an operating procedure: the commands, defaults and failure modes differ, so write the runbook per distribution.
A configuration review of a Windows file server finds services running that nobody can justify. How do you decide what to disable, and which policy and endpoint controls keep it that way?
Sample Answer
Direct answer
Treat every running service as unjustified until someone can name the business function, the dependency, or the product feature that needs it. Inventory what is running and what is listening, classify each service as required by the file-server role, required by a dependency, or unowned, and disable the unowned ones in stages (stop, wait one full business cycle, then lock the start type through Group Policy). Then keep it that way with three layers: Group Policy sets the start type, the host firewall closes the ports the service opened, and endpoint controls plus auditing catch a service that is re-enabled or newly installed.
1. Build the evidence (inventory, then ownership)
Collect, for each server, what is installed and what is actually reachable:
Get-CimInstance -ClassName Win32_Service |
Select-Object Name, DisplayName, State, StartMode, StartName, PathName |
Sort-Object StartMode, Name
Get-NetTCPConnection -State Listen |
Select-Object LocalAddress, LocalPort, OwningProcess |
Sort-Object LocalPort
Illustrative output (a hand-written sample for a hypothetical server, not captured from a real one), trimmed to a few columns:
Name State StartMode StartName PathName
---- ----- --------- --------- --------
LanmanServer Running Auto LocalSystem C:\Windows\system32\svchost.exe -k netsvcs -p
Spooler Running Auto LocalSystem C:\Windows\System32\spoolsv.exe
VendorSync Running Auto LocalSystem D:\Tools\VendorSync\sync.exe
LocalAddress LocalPort OwningProcess
------------ --------- -------------
:: 135 1180
0.0.0.0 445 4
0.0.0.0 8443 4412
Reading a row: each line of the first table is one service, whether it is running now, how it starts, the account it runs as and the program behind it. In the second table each line is one listening port and the process ID that owns it. The third service stands out because its program sits under D:\Tools, outside the Windows and Program Files folders, and it runs as LocalSystem. StartMode is Auto, Manual or Disabled (also Boot and System for drivers). StartName is the account the service runs as: anything running as LocalSystem (a built-in account with full control of the local machine, and the computer's own identity on the network) or as a domain account (an account that works across the domain, so a stolen one reaches beyond this server) deserves extra scrutiny. PathName shows the binary: a service whose binary sits outside the Windows or Program Files folders is a question in itself. Match a listening port to a service through the process ID (OwningProcess against the service's ProcessId). One way to join the two lists, which I ran in PowerShell 7 against sample objects shaped like the real ones (illustrative values):
$byPid = @{}
Get-CimInstance Win32_Service | Where-Object ProcessId -gt 0 |
ForEach-Object { $byPid[[int]$_.ProcessId] += @($_.Name) }
Get-NetTCPConnection -State Listen | Sort-Object LocalPort |
Select-Object LocalAddress, LocalPort, OwningProcess,
@{ Name = 'Services'; Expression = { $byPid[[int]$_.OwningProcess] -join ',' } }
LocalAddress LocalPort OwningProcess Services
------------ --------- ------------- --------
:: 135 1180 RpcEptMapper,RpcSs
0.0.0.0 445 4
0.0.0.0 8443 4412 VendorSync
The first command builds a lookup from process ID to the names of the services living in that process (several services can share one process, so a row can list more than one: on a real server port 135 is the RPC endpoint mapper, which is why the sample shows the RPC services there; all service names in the sample rows are illustrative). The second adds a Services column to each listening port by looking up its owner. Port 8443 belongs to the unowned VendorSync agent, so that is a service worth asking about. A blank cell, as on port 445, means no service process owns the port; port 445 normally shows owner process ID 4, the Windows kernel's own System process, so blank is expected there. Before disabling anything, check what depends on it with Get-Service -Name <name> -DependentServices and what it requires with -RequiredServices; disabling a service that others depend on takes those down too.
Ownership: send the list of unexplained services to the application owner, the backup team, and the security team. "Nobody can justify it" is the starting position, but the answer has to be findable: a vendor agent installed by a team that left is still a service somebody owns.
2. Decide what to disable
Use a short decision rule, applied per service:
- Is it part of the file-server role or a dependency of something that is? Keep (document why).
- Is it a feature this server could use but does not (printing, remote registry access, a web server)? Candidate for disable. Fewer running services means a smaller attack surface (the set of code an attacker can reach), and some of these, such as the Print Spooler, have repeatedly been remote-exploit targets.
- Is it third-party and unowned? Stop it on a pilot server, and escalate to the owner; do not delete it first.
Worked example, one file server:
| Service | Evidence | Decision | How you verify it afterwards |
|---|---|---|---|
Server (LanmanServer) | Hosts the SMB (Server Message Block) file shares; the role depends on it | Keep | Get-Service LanmanServer is Running; shares reachable from a client |
Print Spooler (Spooler) | No printers or print queues on this server | Disable | Get-CimInstance Win32_Service -Filter "Name='Spooler'" shows StartMode Disabled and State Stopped |
Remote Registry (RemoteRegistry) | Nobody can name a tool that reads this server's registry remotely | Disable after asking the monitoring and backup owners | Same Win32_Service query shows Disabled; monitoring still green after a day |
Windows Search (WSearch) | Users search the shares through it | Keep | Get-Service WSearch is Running; a test search returns results |
| Vendor sync agent (unowned) | Runs as LocalSystem from a path outside the Windows and Program Files folders; no owner found | Stop on this server first, escalate to owners | Get-CimInstance Win32_Service -Filter "Name='<agent>'" shows Stopped; ask the owner to confirm in writing before removal |
Staged rollout: stop the service and set the start type to Disabled on one pilot server, watch for a full business cycle (including month-end jobs and the backup window, because a monthly job is exactly what a 24-hour test misses), and only then push the setting to the rest of the file servers. Keep the rollback to one command: Set-Service -Name Spooler -StartupType Manual followed by Start-Service Spooler.
3. Keep it that way: policy and endpoint controls
| Layer | Control | What it prevents or catches |
|---|---|---|
| Configuration | A Group Policy Object linked to the file-server OU, under Computer Configuration, Policies, Windows Settings, Security Settings, System Services, with each approved-off service set to Disabled | Drift: the start type is reapplied when Group Policy next processes the setting, so a local admin change is reverted on a later refresh rather than instantly |
| Network | Host firewall: inbound blocked by default, then explicit allow rules for only SMB and the management ports you use (on a locally managed computer -DefaultInboundAction already defaults to Block, so the point is to enforce it through Group Policy and keep the allow list short; Set-NetFirewallProfile -All -DefaultInboundAction Block is the local equivalent) | A service that is running but should not be reachable stays unreachable |
| Execution | Application allow-listing with App Control for Business or AppLocker, the two application control technologies Windows includes | With App Control for Business, a newly dropped service binary that no rule allows does not run (AppLocker has limits, described below) |
| Detection | Audit service installation: event 4697 "A service was installed in the system" in the Security log (subcategory Audit Security System Extension) | An unexpected new service raises an alert on a high-value server |
| Verification | A scheduled compare of Win32_Service against the approved list, with a report of anything not on it | Services that appear outside the change process |
Which of the two application control technologies to start with: Microsoft's guidance is to use App Control for Business (the newer feature, which decides which programs and drivers may run on the whole machine) wherever you can, because AppLocker still receives security fixes but no new features. AppLocker fits when you have older Windows versions in the mix or need different rules for different users on a shared computer.
Three limits of AppLocker matter for a services review. By default AppLocker policy only applies to code launched in a user's context, so a service running as SYSTEM is outside it unless you turn on its services enforcement (supported on Windows Server 2016 and later). Microsoft describes AppLocker as a defense-in-depth feature rather than a defensible security boundary, and recommends App Control for Business when the aim is robust protection. And AppLocker is not supported on Server Core installations, so check which installation option your file servers use before relying on it.
Two caveats about the audit event: it records the service name, binary path, start type and account at install time, so a later change to the binary path is not logged and has to be caught by process-creation auditing; and the event only exists if the audit subcategory (Audit Security System Extension, one switch in the Advanced Audit Policy settings that covers system-level changes such as installing a service) is enabled in policy.
Why a layered approach and not just Group Policy: the policy only controls the start type. It does not stop a service that is installed fresh, it does not close a port that something else opened, and an administrator or malware with admin rights can stop the policy from applying. The firewall, the allow-list and the audit event each cover a different gap.
Trade-offs and pitfalls
- Disabling is reversible, uninstalling is not. Prefer Disabled plus a recorded decision first, and remove the software later.
- Do not disable by name list copied from a hardening guide without checking this server's role: the same service that is a risk on one server is a dependency on another.
- Triggered services (services set to start only when an event happens, such as a device arriving, and to stop again afterwards) may look idle in a snapshot. Look at the start type, not only State.
- Record each decision (service, owner, date, evidence) so the next review does not start from zero, and review the approved list on a schedule.
You have just deployed a Windows Server that will run IIS for an internal business application. What hardening steps do you apply to the host and the web role after installation, covering the host firewall, service accounts and application pool identities, patching, and a configuration baseline? How would you notice later that it has been misconfigured or compromised?
Sample Answer
Direct answer
Harden in this order: patch the server before it takes traffic, shrink what is installed, close the network path to the one port and the one source network that need it, give the application the least privilege it can run with, tighten what the web server accepts and sends, then apply and verify a configuration baseline (a documented list of required settings). Afterwards, detect trouble with three feeds: configuration drift checks, web and security logs forwarded off the host, and endpoint protection alerts. This web server stays a domain member server and is not a domain controller.
Step by step
| Step | Action | How you know it worked |
|---|---|---|
| 1. Patch | Bring the OS, the .NET runtime and the IIS (Internet Information Services, the Windows web server) components fully up to date through the organization's patch server before opening the firewall to users | Windows Update Agent scan shows no applicable critical or security updates |
| 2. Remove unused role services | List the installed Web Server role services, uninstall the ones the application does not use (for example WebDAV publishing or FTP if present), and use -Remove to delete the payload so nobody re-enables them casually | Get-WindowsFeature listing matches the approved list |
| 3. Host firewall | Firewall on in all three profiles, default inbound action Block, allow only HTTPS from the application's client network and management from the admin network, log blocked traffic | Get-NetFirewallProfile -PolicyStore ActiveStore and the rule's address filter |
| 4. TLS | Disable legacy Transport Layer Security (TLS) protocol versions in Schannel (the Windows TLS provider) and control cipher suites through the cipher suite order, not by editing registry keys by hand | Registry value read-back, then an external TLS scan |
| 5. Identities | One application pool per application, running as the default ApplicationPoolIdentity; content read-only for that identity | The w3wp.exe process owner (Task Manager, or the Common Information Model (CIM) query below) is the pool's own identity, and the access control list (ACL) shows only intended grants |
| 6. Request filtering | Limit verbs, URL length, file extensions and suppress the server header | Request filtering settings in the site's web.config |
| 7. Logging | Keep IIS logs, enable process-creation auditing and forward everything to a collector | Test events arrive at the collector |
| 8. Baseline | Apply the organization's baseline (for example a Center for Internet Security (CIS) Benchmark profile for this Windows Server version, or the vendor's security baseline) through Group Policy and keep IIS configuration as files under source control | The commands in the verification table below, run on a schedule |
Commands
# 1. Review installed Web Server role services, then remove what the application does not need
Get-WindowsFeature | Where-Object { $_.Installed -and $_.Name -like 'Web-*' } |
Select-Object Name, DisplayName
Uninstall-WindowsFeature -Name $unneeded -WhatIf # $unneeded = names chosen from the list above
Uninstall-WindowsFeature -Name $unneeded -Remove # -Remove also deletes the feature files
# 2. Host firewall: on for every profile, inbound default Block, HTTPS only from the application subnet
Set-NetFirewallProfile -Profile Domain, Public, Private -Enabled True -DefaultInboundAction Block `
-LogBlocked True -LogFileName '%SystemRoot%\System32\LogFiles\Firewall\pfirewall.log'
New-NetFirewallRule -DisplayName 'IIS HTTPS from app subnet' -Direction Inbound -Action Allow `
-Protocol TCP -LocalPort 443 -RemoteAddress 10.20.0.0/24 -Profile Domain
# 3. TLS 1.0 off for the server role (a service or application restart may be needed)
$p = 'HKLM:\SYSTEM\CurrentControlSet\Control\SecurityProviders\SCHANNEL\Protocols\TLS 1.0\Server'
New-Item -Path $p -Force | Out-Null
New-ItemProperty -Path $p -Name Enabled -Value 0 -PropertyType DWord -Force | Out-Null
# 4. One application pool per application, default identity, read-only content
& "$env:windir\system32\inetsrv\appcmd.exe" set AppPool HrApp -processModel.identityType:ApplicationPoolIdentity
icacls D:\sites\hr\content /grant "IIS AppPool\HrApp:(OI)(CI)RX"
# 5. Verification
Get-NetFirewallProfile -PolicyStore ActiveStore | Select-Object Name, Enabled, DefaultInboundAction
Get-NetFirewallRule -DisplayName 'IIS HTTPS from app subnet' | Get-NetFirewallAddressFilter
Get-CimInstance Win32_Process -Filter "Name = 'w3wp.exe'" |
ForEach-Object { $o = Invoke-CimMethod -InputObject $_ -MethodName GetOwner; "$($o.Domain)\$($o.User)" }
Reading the less obvious pieces: w3wp.exe is the IIS worker process, the program that runs one pool's application code. In icacls ... /grant "IIS AppPool\HrApp:(OI)(CI)RX", IIS AppPool\HrApp is the pool's virtual account (an automatically managed identity with no password that does not appear in the user list), (OI) object inherit means files inside get the entry, (CI) container inherit means subfolders get it, and RX is read and execute. A security identifier (SID) is the unique ID Windows stores in an ACL in place of the name. In the TLS snippet, New-Item -Force creates the registry key for TLS 1.0 on the server side if it is missing, and New-ItemProperty ... -Name Enabled -Value 0 -PropertyType DWord writes the 32-bit number 0, which tells Schannel that this protocol version is off.
Notes on the commands:
- Firewall. The profile defaults on a computer are enabled and inbound Block, so the explicit settings document the intent and protect against a changed profile. Only the rule you add admits traffic, and its
-RemoteAddressaccepts a subnet in CIDR form. The subnet10.20.0.0/24is a placeholder for the application's client network. The firewall sits on the server itself, so it applies whatever network device is in front of it; a rule for the admin network (remote management) is added the same way with its own port and source. - Role services.
Uninstall-WindowsFeaturedoes not accept wildcards in-Name, so the names come from the review listing. Run with-WhatIffirst. Removing the payload with-Removemeans a later reinstall needs an installation source. - TLS. Schannel protocol settings live under
...\SCHANNEL\Protocols\<protocol>\Serverand the DWORDEnabledset to0disables that version. Microsoft documents that disabling versions can break interoperability, and that services may need a restart to pick the change up, so first inventory the clients of this internal application. Control cipher suites with the cipher suite order (Group Policy or the TLS PowerShell commands), not registry edits. - Identities. Since IIS 7.5 a new pool defaults to
ApplicationPoolIdentity, a virtual account namedIIS AppPool\<pool name>with its own security identifier, so permissions can be granted to one application only. Do not useLocalSystem(extensive privileges) orNetworkService(shared with other services, so a compromise of one can tamper with the others). If the application must reach a database or file share, the pool identity authenticates on the network as the computer account (DOMAIN\SERVER$), so grant that account the access on the remote side. Example: the pool onWEB01reads\\FS01\hr-reports.FS01has never heard of the localIIS AppPool\HrApp; it sees a request from the domain account of the machine,CORP\WEB01$(the trailing$marks a machine account), so the share and NTFS permissions onFS01must grantCORP\WEB01$, on that folder only. Every pool onWEB01shares that network identity, which is why one application that needs its own network identity gets a managed service account (a domain account whose password Active Directory rotates automatically). The managed service account is configured as the pool's custom identity. Grant the pool read and execute on the content folder only, and write only to the specific folders the application must write to (uploads, temporary files), with no execute permission there.
Request filtering and the server header
Request filtering is built in and rejects bad requests before application code runs. A web.config fragment (from the IIS documentation pattern) that denies directory traversal and alternate data streams, denies unlisted file extensions and verbs, and limits URL and query string lengths:
<configuration>
<system.webServer>
<security>
<requestFiltering removeServerHeader="true">
<denyUrlSequences>
<add sequence=".." />
<add sequence=":" />
</denyUrlSequences>
<fileExtensions allowUnlisted="false" />
<requestLimits maxUrl="2048" maxQueryString="1024" />
<verbs allowUnlisted="false" />
</requestFiltering>
</security>
</system.webServer>
</configuration>
Line by line: removeServerHeader="true" stops IIS announcing itself in the Server response header. denyUrlSequences rejects any URL containing .. (walking up out of the web root) or : (which can address an alternate data stream, extra hidden content attached to an NTFS file name). fileExtensions allowUnlisted="false" and verbs allowUnlisted="false" mean only the extensions and HTTP verbs you list are accepted. maxUrl="2048" and maxQueryString="1024" reject URLs longer than 2048 characters and query strings longer than 1024.
Because allowUnlisted="false" rejects everything not listed, add fileExtensions and verbs entries for what the application uses (for example .aspx and GET, POST, or PUT and DELETE for a REST API) or the application breaks. removeServerHeader requires IIS 10 on Windows Server version 1709 or later. Hiding the header removes a convenience for scanners but is not a security boundary.
Patching after day one
Patching has three parts: the OS and IIS through the patch server on a ring schedule (test server first, then production), the .NET runtime and IIS modules through the same channel, and the application's own libraries through its release pipeline. Check each cycle with an automated scan (a Windows Update Agent search such as (New-Object -ComObject Microsoft.Update.Session).CreateUpdateSearcher().Search("IsInstalled=0 and Type='Software'").Updates.Count, where 0 is the passing result) and keep an emergency path for out-of-band fixes.
How you notice a misconfiguration or compromise later
- Configuration drift. A scheduled check compares live settings with the baseline and reports to a central location (the verification table below lists the checks). Add file integrity monitoring (a tool that records a hash of each watched file and alerts when the hash changes) on
applicationHost.config, eachweb.configand the web root, so a new.aspxfile in an upload folder or a changed configuration raises an alert. - Web logs forwarded off the host. Request filtering returns HTTP 404 with a specific substatus when it blocks something:
404.5URL sequence denied,404.6verb denied,404.7file extension denied,404.11double-escaped URL,404.14URL too long. A burst of these from one source is scanning or an attack attempt, and a sudden stop of logs is also a signal. - Security log. Turn on Audit Process Creation (Advanced Audit Policy, Detailed Tracking) and the policy "Include command line in process creation events". Event 4688 then records the command line of every new process. The command line can contain secrets and is readable by anyone who can read the Security log, so restrict that access. A shell or scripting host started on an internal web server at the time of web requests is worth an alert. Also alert on a new service, a new scheduled task, or a new member of the local Administrators group.
- Endpoint protection. The endpoint protection or EDR (endpoint detection and response) agent reports malware detections and, where it records them, parent and child process chains.
- Forwarding. Send IIS logs and the Security log to a central collector as events happen (Windows Event Forwarding or an agent), so an attacker on the host cannot erase the evidence after the fact.
- Outside-in checks. A weekly external TLS and configuration scan confirms what clients really see.
Verification table
Every row in the baseline has a command that proves it:
| Setting | Required value | Verify with |
|---|---|---|
| Firewall profiles | Enabled, default inbound Block | Get-NetFirewallProfile -PolicyStore ActiveStore | Select-Object Name, Enabled, DefaultInboundAction |
| HTTPS inbound rule | Port 443, remote address is the app subnet only | Get-NetFirewallRule -DisplayName 'IIS HTTPS from app subnet' | Get-NetFirewallAddressFilter |
| Role services | Only the approved list installed | Get-WindowsFeature | Where-Object { $_.Installed -and $_.Name -like 'Web-*' } compared with the approved list |
| TLS 1.0 server | Enabled is 0 | Get-ItemProperty -Path 'HKLM:\SYSTEM\CurrentControlSet\Control\SecurityProviders\SCHANNEL\Protocols\TLS 1.0\Server' -Name Enabled |
| Pool identity | Worker process runs as the pool's virtual account | the Get-CimInstance Win32_Process owner query above |
| Content permissions | Pool has read and execute only | icacls D:\sites\hr\content |
| Command-line auditing | Audit Process Creation and the command-line policy enabled | auditpol /get /category:"Detailed Tracking" and a test event in the Security log |
What a correct result looks like (described from the commands' documented properties, not run, because they need a Windows host):
Get-NetFirewallProfile -PolicyStore ActiveStore: three rows (Domain, Private, Public), each withEnabledTrue andDefaultInboundActionBlock. A False or an Allow in any row is wrong.Get-NetFirewallRule ... | Get-NetFirewallAddressFilter: the remote address lists only the application subnet.Anymeans the rule admits everyone.- The
Get-CimInstance Win32_Processowner query: prints the pool's own virtual account, withHrAppafter theIIS APPPOOLprefix.SYSTEMorNETWORK SERVICEmeans the pool is running under a broader identity. icacls D:\sites\hr\content: an entry forIIS AppPool\HrAppending in(RX), and no(F),(M)or(W)for it.auditpol /get /category:"Detailed Tracking": the Process Creation line shows Success (or Success and Failure).No Auditingis wrong.
Pitfalls
- Locking down request filtering with
allowUnlisted="false"without listing what the application needs causes a production outage that looks like an application bug. - Disabling TLS versions without checking the clients breaks an internal client that is still on an old stack.
- A firewall rule that allows HTTPS from "Any" source makes the hardening cosmetic; the source network is the control.
- Logs kept only on the web server are lost when that server is compromised.
- A baseline that is applied once and never rechecked drifts within months; the recurring check is part of the control.
You have 300 pending patches this month and capacity for a fraction of them across a mixed Windows and Linux estate. How do you decide what goes first?
Sample Answer
Direct answer
I do not rank 300 patches by CVSS score (the Common Vulnerability Scoring System, a 0 to 10 measure of how severe a flaw is in the abstract). I rank them by how likely they are to hurt us this month: first evidence the flaw is being exploited, then whether the vulnerable thing is reachable by an attacker, then how much the host matters, then what already protects it, and last the cost and risk of the patch itself. That gives four tiers. The top tiers fill the capacity; everything else is deferred on purpose, with an owner and a date, and the ranking is re-run weekly because exploitation news changes it.
Signals
| Signal | What it tells you | Source |
|---|---|---|
| Known exploited | Attackers are using this flaw now | CISA's Known Exploited Vulnerabilities (KEV) catalog, which CISA says organisations should use as an input to vulnerability prioritisation; vendor advisories stating active exploitation |
| Likelihood of exploitation | Statistical estimate of exploitation soon | EPSS (Exploit Prediction Scoring System, from FIRST), a 0 to 1 probability that a published CVE will be exploited in the wild in the next 30 days |
| Severity | How bad it is if it works | CVSS base score; useful for tie-breaks, poor as the primary sort because it ignores whether anyone is exploiting it |
| Exposure | Can an attacker reach it? | Internet-facing, reachable from partner networks, internal only, isolated. "Reachable" means an attacker's traffic can actually get to the vulnerable service |
| Asset value | What the host holds or controls | Domain controllers, identity systems, database servers, jump hosts rank above test boxes |
| Compensating controls | Does something already block exploitation? | WAF (web application firewall) rule, network segmentation (splitting the network into zones with firewalls between them, so a host in one zone cannot be reached from the others), feature disabled, EDR (endpoint detection and response) with blocking enabled for that technique. A compensating control is any such measure that blocks the same attack when you cannot patch yet; detection-only coverage shortens response but blocks nothing, so it does not count |
| Patch cost and risk | Effort, downtime, chance of regression | Reboot needed, vendor warns of known issues, cluster failover required |
A CVE (Common Vulnerabilities and Exposures) identifier is just the public name of one flaw; one patch commonly fixes many CVEs, so I rank patches by the worst CVE they contain.
Tier rules
| Tier | Rule | Target |
|---|---|---|
| P0 | Known exploited AND the vulnerable component is present on an internet-facing host or a critical host | Out-of-band, within days, ahead of the normal ring schedule |
| P1 | Known exploited on internal hosts, OR high EPSS or critical CVSS on an internet-facing or critical host | This monthly cycle, first rings |
| P2 | High or critical CVSS on internal, non-critical hosts (a compensating control puts the patch at the back of the tier), or moderate flaws on exposed hosts | Next cycle unless EPSS or KEV changes |
| P3 | Low severity, no exploitation evidence, isolated or low-value hosts | Routine bundle, quarterly |
The numeric cut-off for "high EPSS" is a team decision that should be written down and revisited; one defensible way is to set it from how many patches you can absorb in a cycle rather than from a magic number.
Worked example with the numbers shown
Month backlog: 300 patches across Windows and Linux. Triage with the rules above gives:
| Tier | Patches |
|---|---|
| P0 | 12 |
| P1 | 41 |
| P2 | 97 |
| P3 | 150 |
| Total | 300 |
Capacity this month, from the team's test and reboot-window budget: 60 patches. Illustrative derivation: the test lab can validate 15 patches a week, and the month has 4 weeks, so 15 x 4 = 60 patches can be tested. There are 4 reboot windows a month and each can absorb 25 patches, so 4 x 25 = 100 can be rebooted. Capacity is the smaller of the two limits, which is 60, set by testing. P0 and P1 together are 12 + 41 = 53, which fills 53 of the 60 slots. The remaining 7 go to the top of P2, chosen by asset value (domain controllers and identity servers first). Deferred: 97 - 7 = 90 P2 patches plus 150 P3 patches = 240, and 60 + 240 = 300. Each deferred patch has a recorded reason, an owner and a review date, and the 240 are re-scored every week. If a P2 CVE lands on the KEV list on the 12th, it is re-tiered that day instead of waiting for the next cycle: P0 if it sits on an internet-facing or critical host, otherwise P1 under the rules above (known exploited on internal hosts).
Two illustrative patches show how the signals combine. Patch A fixes a remote code execution flaw in a web server component. It is on the KEV list (exploited now) and the component is installed on an internet-facing host, so it meets both halves of the P0 rule (known exploited, and present on an internet-facing host), and CVSS and EPSS are not needed to decide. Patch B fixes a privilege escalation flaw in a desktop library. It has a CVSS score of 8.8 (high), is not on the KEV list, has a low EPSS probability (say 0.02, a 2 percent chance of exploitation in the next 30 days), and is installed only on internal, non-critical hosts, behind a segment boundary with no route from the internet. No P0 or P1 rule matches, because it is not known exploited and it is not on an internet-facing or critical host. The P2 rule does (high CVSS on internal, non-critical hosts), so it is P2, and the segmentation (a compensating control) places it behind other P2 patches for the leftover slots, which are chosen by asset value; otherwise it waits for the next cycle.
Notice that capacity is the number of patches the team can safely test and reboot, not the number the tools can install. Grouping patches that need the same reboot, for example a kernel and its libraries together on Linux or the monthly cumulative update on Windows, raises effective capacity without raising risk.
Mixed Windows and Linux
Do not keep two ranked lists. Put both into one backlog with the same tier rules, then split by operating system only at scheduling time because the mechanics differ: Windows patches arrive in a monthly cumulative bundle, so you choose rings and timing, whereas Linux fixes arrive per package and can be selected individually, so you can pull a single critical package forward. Third-party applications on either platform go into the same ranking.
Pitfalls
- Sorting by CVSS alone sends the team after unexploited "critical" flaws while an exploited "medium" stays open.
- Ranking by CVE count per patch rewards bundles, not risk.
- Treating deferred items as accepted: deferral without an owner and date is how a backlog becomes permanent.
- Trusting scanner severity for a flaw in a component that is installed but never loaded; check reachability before spending an emergency slot.
- Using EPSS as an on/off switch. It is a probability, so a low score reduces priority but does not make a KEV-listed flaw safe.
You are handed a freshly provisioned Linux server that will host a production web service. How do you bring it to a secure baseline before it takes traffic, and how do you make that work repeatable rather than a one-off?
Sample Answer
Direct answer
I treat the first hour of a server's life as the cheapest time to harden it: it carries no traffic, no one depends on it, and a mistake costs a rebuild rather than an outage. In order, I patch it, remove or disable everything the service does not need (accounts, packages, listening ports), lock down remote access (key-only SSH, scoped sudo, default-deny firewall), set safe kernel and file permission defaults, turn on logging that leaves the host, and then prove the result with commands and a scanner. To make it repeatable I never do any of that by hand on the real server: the steps live in an image build and in configuration management (a tool such as Ansible that describes the desired state and applies it idempotently, meaning a second run changes nothing), and a scheduled run reports drift (a host that no longer matches the desired state).
The point of every step is to shrink the attack surface: the set of things an attacker can reach or abuse. An assessor looking at the host should see few open ports, few accounts that can log in, and no weak defaults left to find.
The eight steps, why each matters, and how I check it
| # | Step | Why it matters | How I check it |
|---|---|---|---|
| 1 | Apply pending security updates, then reboot if a kernel or libc (the C standard library nearly every program uses) changed | A fresh image is usually weeks old; known flaws are the easiest way in | On Ubuntu, apt list --upgradable should list nothing security-related afterwards |
| 2 | Accounts and sudo: remove unused accounts, give every person their own login, grant sudo to named commands rather than blanket root | Shared or unused accounts destroy accountability and widen the attack surface; narrow sudo limits what a stolen account can do | sudo -l -U <user> shows exactly what each person may run |
| 3 | SSH: keys only, no root login, no password or keyboard-interactive login | Password guessing against an internet-facing port is constant background noise | sshd -T prints the effective settings the daemon will use |
| 4 | Network exposure: stop services that are not part of the web service, then default-deny inbound with the firewall allowing only 443 and an admin source range | A port nobody listens on cannot be attacked | ss -tln before and after: only expected sockets remain |
| 5 | Kernel network parameters via a sysctl drop-in (sysctl reads and sets kernel settings; a drop-in is a small file in /etc/sysctl.d/ that is loaded at boot) | Stops the host honouring redirects (messages telling it to send traffic through a different router, which an attacker on the network can forge to divert traffic) and source-routed packets (packets that dictate their own path, which can steer around filtering), and logs impossible source addresses (a packet arriving from outside that claims a source such as 127.0.0.1 is almost certainly spoofed; the kernel calls these martians and the setting is log_martians) | Read each value back with sysctl -n <name> |
| 6 | File permissions: least-privilege owners and modes on the app tree, no world-writable files, review of setuid files (programs that run with their owner's rights) | A writable file in the wrong place turns a small bug into code execution | find listings (below) return nothing unexpected |
| 7 | Logging and audit: ship auth and system logs to a central collector, add audit rules for changes to /etc/ssh, /etc/sudoers.d and the account files | Logs kept only on the compromised host can be erased by the attacker | Generate a test event with logger and find it on the collector; make a test edit and find the audit record |
| 8 | File integrity: record a baseline of hashes for system binaries and /etc with a tool such as AIDE, check on a schedule | Detects a changed binary or config that no patch or ticket explains | A scheduled run that alerts on changes outside patch windows |
Repeatable: the same decisions as code
sshd reads the files in /etc/ssh/sshd_config.d/ in alphabetical order and, for each setting, the first value it finds wins. That is why the file is named 00-hardening.conf: it sorts before vendor or cloud-image drop-ins such as 50-cloud-init.conf, so its values cannot be overridden by them. The play's become: false means it does not escalate to root with sudo, because in the test container it already runs as root; on real hosts the play would use become: true.
This playbook was executed with ansible-core 2.16.3 inside an ubuntu:24.04 container (root, ansible-playbook baseline.yml). It writes the SSH drop-in, a validated sudoers rule and the sysctl drop-in, then fails the run if the effective SSH configuration is not what was intended.
- hosts: localhost
connection: local
become: false
tasks:
- name: sshd hardening drop-in (read first because of the 00 prefix)
ansible.builtin.copy:
dest: /etc/ssh/sshd_config.d/00-hardening.conf
mode: "0644"
content: |
PasswordAuthentication no
KbdInteractiveAuthentication no
PermitRootLogin no
- name: sudoers rule for the deploy user, syntax-checked before install
ansible.builtin.copy:
dest: /etc/sudoers.d/deploy-webapp
mode: "0440"
validate: /usr/sbin/visudo -cf %s
content: |
deploy ALL=(root) NOPASSWD: /usr/bin/systemctl restart webapp
- name: kernel network hardening drop-in
ansible.builtin.copy:
dest: /etc/sysctl.d/90-hardening.conf
mode: "0644"
content: |
net.ipv4.conf.all.accept_redirects = 0
net.ipv4.conf.all.send_redirects = 0
net.ipv4.conf.all.accept_source_route = 0
net.ipv4.conf.all.log_martians = 1
- name: refuse to continue if the effective sshd config is not what we intend
ansible.builtin.command: /usr/sbin/sshd -T
register: eff
changed_when: false
failed_when: "'passwordauthentication no' not in eff.stdout"
Run output (first run, then second run on the same host):
TASK [sshd hardening drop-in (read first because of the 00 prefix)] ************
changed: [localhost]
TASK [sudoers rule for the deploy user, syntax-checked before install] *********
changed: [localhost]
TASK [kernel network hardening drop-in] ****************************************
changed: [localhost]
TASK [refuse to continue if the effective sshd config is not what we intend] ***
ok: [localhost]
localhost : ok=5 changed=3 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
--- second run
localhost : ok=5 changed=0 unreachable=0 failed=0 skipped=0 rescued=0 ignored=0
The check task runs sshd -T, which refuses to start with "Missing privilege separation directory: /run/sshd" on a host where the SSH service has never run; on a real server ssh.service creates that directory, and in the bare test container it was created first with mkdir -p /run/sshd. The recap line says ok=5 although the play lists four tasks: Ansible runs an automatic Gathering Facts step first (it collects details about the host), and that counts as the fifth. In the first run three tasks changed something and the fourth only read the result, which is why it shows ok and changed=3. changed=0 on the second run is the idempotence proof. The sysctl file only takes effect once the settings are loaded (sysctl --system reads every system directory, -p reads one file), which I did not run because a container does not let you change host kernel settings. The kernel's own documentation (ip-sysctl) gives tcp_syncookies a default of 1 and rp_filter a default of 0; on the test kernel the defaults read back as rp_filter = 0, tcp_syncookies = 1, accept_redirects = 0, send_redirects = 1 and accept_source_route = 0, which is why the drop-in sets send_redirects explicitly rather than trusting the default.
Reverse-path filtering (rp_filter, which drops packets whose source address would not be routed back out the same interface) is deliberately not in the drop-in. Strict mode is correct for a single-homed web server but breaks hosts with asymmetric routing, and the kernel uses the larger of the all and per-interface values (the kernel documentation: "The max value from conf/{all,interface}/rp_filter is used"; the values are 0 off, 1 strict, 2 loose), so setting all to strict cannot be undone for one interface, and I decide it per host class. Asymmetric routing means replies leave by a different network interface than the one the request arrived on. Illustrative case: a server has a public interface and a management interface, and a request from a monitoring system arrives on the management interface while the server's default route sends replies via the public one. Strict mode checks "would I send a reply to this source out of the interface it came in on?", answers no, and silently drops the request.
File permissions on the application tree, as run in the same kind of container (users webapp and deploy created first):
$ chmod 750 /srv/app; chmod 2770 /srv/app/shared
$ setfacl -m u:deploy:rx /srv/app
$ getfacl -p /srv/app
# file: /srv/app
# owner: webapp
# group: webapp
user::rwx
user:deploy:r-x
group::r-x
mask::r-x
other::---
$ stat -c '%A %U:%G %n' /srv/app /srv/app/shared
drwxr-x--- webapp:webapp /srv/app
drwxrws--- webapp:webapp /srv/app/shared
$ touch /srv/app/oops; chmod 666 /srv/app/oops
$ find /srv -xdev -type f -perm -0002 -print
/srv/app/oops
The s in drwxrws--- is the setgid bit on a directory, so files created inside inherit the directory's group. The ACL (access control list, a per-user permission beyond owner, group and other) lets deploy read the tree without joining the webapp group, and other has no access at all. The find line is the audit: any file under the application tree that anyone can write is a finding. find /usr/bin -xdev -type f -perm -4000 lists setuid files; on the test image it printed chfn, chsh, gpasswd, mount and newgrp among others, and the review question is whether each is expected for this server's role.
Steps 4, 7 and 8 as commands
These three steps are not part of the playbook above; each was run in an ubuntu:24.04 container. The firewall needs --cap-add NET_ADMIN to run there.
Step 4, default-deny inbound with ufw (Ubuntu's firewall front end), allowing SSH only from an illustrative admin range (203.0.113.0/24 is a documentation address block) and 443 from anywhere:
ufw default deny incoming
ufw default allow outgoing
ufw allow from 203.0.113.0/24 to any port 22 proto tcp
ufw allow 443/tcp
ufw enable
ufw status verbose
Status: active
Logging: on (low)
Default: deny (incoming), allow (outgoing), deny (routed)
New profiles: skip
To Action From
-- ------ ----
22/tcp ALLOW IN 203.0.113.0/24
443/tcp ALLOW IN Anywhere
443/tcp (v6) ALLOW IN Anywhere (v6)
Run over an SSH session, ufw enable first asks "Command may disrupt existing ssh connections. Proceed with operation (y|n)?"; the rule for port 22 above is added before enabling so that your own session survives the prompt. Read the table as: anything inbound that has no row is dropped; port 22 answers only the admin range; 443 answers everyone, on IPv4 and IPv6.
Step 7, audit rules. Create /etc/audit/rules.d/50-hardening.rules (the augenrules tool merges every file in that directory into /etc/audit/audit.rules, which auditd, the Linux audit daemon, loads):
-a always,exit -F arch=b64 -F dir=/etc/ssh -F perm=wa -F key=ssh_config
-a always,exit -F arch=b64 -F dir=/etc/sudoers.d -F perm=wa -F key=sudoers
-a always,exit -F arch=b64 -F path=/etc/passwd -F perm=wa -F key=accounts
-a always,exit -F arch=b64 -F path=/etc/shadow -F perm=wa -F key=accounts
Each line says: always log, on exit from a system call, when a file in this directory (dir=) or this exact file (path=) is written to or has its attributes changed (perm=wa: w is write, a is attribute change; r and x exist too), and tag the record with a key so ausearch -k ssh_config finds it later. Running augenrules in the container merged the file into /etc/audit/audit.rules; loading rules into a kernel was not possible there, so the test edit that produces an audit record needs a real host. Shipping logs off the host (the other half of step 7) is done with your log forwarder, such as rsyslog or a vendor agent; logger "test" then proves the path.
Step 8, file integrity with AIDE (Advanced Intrusion Detection Environment). Illustrative minimal setup, using a throwaway configuration file that watches one directory and records permissions, owner, group, size and SHA-256 hash:
# /tmp/aide.conf
database_in=file:/tmp/aide.db
database_out=file:/tmp/aide.db.new
database_new=file:/tmp/aide.db.new
report_url=stdout
Rules = p+u+g+s+sha256
/srv/demo Rules
aide -c /tmp/aide.conf --init builds the baseline database as aide.db.new; copy it to aide.db (the file named by database_in) to make it the baseline, then aide -c /tmp/aide.conf --check compares the disk with it and a clean run exits 0. After echo "more text" >> /srv/demo/app.conf (the file held hello plus a newline, 6 bytes), the check printed the following, abridged: the real report also has dashed section headings, shows the new SHA-256 value beside the old one, and ends with a block of database checksums.
AIDE found differences between database and filesystem!!
Summary:
Total number of entries: 2
Added entries: 0
Removed entries: 0
Changed entries: 1
Changed entries:
f > ... H : /srv/demo/app.conf
Detailed information about changes:
File: /srv/demo/app.conf
Size : 6 | 16
The summary counts entries added, removed and changed; the detail block shows the stored value on the left of | and the current value on the right (the file grew from 6 to 16 bytes, and its SHA-256 differs as well). The exit status is a sum: 1 if files were added, 2 if removed, 4 if changed, so this run exited 4 and a clean run exits 0, which makes it easy to alert on from a schedule.
Worked example: attack surface before and after
Before step 4 a default image running only SSH listens on one port. On the test container, with sshd started on port 2222, ss -tln printed:
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 0 128 0.0.0.0:2222 0.0.0.0:*
LISTEN 0 128 [::]:2222 [::]:*
On a real web host the target list is explicit: SSH (restricted by the firewall to the bastion or admin range) and 443. Anything else that appears in ss -tln after provisioning (a database listening on all interfaces, a metrics agent, a leftover test service) is either bound to localhost, firewalled, or removed. That before and after listing is also what I hand a security assessor.
Trade-offs and pitfalls
- Developer workflow. Developers lose "just sudo to root and fix it". The fix is to give them what they actually do as named sudo rules (restart the service, read its logs), to deploy through the pipeline instead of by hand, and to give them a read-only way to see logs. Pushing back on blanket root is easier when the three commands they need work on day one.
- Emergency access. Key-only SSH plus a firewall can lock out the people who must fix an incident. Keep an out-of-band path (cloud serial console or the server's management controller) and a documented break-glass credential that is sealed, audited and rotated after use. Test it before you need it.
- Hardening that breaks the app. Apply the baseline to a staging host built from the same image, run the application's smoke tests, and only then promote the image. A baseline that is only ever tested on production eventually causes an outage.
- One-off hand edits. A hardened server that someone then tweaks by hand drifts. Either rebuild instead of editing, or have configuration management revert and report the change.
- File integrity is noisy until it is told about patch windows: baseline after patching, and alert on changes outside them.
Unlock Full Question Bank
Get access to all 14 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.