System and Endpoint Hardening Questions
Making operating systems, hosts, and endpoints resistant to compromise. Covers secure baseline configuration (CIS Benchmarks, Microsoft security baselines) and drift against the baseline, including detecting drift and deciding what to report versus auto-correct, OS and application hardening for Linux and Windows (SSH, host firewalls, service minimization, SELinux and AppArmor, file permissions, least privilege, application allow-listing, local administrator accounts), patch management and rollout (asset inventory, prioritisation, patch cadence, deployment rings and canaries, maintenance windows, emergency and out-of-cycle patching, post-patch verification, rollback, patch compliance metrics, immutable images, Windows and Linux update tooling such as Windows Update for Business, Intune, WSUS, Configuration Manager and Azure Update Manager), scripted audits and enforcement of host settings (Ansible, PowerShell, shell), and the host-side conditions that protect an endpoint (device posture checks, disk encryption, protection agent status). The host-level preventive layer. Detecting and investigating attacks, vulnerability scanning and scoring, network device and perimeter security, identity and key management, Active Directory attack hardening, operating WSUS or ConfigMgr as server roles, and container platform security are covered elsewhere.
Corporate access from personal laptops and phones must depend on the device being in a safe state. Which host-side conditions would you check, what happens to a device that fails, and how do you avoid pushing users toward workarounds?
Sample Answer
Direct answer
Check a short list of host conditions that each map to a real attack path (disk encryption, operating system patch level, a protection agent that is installed and healthy, a firewall and screen lock, and not being rooted or jailbroken), plus whether the device is managed or not. Evaluate them at sign-in and again while the session lives. A device that fails is not simply blocked: it moves down a ladder (grace period, then restricted access, with an outright block reserved for a device that cannot be trusted at all, such as a rooted or jailbroken one), and an unmanaged personal device gets a lower access tier (a level of access, such as full, browser-only or none) by design rather than being asked to meet managed-device rules. Users go looking for workarounds when failure is opaque and the fix is slow, so the design effort goes into clear messages, self-service fixes and a pilot that measures the failure rate before anything is enforced. The identity rules that assign each tier are set by access policy; the host signals and the handling of a failing device are what the endpoint side owns.
Which host conditions to check, and how the signal is read
| Condition | What it protects against | Example of how the signal is read | Signal quality |
|---|---|---|---|
| Disk encryption on | Data exposure from a lost or stolen device | Windows: Get-BitLockerVolume reports VolumeStatus (FullyEncrypted) and ProtectionStatus (On) for the OS volume. macOS FileVault, Linux LUKS and mobile device encryption are reported by the management agent | Good on a managed device |
| OS patch level | Known exploited vulnerabilities | Build or version at or above a minimum that you set (for example "within the last two monthly cumulative updates", which is a policy choice you tune, not a standard) | Good, hard to fake on a managed device |
| Protection agent present, running, current | Commodity malware, missing detection | Microsoft Defender: Get-MpComputerStatus exposes AMServiceEnabled, RealTimeProtectionEnabled and AntivirusSignatureAge. For a third-party endpoint detection and response (EDR) agent, use the agent's own health report | Moderate: an attacker with admin rights can stop an agent, so combine with a server-side "last seen" time |
| Firewall on, screen lock or passcode set | Opportunistic network and physical access | MDM (mobile device management) compliance report | Good on a managed device |
| Not rooted or jailbroken | Local controls bypassed by the user or malware | Platform integrity or attestation result via MDM (attestation is a signed statement, produced by the device's own platform and checked by a vendor or management service, that the operating system started unmodified; a rooted or jailbroken device cannot give a clean one) | Moderate to good; self-reported checks on a compromised device are not trustworthy |
| Managed or unmanaged | Decides the access tier rather than pass or fail | Device is enrolled in MDM or presents a device certificate (a certificate the management system installs on an enrolled device, so sign-in can tell a known device from an unknown one) | Strong when certificate-based |
Keep the list short. Every extra condition adds false failures, and each one should be defensible as "this blocks a specific attack". The strength of the signal matters more than the number of signals: a check that the device reports about itself proves less on an unmanaged phone than a hardware-backed result on a managed laptop, so unmanaged devices get less access instead of the same checks with less trust.
What happens to a device that fails
| State | Access | Why |
|---|---|---|
| Managed and healthy | Full access to the apps the user's role allows | The normal case |
| Managed, out of compliance (late patch, signatures stale, agent stopped or missing) | Grace period with a banner, then restricted access until fixed. For example 7 days for patch age, none for a stopped agent | Gives people time to apply a patch without punishing a vacation, but does not wait on an agent that is off |
| Unmanaged personal device | Web applications in the browser only, no downloads, no local sync, no client apps | You cannot verify or fix the device, so you limit what can leave through it |
| Rooted or jailbroken | Blocked for corporate resources | The device can no longer be trusted to enforce anything |
Written as access rules, in plain words (illustrative, not any product's syntax):
- If the device presents a valid device certificate and its compliance report says compliant, allow the apps the user's role needs.
- If it is enrolled but out of compliance, allow with a banner for the grace period, then restrict access until it is fixed.
- If it presents no device certificate (unmanaged), allow only browser sessions with downloads off.
- If the attestation result says rooted or jailbroken, block corporate resources.
How to avoid pushing users to workarounds
- Say exactly what failed and how to fix it. The block page should name the condition ("disk encryption is off"), give the fix in one step, and show a self-service button or link where the platform allows. A generic "device not compliant" message sends people to the help desk or to personal email.
- Make the fix quick. Push patches and enable encryption through management where you can, and keep a short path for people who cannot self-fix (a help desk queue with a target time).
- Grace periods tied to the cause, not a single number. A newly released patch needs days. A disabled protection agent does not.
- Leave an audited emergency path for the people who repair devices (help desk, on-call), so enforcement never blocks the fix.
- Respect privacy on personal devices. Collect only the signals in the table, say what is collected, and on an unmanaged device remove only corporate app data (selective wipe), never the whole device.
- Watch for workaround signals after launch: corporate mail forwarded to personal accounts, files moved through unsanctioned sharing tools, spikes in help desk tickets about access.
- Pilot in report-only mode first (the policy is evaluated for every sign-in and the result is logged, but nothing is blocked). Log who would have failed and why without enforcing, read the numbers, fix the noisy conditions, then enforce.
Worked example
A sales manager signs in from a personal Windows laptop whose disk is not encrypted and which is four months behind on updates. The device is unmanaged, so the rule that applies is the unmanaged tier, not a compliance check: the CRM works in the browser, downloads and desktop sync are off, and the page tells them that enrolling the laptop would earn full access. A different user's managed laptop first failed the patch check yesterday: the user sees a banner each day and keeps full access, and on the eighth day after that first failure access drops to the restricted tier until the update installs. The grace clock starts at the first failed check, not at the patch release date.
Planning the help desk load for the pilot (illustrative arithmetic, not a measurement): if the report-only run shows 6% of 1,000 managed devices would fail on patch age, that is 60 devices; at 15 minutes of help desk time each, enforcement day costs 900 minutes, or 15 hours. If that is too much, extend the grace period or push the patch first, then enforce.
Trade-offs and what would change the answer
- A stricter model (managed devices only, personal devices through a virtual desktop) is simpler to reason about and right for regulated data, but it costs more and hurts contractors. The tiered model is the better default when most unmanaged users only need web apps.
- If the workforce is mostly contractors on their own devices, protect the applications and data (app-level controls) rather than depending on host posture you cannot verify.
- Continuous evaluation (re-running these rules while the session is open, not only at sign-in) catches a device that drifts after sign-in, at the price of more disruptive mid-session restrictions. Use it for the conditions that change quickly (agent stopped) and not for slow ones (patch age).
A business-critical internal application depends on an old protocol version your baseline forbids, and it cannot be upgraded for six months. How do you contain the risk on the hosts involved while it stays up, and how do you stop the exception becoming permanent?
Sample Answer
Direct answer
Keep the forbidden protocol alive only on the one service that needs it, shrink who can reach that service to a named list of clients, wrap the process so a compromise of it goes nowhere, and record the whole arrangement as an exception with an owner and a hard expiry date that a script checks every night. The exception becomes permanent when nobody is forced to look at it, so the design goal is that the expiry date is enforced by the system, not remembered by a person.
The running example below is a billing application that still needs TLS 1.0 (an obsolete version of the Transport Layer Security protocol) on port 8443, approved on 2026-11-01 for six months, so it expires 2027-04-30. Network segmentation would add another layer, but the controls below work on the host itself, so they hold even if the network enforces nothing.
Step 1: Contain the risk on the host
| Layer | What it does | Where it sits |
|---|---|---|
| Scoped weakness | The legacy setting lives in this one application's own listener config, never in host-wide crypto settings. Every other service on the box keeps the baseline. | Application config |
| Host firewall allow-list | Only the two known client addresses may reach port 8443. Everyone else is dropped and logged. | The host's own packet filter, on the input path, so any packet addressed to this host must pass it |
| Service sandbox | The process runs as a dedicated unprivileged user, cannot gain privileges, sees a read-only filesystem except its data directory, and can only talk to listed addresses. | systemd unit |
| Reduced exposure | Remove every package, shell, compiler and unused listener the app does not need. Disable remote admin paths that are not required. | Host build |
| Watching | Dropped attempts are logged, and the vulnerability scanner finding is linked to the exception ID. | Logging and scanner |
Host firewall (nftables, the Linux packet filter). Rules run top to bottom and the first match decides, so the allow rule comes before the drop. The policy stays accept so only port 8443 is restricted and nothing else on the host changes:
table inet legacy_app {
set allowed_clients {
type ipv4_addr
elements = { 10.20.4.11, 10.20.4.12 }
}
chain input {
type filter hook input priority 0; policy accept;
tcp dport 8443 ip saddr @allowed_clients accept
tcp dport 8443 counter log prefix "legacy-app-denied " drop
}
}
Loaded with nft -f legacy.nft and then listed with nft list table inet legacy_app, the ruleset is accepted and prints back as:
table inet legacy_app {
set allowed_clients {
type ipv4_addr
elements = { 10.20.4.11, 10.20.4.12 }
}
chain input {
type filter hook input priority filter; policy accept;
tcp dport 8443 ip saddr @allowed_clients accept
tcp dport 8443 counter packets 0 bytes 0 log prefix "legacy-app-denied " drop
}
}
Reading the listing: the source said priority 0, and nft prints it back as priority filter, its name for 0 (the ordinary position for filtering rules among other chains on the same hook). hook input means the chain sees packets addressed to this host. ip saddr @allowed_clients matches a source address in the named set. counter packets 0 bytes 0 was added by the word counter in the drop rule: it is a tally that starts at zero and increments each time a packet matches that rule, so it is how you see denied attempts. log prefix "legacy-app-denied " writes a kernel log line with that tag for each match, and drop discards the packet silently.
That shows the syntax loads and the rules are in the order intended; it does not prove live traffic behaviour, so after deployment send a test connection from an allowed and a non-allowed host and confirm the drop counter moves only for the second. Persist the rules through the host's normal nftables configuration so they survive a reboot.
Service sandbox (systemd unit; once the program named in ExecStart exists on disk the file passes systemd-analyze verify with no output, because verify also checks that path and prints a "not executable" error while it is missing; a deliberately misspelled directive is reported, so the check does catch typos):
[Unit]
Description=Legacy billing app (TLS 1.0 exception, expires 2027-04-30)
[Service]
ExecStart=/opt/legacy-app/bin/server
User=legacyapp
NoNewPrivileges=yes
ProtectSystem=strict
ReadWritePaths=/var/lib/legacy-app
ProtectHome=yes
PrivateTmp=yes
CapabilityBoundingSet=
RestrictAddressFamilies=AF_INET AF_UNIX
IPAddressAllow=10.20.4.11 10.20.4.12 10.20.9.0/28
IPAddressDeny=any
$ systemd-analyze verify /etc/systemd/system/bad.service
/etc/systemd/system/bad.service:7: Unknown key name 'NoNewPrivilegez' in section 'Service', ignoring.
Why each line: ProtectSystem=strict mounts the whole filesystem read-only except the paths named in ReadWritePaths; an empty CapabilityBoundingSet= removes every Linux capability from the process (Linux splits the powers of the root user into separate units called capabilities, such as CAP_NET_ADMIN for changing firewall and network settings and CAP_SYS_ADMIN for mounting filesystems, so a process with none of them cannot do any of those root-only things even if it is compromised); NoNewPrivileges=yes stops it gaining privilege through execve() (the system call that starts a program), for example by launching a setuid program (a program that runs with its owner's rights); IPAddressAllow/IPAddressDeny filter both incoming and outgoing packets of the service, with the allow list checked first, so IPAddressDeny=any plus an allow list is a default-deny. The 10.20.9.0/28 entry stands for the application's own database or dependency range; add the local resolver too if the app needs one. This is a second, independent layer: if someone edits the nftables rules, the unit still limits where the process can talk.
Step 2: Stop the exception becoming permanent
- A written exception record: ID, host, control being waived, compensating controls (the table above), business owner, technical owner, approver, and an expiry date. The expiry is the vendor's committed upgrade date, not an open "until fixed".
- Expiry that fails loudly. A nightly job reads the exception register and exits non-zero for any expired row, which opens a ticket and pages the owner.
- Checkpoints on the project, not only on the exception. At the midpoint (2027-01-31) the upgrade must have a test build in hand; if it does not, the business owner and the approver meet to decide, with the real dates, whether to fund an alternative.
- Renewal is harder than the first approval. One extension of at most 30 days, signed by a more senior approver than the original, with the reason the date slipped written down. A second extension is a new exception.
- Scanner and baseline linkage. The scanner suppression for this host carries the exception ID and the same expiry, so when the exception lapses the finding reappears on its own.
Register and checker (bash, run with the date pinned so the result is repeatable). The register is the CSV file shown after the script, one exception per row. The billing exception was approved on 2026-11-01, so the check below is pinned to 2026-11-05, a few days after approval:
#!/usr/bin/env bash
# Fail (exit 1) if any approved exception has passed its expiry date.
set -euo pipefail
today=${TODAY:-$(date +%F)}
rc=0
while IFS=, read -r id host control expires owner; do
[[ $id == id ]] && continue
if [[ $expires < $today ]]; then
printf 'EXPIRED %s host=%s control=%s expired=%s owner=%s\n' "$id" "$host" "$control" "$expires" "$owner"
rc=1
fi
done < "${1:?usage: check-exceptions.sh exceptions.csv}"
exit "$rc"
id,host,control,expires,owner
EX-101,billing01,legacy-tls,2027-04-30,jlee
EX-102,billing02,legacy-tls,2026-09-30,jlee
EX-103,hr07,smb1,2026-12-31,mpatel
$ TODAY=2026-11-05 ./check-exceptions.sh exceptions.csv
EXPIRED EX-102 host=billing02 control=legacy-tls expired=2026-09-30 owner=jlee
$ echo $?
1
Only EX-102 is reported because its date is before the pinned day; the other two have later dates. The comparison works because ISO dates (YYYY-MM-DD) sort the same alphabetically as chronologically. A row whose expiry equals the checked day is not yet reported (the test is "earlier than today"), and it is reported the following day.
Reading the checker, line by line:
| Line | What it does |
|---|---|
set -euo pipefail | Stop on any failed command (-e), treat an unset variable as an error (-u), and count a pipeline as failed if any stage fails (-o pipefail) |
today=${TODAY:-$(date +%F)} | Use the TODAY environment variable if set (that is how the run above is pinned); otherwise ask date for today in YYYY-MM-DD form |
rc=0 | Remember the exit code to return: 0 means no expired rows |
while IFS=, read -r id host control expires owner; do | Read the file one line at a time. Setting IFS=, makes a comma the field separator, so a line like EX-102,billing02,legacy-tls,2026-09-30,jlee is split into the five named variables; -r stops backslashes being interpreted |
[[ $id == id ]] && continue | The first line of the CSV is the header, where the first field is literally id; skip it |
[[ $expires < $today ]] | Inside [[ ]] the < compares the two strings alphabetically, which is correct for ISO dates |
printf 'EXPIRED ...', rc=1 | Print one line per expired row and remember that the run must fail |
done < "${1:?usage: ...}" | The < feeds the file named by the first argument into the loop. ${1:?message} stops the script with that usage message if no argument was given |
exit "$rc" | Exit 0 if nothing expired, 1 otherwise, so a scheduler or pipeline can open a ticket on failure |
Trade-offs and pitfalls
- Allow-list over "monitor only". Monitoring tells you after the fact; the allow-list means a compromised neighbour on the same network segment still cannot open a session to the weak listener.
- Do not lower host-wide settings to make the app work. That silently weakens every other service and is the usual way a "temporary" exception spreads.
- A compensating control is not a fix. It reduces the likelihood of exploitation; the exception record must say that residual risk remains and who accepted it.
- What would change the plan: if the vendor cannot commit to a date at all, treat it as a replacement project, not an exception; if clients are too many to list, the allow-list loses value and the risk acceptance must be escalated rather than quietly widened.
A small Linux web server serves only HTTPS to the public and takes SSH from a few admins. Design its host firewall policy, show the ruleset, and explain how you would test it and make it survive a reboot without cutting your own session.
Sample Answer
Direct answer
Use the kernel's own packet filter, nftables, on the server with a default-deny inbound policy: accept packets that belong to connections already allowed, accept loopback, accept a minimal set of ICMP, accept TCP 443 from anywhere, accept TCP 22 only from a named set of admin addresses, and drop everything else. Do not filter outbound for this host unless there is a reason (see the trade-off below). Test the file for syntax before applying it, apply it with a timed automatic rollback armed, confirm from a second, new session before cancelling the rollback, and make it permanent by putting it in the file the distribution's nftables service loads at boot.
If the host is already managed with ufw or firewalld, use that tool and express the same policy in it. Two firewall managers writing rules on one host is the failure to avoid.
Policy
| Traffic | Decision | Why |
|---|---|---|
| Inbound packets of an already-established or related connection | accept (first rule) | replies to the server's own outbound requests, and your existing SSH session, keep working while rules change |
Inbound packets in connection-tracking state invalid | drop | not part of any valid connection |
| Loopback interface | accept | local services talk to each other over it |
| ICMP and ICMPv6 error messages (destination-unreachable, time-exceeded, parameter-problem, packet-too-big) | accept | path MTU discovery (how hosts learn the largest packet that fits along the path) and error reporting fail without them. RFC 4890 section 4.3.1 says Destination Unreachable (all codes), Packet Too Big, Time Exceeded code 0 and Parameter Problem codes 1 and 2 must not be dropped, and section 4.3.2 says Time Exceeded code 1 and Parameter Problem code 0 normally should not be dropped; the ruleset accepts every code of all four types |
| ICMPv6 neighbour discovery (neighbour solicitation and advertisement, router advertisement) | accept | IPv6 hosts use them to find the other hosts on the link and the default gateway; blocking them breaks IPv6 on this host |
| Echo request (ping), IPv4 and IPv6 | accept, limited to 5 per second | lets monitoring and humans check reachability, with a cap on abuse |
| TCP 443 | accept from anywhere | the service this host exists to provide |
| TCP 22 | accept only from admin_v4 / admin_v6 sets | SSH exposed to the internet is the most attacked port; a handful of known addresses removes that exposure |
| Anything else inbound | drop (policy drop) | default deny: a service nobody planned for is not reachable by accident |
| Forwarding | drop | this is a server, not a router |
| Outbound | accept (see below) |
The firewall runs in the input hook of the server itself, so every packet addressed to the host passes it before any local program sees it. It complements, and does not replace, a cloud security group or a perimeter firewall: the host firewall is what still protects the server from other machines on the same network that have already been compromised.
The ruleset
#!/usr/sbin/nft -f
# Host firewall for a public HTTPS server with SSH from a few admin addresses.
# Replace only our own table, so rules from other tools (Docker, fail2ban) survive.
table inet filter
delete table inet filter
table inet filter {
set admin_v4 {
type ipv4_addr
elements = { 10.77.0.10, 10.77.0.11 }
}
set admin_v6 {
type ipv6_addr
elements = { fd00:77::10 }
}
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iifname "lo" accept
# IPv6 does not work without neighbour discovery
icmpv6 type { nd-neighbor-solicit, nd-neighbor-advert, nd-router-advert, destination-unreachable, packet-too-big, time-exceeded, parameter-problem } accept
icmp type { destination-unreachable, time-exceeded, parameter-problem } accept
icmp type echo-request limit rate 5/second accept
icmpv6 type echo-request limit rate 5/second accept
tcp dport 443 accept
tcp dport 22 ip saddr @admin_v4 accept
tcp dport 22 ip6 saddr @admin_v6 accept
}
chain forward {
type filter hook forward priority filter; policy drop;
}
chain output {
type filter hook output priority filter; policy accept;
}
}
Reading the ruleset, line by line
| Line | What it means |
|---|---|
table inet filter | A table is a named container for chains and sets. The family inet makes one table cover IPv4 and IPv6 together; filter is only the table's name |
set admin_v4 { type ipv4_addr; elements = { ... } } | A named list of addresses. Rules refer to it as @admin_v4, so the allowed admins live in one place |
chain input { | A chain is an ordered list of rules. Packets are checked against the rules from top to bottom |
type filter hook input priority filter; policy drop; | type filter: this chain accepts or drops packets. hook input: attach it where packets addressed to this host arrive (forward is for packets passing through, output for packets the host sends). priority filter: the order among chains on the same hook; filter is the standard name for priority 0. policy drop: a packet that reaches the end without being accepted is dropped |
ct state established,related accept | ct is connection tracking, the kernel's memory of each flow. A packet that belongs to a connection already allowed (a reply, or a related error message) is accepted straight away, which is why this rule comes first |
ct state invalid drop | Packets that fit no known connection are dropped |
iifname "lo" accept | iifname is the input interface name. lo is the loopback interface, used by programs on this host talking to each other |
icmpv6 type { nd-neighbor-solicit, ... } accept | Accept the listed ICMPv6 message types: neighbour discovery (nd-...) plus the error types |
icmp type echo-request limit rate 5/second accept | Accept ping requests only while fewer than 5 per second arrive; the excess does not match this rule, falls through to the end and meets policy drop |
tcp dport 443 accept | dport is the destination port: anyone may connect to HTTPS |
tcp dport 22 ip saddr @admin_v4 accept | Port 22 and the source address (saddr) is in the set. The ip6 saddr line does the same for IPv6 |
chain forward / chain output | forward drops everything (not a router); output is policy accept, so outbound is unfiltered |
Each term in the policy table has a line that implements it:
| Policy row | Ruleset line |
|---|---|
| connection tracking (replies keep working) | ct state established,related accept and ct state invalid drop |
| path MTU discovery | packet-too-big in the icmpv6 line, destination-unreachable in the icmp line |
| neighbour solicitation and advertisement | nd-neighbor-solicit, nd-neighbor-advert, nd-router-advert in the icmpv6 line |
| SSH from admins only | the two tcp dport 22 ... saddr @admin_... lines |
Notes on the file:
- The first two lines of the table section (
table inet filterthendelete table inet filter) make the file re-loadable: the first creates the table if it does not exist (so the delete never fails), the second removes it, and the definition that follows rebuilds it. Only this table is replaced, so tables created by other tools (Docker's, fail2ban's) are left alone.flush rulesetwould wipe those too. I loaded the file onto an empty ruleset, then again onto the existing table, and it succeeded both times; a tableinet othercreated beforehand survived. - The family
inetcovers IPv4 and IPv6 in one table, so there is no forgotten IPv6 path with no rules. The two address sets are separate because the address types differ. The example addresses are a documentation-style private range (10.77.0.0/24) and a unique-local IPv6 address; replace them with your admin or bastion addresses. - Sets can be changed without reloading the file:
nft add element inet filter admin_v4 '{ 10.77.0.99 }'adds an address. To make it permanent, add it to the file as well, or the next reload forgets it. ct stateis connection tracking: the kernel remembers each flow, which is why one rule accepts all replies.
Test it before trusting it
-
Syntax check without applying:
nft -c -f /etc/nftables.confparses the file and reports errors without changing the running ruleset. -
Behaviour test from outside, using throwaway containers. This needs Docker on a Linux machine. Three containers share one Docker network: a server with the ruleset loaded and listeners on 22, 443 and 8080, one probe at an admin address and one at an address that is not in the set (the network, names and addresses are illustrative):
bash# setup, run from a directory holding web-fw.nft (the ruleset above), echo_server.py, server.sh and probe.sh docker network create --subnet 10.77.0.0/24 fw-test-net docker run -d --name fw-server --network fw-test-net --ip 10.77.0.2 --cap-add NET_ADMIN \ -v "$PWD":/w ubuntu:24.04 bash /w/server.sh docker run -d --name fw-admin --network fw-test-net --ip 10.77.0.10 -v "$PWD":/w ubuntu:24.04 sleep 300 docker run -d --name fw-other --network fw-test-net --ip 10.77.0.50 -v "$PWD":/w ubuntu:24.04 sleep 300 # wait for the server to finish installing and loading (about 30 seconds), then probe from each client docker exec fw-admin bash /w/probe.sh docker exec fw-other bash /w/probe.sh # cleanup docker rm -f fw-server fw-admin fw-other && docker network rm fw-test-net--cap-add NET_ADMINlets the server container change its own firewall; the clients do not need it. The echo server is a few lines of Python that listens on port 22 and sends back whatever it receives:python# echo_server.py import socket s = socket.socket() s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) s.bind(("0.0.0.0", 22)) s.listen() while True: c, _ = s.accept() while True: d = c.recv(100) if not d: break c.sendall(d)bash# server.sh (runs in the server container: Ubuntu 24.04) apt-get update -qq && apt-get install -y -qq nftables python3 > /dev/null python3 /w/echo_server.py & # echo server on TCP 22 for p in 443 8080; do python3 -m http.server "$p" --bind 0.0.0.0 > /dev/null 2>&1 & done nft -c -f /w/web-fw.nft && nft -f /w/web-fw.nft wait # keep the container (and the listeners) alive # probe.sh (runs in each client container; the server is 10.77.0.2) for p in 22 443 8080; do timeout 2 bash -c "exec 3<>/dev/tcp/10.77.0.2/$p" 2>/dev/null \ && echo "port $p open" || echo "port $p blocked/timeout" doneReading the probe:
/dev/tcp/HOST/PORTis not a real file but a bash feature, andexec 3<>/dev/tcp/10.77.0.2/22makes bash try to open a TCP connection to that address and port as file descriptor 3. If the connection is accepted the command succeeds, so&&prints "open". A dropped packet gets no reply, so the attempt hangs untiltimeout 2kills it after 2 seconds and||prints "blocked/timeout". A closed port that is not firewalled would answer at once with a refusal and print the same blocked message. Without the finalwaitthe script ends, the container exits and every probe prints blocked/timeout. With it, the probes printed exactly the table below.Probe from port 22 port 443 port 8080 admin address 10.77.0.10 open open blocked/timeout other address 10.77.0.50 blocked/timeout open blocked/timeout This confirms all three decisions: the public service is reachable by everyone, SSH only by the admin set, and an unplanned port by nobody.
-
Check that reloading does not drop a live session. A client at an admin address opens a TCP session to port 22, the server's ruleset is reloaded with
nft -fwhile the session is open, and the client sends again:python# session_client.py <server-ip> import socket, sys, time s = socket.create_connection((sys.argv[1], 22), timeout=5) s.sendall(b"before"); print("echo 1:", s.recv(100)) time.sleep(6) # the server reloads its ruleset during this pause s.sendall(b"after-reload"); print("echo 2:", s.recv(100))Output:
echo 1: b'before'thenecho 2: b'after-reload'. The established-connection rule is what carries the session across the reload.
Applying it on a real host without cutting yourself off
Two ways lock-outs happen: a typo in the admin address set, and rules that load fine but block the path you are using. Guard against both:
# 1. save our own table so it can be put back (or removed again if it did not exist)
{ echo 'table inet filter'; echo 'delete table inet filter'; nft list table inet filter 2>/dev/null; } > /root/fw-before.nft
# 2. arm a rollback in 2 minutes (systemd transient timer)
systemd-run --unit=fw-revert --on-active=2min /usr/sbin/nft -f /root/fw-before.nft
# 3. validate, then apply
nft -c -f /etc/nftables.conf && nft -f /etc/nftables.conf
# 4. from a SECOND terminal, open a new SSH connection from an admin address
# and request the HTTPS page from outside; only then cancel the rollback:
systemctl stop fw-revert.timer
Reading step 1: the curly braces form a brace group, which runs the three commands one after the other and lets one > send their combined output to a single file. The file therefore starts with table inet filter (create the table if it is missing, so the next line cannot fail) and delete table inet filter (remove it), followed by nft list table inet filter (the current definition of that table in loadable form, or nothing if there is no such table). Loading the file later first removes our table and then re-creates the old one if there was one. Only our own table is touched, which matches how the ruleset file itself is loaded. In step 2, systemd-run --on-active=2min schedules that load to run once, 2 minutes from now.
If the new rules strand you, you do nothing and the timer restores the old behaviour in 2 minutes. The saved file begins with the create-then-delete pair because loading a file does not remove rules that are already present; I checked that loading an empty file left the new rules in place, that the saved file removed the table again when none existed beforehand, and that, when one did exist, adding a rule and then loading the saved file brought back exactly the saved rules. I did not save the whole ruleset with flush ruleset because that is not safe on a host that also runs Docker: nft list ruleset prints the iptables-nft compatibility tables too, and in a test container nft -f refused to load them back with unsupported xtables compat expression. A file loads as one transaction, so that error would have meant the revert changed nothing and the lock-out stayed in place, and a successful flush would also have wiped the other tools' tables that the ruleset file takes care to leave alone. systemd-run --on-active= creates a transient timer, and with --unit= the timer is named fw-revert.timer; the revert command itself was not run here, because the test containers have no systemd. Keep out-of-band access (a cloud serial console or hypervisor console) ready, since it is the only route back if even the timer fails.
Surviving a reboot
On Debian and Ubuntu the nftables package ships nftables.service, whose ExecStart is /usr/sbin/nft -f /etc/nftables.conf, so the file above goes in /etc/nftables.conf and systemctl enable nftables makes it load at boot. On the RHEL family the service loads /etc/sysconfig/nftables.conf instead (in the AlmaLinux 9 package, whose ExecStart is /sbin/nft -f /etc/sysconfig/nftables.conf, that file ships with the sample include commented out), so use include there or put the rules in that file. To test persistence without waiting for a reboot, run nft flush ruleset from a console session and then systemctl restart nftables and list the ruleset; then do one real reboot in a maintenance window, with console access open, before calling it done.
Trade-offs and what would change the design
- Outbound
accept: a web server that can reach anywhere makes data theft and reverse shells easier after a compromise. The alternative is an outbound allow-list (DNS, time sync, package repositories, the specific APIs the application calls). I would take that for a server holding sensitive data, and accept the maintenance cost (every new dependency is a ticket). For a small server whose compromise has limited blast radius, accept-all outbound plus egress monitoring is a defensible choice. - SSH allow-list versus a bastion or VPN: if admins have changing addresses, an allow-list becomes constant churn and lock-outs. Allow SSH only from the bastion's address instead, and control who can use the bastion.
- Rate-limiting echo and new connections is a convenience, not a security boundary. Real denial-of-service protection happens upstream.
- Containers on the same host (Docker, Kubernetes) add their own rules and forwarding behaviour. The scoped
delete tableapproach avoids deleting them, but you must test the combination.
Pitfalls
- Putting
ct state established,related acceptafter the SSH rule, or omitting it, and cutting off the session when the policy loads. - Forgetting IPv6: an IPv4-only ruleset leaves the host wide open over IPv6 if it has an address, which is why the table is
inet. - Setting
policy dropon the input chain before confirming the ICMPv6 neighbour-discovery rules, so the host loses its default gateway and drops off the network. - Testing only from inside the host, which shows nothing about what the network can reach.
Explain what permissive and enforcing modes actually do in SELinux and how policy types differ. Then describe how you would let an in-house service bind a non-standard port and write to a custom log directory without weakening policy.
Sample Answer
Direct answer
SELinux (Security-Enhanced Linux) is mandatory access control: the kernel checks every access against policy in addition to ordinary file permissions. In enforcing mode it denies what policy does not allow and logs the denial. In permissive mode it allows everything and only logs what would have been denied (the log entries are called AVC messages, access vector cache denials). The policy type decides which processes are constrained: the targeted policy confines specific services and leaves other processes unconfined, while MLS (multi-level security) adds sensitivity levels for classified-style environments. To let an in-house service use a custom port and log directory without weakening policy, do not disable SELinux and do not stretch a shared type. Give the service its own confined domain and give the port and the log directory their own labels that only that domain may use.
A domain is the label a process runs under, and a type is the label on a file, port or other object; policy rules say which domains may do what to which types. Both are written the same way, as a name ending in _t (the process label httpd_t for Apache, the file label httpd_sys_content_t for web content), so "domain" is just a type attached to a running process. "Mandatory" means a process cannot opt out or edit these rules the way a file's owner can chmod their own file.
Modes
| Mode | Policy enforced? | What happens on a violation |
|---|---|---|
| Enforcing | Yes | Access denied and logged |
| Permissive | No | Access allowed, AVC logged |
| Disabled | No | SELinux off (getenforce reports Enforcing, Permissive or Disabled) |
setenforce 0 or setenforce 1 switches between permissive and enforcing at runtime, and sestatus shows the current mode and loaded policy type. Use permissive on the whole machine only briefly for diagnosis. The finer tool is semanage permissive --add <type>, which makes one domain permissive while everything else stays enforcing.
Policy types
- Targeted: the default on RHEL (Red Hat Enterprise Linux), protects the processes that policy defines a domain for. A daemon with no policy of its own commonly runs as
unconfined_service_t, which means SELinux is not constraining it. Check withps -eZ. - MLS: adds multi-level security labels on top of types and is used where information must be separated by classification.
The task: a custom port and a custom log directory
First look at what is really happening: ps -eZ | grep reportd (-Z adds the security label as the first column; reportd is an illustrative name for the in-house service) shows the domain the service runs in. Illustrative output (the label format is real, the PID and name are made up):
system_u:system_r:unconfined_service_t:s0 2210 ? 00:00:03 reportd
system_u:system_r:httpd_t:s0 1187 ? 00:00:09 httpd
Read the first column as user:role:type:level; the third part is the domain. httpd_t is a confined domain (Apache has policy written for it). unconfined_service_t is what a daemon started by systemd gets when no policy names it, and SELinux is not constraining it. If your service shows that, SELinux is not what stops it from binding the port or writing the log; look at ordinary permissions and the firewall instead. Confining it is the job of the module in step 3 below. The question assumes SELinux is involved, so assume the service runs confined as reportd_t.
Which tool to reach for depends on whether an existing type already fits. The everyday path is labels and booleans (an SELinux boolean is a built-in policy switch you turn on or off without writing rules). Writing a policy module is the occasional path, for a daemon policy has never heard of. Red Hat's guidance is to check labels and booleans before writing policy, because they are the smallest change. For an in-house service neither shipped type fits well (step 2 shows why), so the dedicated types come from the module and are then applied with the label commands.
-
Find the denial:
ausearch -m AVC,USER_AVC,SELINUX_ERR,USER_SELINUX_ERR -ts recentlists the denial records from the audit log from the last ten minutes (-mpicks message types,-ts recentthe time window), andsealert -l "*"prints each as a readable explanation. A raw record looks like this Red Hat documentation example (a different service, same format):texttype=AVC msg=audit(1226874073.147:96): avc: denied { getattr } for pid=2465 comm="httpd" path="/var/www/html/file1" dev=dm-0 ino=284133 scontext=unconfined_u:system_r:httpd_t:s0 tcontext=unconfined_u:object_r:samba_share_t:s0 tclass=fileRead it as:
{ getattr }is the action refused;comm="httpd"andscontextname who asked (domainhttpd_t);pathandtcontextname what was touched (typesamba_share_t);tclass=fileis the kind of object. Who was denied what is therefore: domainhttpd_twas refusedgetattron a file labelledsamba_share_t. For our service the records you expect are an actionname_bindon classtcp_socketwith a target port type such asunreserved_port_t(the default label for unassigned high ports), and awriteoradd_nameon classdirorfilewith a target type such asvar_t(the default label under/srv). -
Check the existing types and booleans. The port could be given
http_port_tand the log directory a shared log type, but both are shortcuts that weaken policy: every web-server domain on the host could then bind that port, and every daemon allowed to write the shared log type could read or write the service's logs. A dedicated type per object avoids that, which is why the module comes next and the labels after it. -
Write a small policy module for the service:
sepolicy generate --init /usr/local/bin/reportdcreates the type enforcement (reportd.te), interface (reportd.if) and file-context (reportd.fc) files plus a spec file and a build scriptreportd.sh. The generated policy starts with a permissive domain (the linepermissive reportd_t;), which only logs what would be denied. Declare the dedicated types first, apply the labels in step 4, and only then run the service under real load (step 5). Denials gathered before the labels exist would name the shared types (unreserved_port_tfor the port,var_tfor the log directory), and allowing those is exactly the weakening this design avoids. This module text was compiled in an AlmaLinux 9 container with the SELinux policy development package; it builtreportd.ppwithout errors, but it was not loaded, because a container has no SELinux to load it into:text# added to reportd.te (the generated file declares reportd_t and reportd_exec_t) type reportd_log_t; logging_log_file(reportd_log_t) # marks it as a log file type type reportd_port_t; corenet_port(reportd_port_t) # marks it as a network port type manage_files_pattern(reportd_t, reportd_log_t, reportd_log_t) allow reportd_t reportd_log_t:dir { add_name remove_name search write getattr }; allow reportd_t reportd_port_t:tcp_socket name_bind; allow reportd_t self:tcp_socket create_stream_socket_perms; corenet_tcp_bind_generic_node(reportd_t)Line by line: the two
typelines create the dedicated labelsreportd_log_tandreportd_port_t; themanage_files_patternline lets the domain create, write and delete files of the log type; thedirrule lets it add files to the log directory; thename_bindrule lets the domain bind a TCP port labelled with the port type and no other domain gets it; the last two lines let it create TCP sockets and bind to local addresses. Build and load it with the generatedreportd.shscript (it builds and loads the module); keeppermissive reportd_t;for now. -
Label the objects with the dedicated types. For the custom log directory, map a file context (a rule that says which label files at a path get) and apply it:
semanage fcontext -a -t reportd_log_t "/srv/reportd/logs(/.*)?"thenrestorecon -Rv /srv/reportd/logs(to ship the mapping inside the module instead, put an entry inreportd.fcin that file's own syntax, as in the generated line for the executable:/srv/reportd/logs(/.*)? gen_context(system_u:object_r:reportd_log_t,s0); thesemanagecommand line itself is not valid there). The pattern is a regular expression:(/.*)?means an optional slash followed by anything, so the rule covers/srv/reportd/logsitself and everything below it (checked withgrep -E:/srv/myweband/srv/myweb/a/b.htmlmatch^/srv/myweb(/.*)?$while/srv/mywebsitedoes not). Check withmatchpathcon -V(it compares the label on disk with what policy says it should be). For the port:semanage port -a -t reportd_port_t -p tcp 9876(list existing assignments withsemanage port -l). The documented example in Red Hat's guide useshttp_port_t, which is right for a web server. Illustrative result of a correct labelling, with the shape of the output fromls -Zd:system_u:object_r:reportd_log_t:s0 /srv/reportd/logs; before the fix the same command would showsystem_u:object_r:var_t:s0, the default under/srv. -
Run the service under real load with the labels in place and collect denials with
ausearch -m AVC -ts recent, then reviewaudit2allow -Rsuggestions (it turns denials into rules). Red Hat says audit2allow should not be the first option for a denial and warns the suggestions can be wrong in some cases, so read each rule. A suggestion that grants the domain access to a shared type such asvar_torunreserved_port_tmeans a label is missing or wrong, so fix the label instead of adding the rule. When the denials stop, delete thepermissive reportd_t;line, rebuild and reload withreportd.sh, and confirm there are no remaining denials in enforcing mode.
Because the port and log types belong only to the service's domain, no other process gains access. That is "without weakening policy".
Pitfalls
setenforce 0as a fix: it hides every problem on the host and is lost or reverted later.- Labelling the log directory with a broad type such as the system log type: other daemons may then read or write it.
- Pasting
audit2allowoutput without reading it. - Running the label commands before the module that defines the type exists:
semanagecan only use a type the loaded policy already contains. - Forgetting
restorecon: the mapping exists but files keep the old label.
Write PowerShell that audits the local Administrators group across domain-joined Windows servers and reports hosts with unexpected members. How do you handle credentials and unreachable hosts?
Sample Answer
Direct answer
Fan out one read-only PowerShell remoting call (PowerShell's way of running a command on another computer over WinRM, Windows Remote Management) to every member server, have each server return exactly one object (either its member list or the error it hit), compare the members to an allow-list on the auditing side, and report four statuses: OK, UNEXPECTED, ERROR and UNREACHABLE. The two design rules are that the script never handles a password (it runs under the operator's or a service account's Kerberos identity: Kerberos is the Windows domain sign-in protocol, and the script simply borrows the ticket of whoever runs it), and that a host which did not answer is reported as UNREACHABLE rather than silently counted as clean. Silence must never read as a pass.
The script
#Requires -Version 5.1
<#
Audit the local Administrators group on domain-joined member servers.
Runs under the caller's Kerberos identity (no passwords handled).
#>
[CmdletBinding()]
param(
[string[]]$ComputerName,
# Principals that are expected in local Administrators on every server.
[string[]]$Allowed = @('LOCAL\Administrator', 'CORP\Domain Admins', 'CORP\SRV-Local-Admins'),
[int]$ThrottleLimit = 32,
[int]$ConnectTimeoutSec = 20
)
# Pure comparison logic: no remoting, so it can be tested anywhere.
function Compare-AdminMember {
param(
[Parameter(Mandatory)][string[]]$Actual,
[Parameter(Mandatory)][string[]]$Allowed,
# Host being audited: HOST\name is rewritten to LOCAL\name so one allow-list fits every server.
[string]$HostName
)
$allowedSet = [System.Collections.Generic.HashSet[string]]::new(
[string[]]$Allowed, [System.StringComparer]::OrdinalIgnoreCase)
# Local accounts are prefixed with the short (NetBIOS) computer name, even when the host is addressed by FQDN.
$shortName = $HostName.Split('.')[0]
foreach ($member in $Actual) {
$normalised = $member -replace ('^' + [regex]::Escape($shortName) + '\\'), 'LOCAL\'
if (-not $allowedSet.Contains($normalised)) { $member }
}
}
if (-not $ComputerName) {
$dcNames = (Get-ADDomainController -Filter *).Name
$ComputerName = (Get-ADComputer -Filter 'OperatingSystem -like "*Server*" -and Enabled -eq $true').Name |
Where-Object { $_ -notin $dcNames }
}
$probe = {
# One object per host, success or failure, so silence can never mean "clean".
try {
$members = Get-LocalGroupMember -SID 'S-1-5-32-544' -ErrorAction Stop
[pscustomobject]@{
Error = $null
Members = @($members | ForEach-Object { $_.Name })
}
}
catch {
[pscustomobject]@{ Error = $_.Exception.Message; Members = @() }
}
}
$so = New-PSSessionOption -OpenTimeout ($ConnectTimeoutSec * 1000) -NoMachineProfile
$raw = Invoke-Command -ComputerName $ComputerName -ScriptBlock $probe `
-SessionOption $so -ThrottleLimit $ThrottleLimit `
-ErrorAction SilentlyContinue -ErrorVariable remotingErrors
$remotingErrors | ForEach-Object { Write-Verbose $_.Exception.Message }
$reached = @($raw | Select-Object -ExpandProperty PSComputerName -Unique)
$report = foreach ($r in $raw) {
$extra = if ($r.Error) { @() } else {
@(Compare-AdminMember -Actual $r.Members -Allowed $Allowed -HostName $r.PSComputerName)
}
[pscustomobject]@{
Host = $r.PSComputerName
Status = if ($r.Error) { 'ERROR' } elseif ($extra.Count) { 'UNEXPECTED' } else { 'OK' }
Unexpected = $extra -join '; '
Detail = $r.Error
}
}
$unreachable = $ComputerName | Where-Object { $_ -notin $reached } | ForEach-Object {
[pscustomobject]@{ Host = $_; Status = 'UNREACHABLE'; Unexpected = ''; Detail = 'no result returned (see -Verbose)' }
}
@($report) + @($unreachable) | Sort-Object Status, Host
Run it as .\Audit-LocalAdmins.ps1 -Verbose | Export-Csv admins.csv -NoTypeInformation from a machine with the Active Directory PowerShell module (installed with RSAT, the Remote Server Administration Tools) and an account that is allowed to open remoting sessions to the servers. Pass -ComputerName for a subset; with no parameter it discovers every enabled server computer account in the Active Directory (AD) domain except domain controllers.
What the report looks like. The four rows below are illustrative: the host names are made up, and I produced them by running the script's own comparison and report-building lines (unchanged, in PowerShell 7) on sample probe results, because the remoting itself needs Windows servers. The CSV is sorted by Status then Host:
"Host","Status","Unexpected","Detail"
"sql01","ERROR","","Access is denied."
"app01","OK","",
"file01","UNEXPECTED","CORP\jsmith; FILE01\svc_backup",
"web02","UNREACHABLE","","no result returned (see -Verbose)"
Reading it: OK means the host answered and every member is on the allow-list; UNEXPECTED lists the members that are not (Unexpected column), which is what you investigate; ERROR means the host answered but the group read failed, with the reason in Detail (the error text in the sample row is made up); UNREACHABLE means no result came back from that host: it did not answer, or it answered but refused the connection (for example access denied at logon), so nothing is known about its group. The reason, where Windows gave one, is in $remotingErrors and is printed with -Verbose. The empty cells are normal.
Why it is built this way
Reading the group. Get-LocalGroupMember -SID 'S-1-5-32-544' addresses the local Administrators group by its well-known SID (security identifier), so it still works on a server whose OS language renames the group. Note on scope: the Microsoft.PowerShell.LocalAccounts module is not available in 32-bit PowerShell on a 64-bit system. Domain controllers are excluded because they have no local account database of their own: the Administrators group there is a domain-wide built-in group, which needs a separate AD audit.
Credentials.
- No password appears in the script, a parameter, or a file.
Invoke-Commanduses the caller's current Kerberos credentials, and the script block touches nothing beyond the local server, so no credential has to be forwarded to a third machine, and CredSSP is not needed. The "second hop" picture: you run the script on workstation A, and it connects to server B (the first hop). If code running on B then had to reach a file share or server C (the second hop), B would need your credentials to authenticate there, and Kerberos does not hand B a reusable copy of them by default, so that call fails. CredSSP is the setting that works around this by sending your actual credentials to B, which is why it is best avoided. Here B only reads its own local group, so there is no second hop. - For a one-off with a different account,
Invoke-Command -Credential (Get-Credential)prompts and keeps the secret in aPSCredentialobject only in memory. - For a scheduled run, run the task as a group managed service account (gMSA): Active Directory rotates its password, so no secret is stored anywhere. Give that account the least privilege that allows opening a remoting session and reading local groups, and test it. If local administrator rights turn out to be required, treat the account as a privileged identity: it should not log on interactively anywhere, and its use should be alerted on, because an account that can read every server's admins can usually do more.
Unreachable hosts. A connection failure in Invoke-Command is a non-terminating error: the hosts that answer still return their results and the command continues. That is convenient, but it also means a dead host contributes nothing to the output. The script therefore computes UNREACHABLE as the requested names minus the names that returned an object (PSComputerName is added to every remote result). Reading the Invoke-Command call: -ComputerName and -ScriptBlock say where to run the probe and what to run; -SessionOption $so applies the timeout and no-profile settings; -ErrorAction SilentlyContinue stops each connection failure from printing a red error and halting the loop, so reachable hosts still return results; -ErrorVariable remotingErrors (written without a $) captures those same errors into the variable $remotingErrors so they are not lost, and -Verbose prints them. Errors are silenced, but hosts are never lost, because the UNREACHABLE list is computed afterwards from the names that did not answer. New-PSSessionOption -OpenTimeout is in milliseconds and defaults to 180000 (three minutes); lowering it to 20 seconds keeps dead hosts from stalling the run, and -ThrottleLimit caps concurrent connections (default 32). Error text from failed connections is available with -Verbose.
Errors on reachable hosts. The probe wraps the cmdlet in try/catch with -ErrorAction Stop and returns the message, so a host where the call failed shows as ERROR with its reason instead of an empty, apparently clean list.
The comparison is separate and testable. Compare-AdminMember takes the actual members, the allow-list and the host name, rewrites HOST\name to LOCAL\name (the local prefix is the short computer name, even when the host is addressed by its FQDN, fully qualified domain name) and returns whatever is not on the list. Testing it needs no Windows host. The script cannot simply be loaded into a test session, because loading it would run the remoting. So the first line uses Parser.ParseFile to read the script file and build a syntax tree (an AST, abstract syntax tree) without executing anything; the second line asks the tree for the first function definition in it, which is Compare-AdminMember; and the third wraps that function's text in a script block and runs it with a leading dot, which defines just that function in the current session:
$ast=[System.Management.Automation.Language.Parser]::ParseFile("$PSScriptRoot/Audit-LocalAdmins.ps1",[ref]$null,[ref]$null)
$fn=$ast.Find({param($n) $n -is [System.Management.Automation.Language.FunctionDefinitionAst]},$true)
. ([scriptblock]::Create($fn.Extent.Text))
$allowed='LOCAL\Administrator','CORP\Domain Admins','CORP\SRV-Local-Admins'
$actual='FILE01\Administrator','CORP\Domain Admins','CORP\jsmith','FILE01\svc_backup','corp\srv-local-admins'
Compare-AdminMember -Actual $actual -Allowed $allowed -HostName file01.corp.example.com
CORP\jsmith
FILE01\svc_backup
FILE01\Administrator matched LOCAL\Administrator, CORP\Domain Admins matched exactly, and corp\srv-local-admins matched despite the case difference. Only the two genuinely unexpected members came back. The whole script was parse-checked and its cmdlets and parameters were checked against Microsoft Learn; the remoting path itself needs Windows servers and was not executed here.
Complexity and edge cases
- Work is one remote call per server, run up to
-ThrottleLimitat a time, so the cost grows linearly with the server count; the comparison is O(members) per host with a hash set lookup. - Both lists are compared as strings, so the requested names and the reported
PSComputerNamemust use the same form. The AD discovery path uses theNameproperty throughout; if you supply FQDNs yourself, supply them consistently. - The allow-list matches by name. A renamed built-in Administrator account would be flagged, which is a safe failure but noisy. A hardened version compares SIDs.
- Domain groups in the allow-list are trusted as a unit: the cmdlet reports direct members, so who sits inside
SRV-Local-Adminsmust be audited in AD. - Servers that are powered off, in a different forest, or with WinRM (Windows Remote Management, the transport that PowerShell remoting uses) blocked all land in UNREACHABLE. Follow up on that list; a host that stays unreachable for several runs is a finding in itself.
PrincipalSource(Local, Active Directory, Microsoft Entra group, Microsoft Account) is only populated on Windows 10, Windows Server 2016 and later, so do not rely on it for older servers.
Trade-offs
A scheduled audit like this detects drift but does not prevent it. The preventive control is a Group Policy "Restricted Groups" or Local Users and Groups preference that sets the membership, with this report as the independent check that the policy is applying. Reporting only is deliberate here: an auditing script that also removes members can lock out an admin during an incident.
Unlock Full Question Bank
Get access to all 21 System and Endpoint Hardening interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.