Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
As a Systems Administrator, configure sudo so members of group 'ops' can run only /usr/bin/systemctl restart httpd and /usr/sbin/ethtool without a password on all servers. Provide the exact /etc/sudoers lines (preferably with a Cmnd_Alias) and explain how to safely edit and test the configuration.
Sample Answer
Direct answer
Define a Cmnd_Alias naming exactly the two allowed commands with their full absolute paths and no arguments left open ended, then grant the ops group NOPASSWD access to only that alias, never to ALL. Full paths matter because sudo matches against the literal command path given in the rule; if a user's $PATH resolves systemctl to a different binary than the one named in sudoers, sudo simply will not match, which is a safety feature here, not a bug to work around.
The sudoers entry
Cmnd_Alias OPS_CMDS = /usr/bin/systemctl restart httpd, /usr/sbin/ethtool
%ops ALL=(root) NOPASSWD: OPS_CMDS
Cmnd_Alias OPS_CMDS = ...: names the exact set of allowed invocations./usr/bin/systemctl restart httpdonly matches that specific subcommand and unit; it does not grantsystemctl restarton any other unit, and does not grantsystemctl stop httpdor any other verb, because sudoers command matching is against the full command line as written, not just the binary./usr/sbin/ethtoolwith no arguments listed grants it with any arguments, since an alias entry with no trailing arguments matches the command run with any arguments at all; if the intent were to restrictethtoolto read only queries and forbid it from changing NIC settings, the entry would need to spell out the specific allowed flags explicitly (for example/usr/sbin/ethtool -S *for statistics only) rather than leaving it open.%ops: the%prefix means this line applies to a group, not a single user;opsmust exist as a real group (groupadd ops) with the intended members already added to it.ALL=(root): this rule applies regardless of which host the sudoers file is deployed to (ALL), and the command runs asrootspecifically, not as any user the invoking user chooses; if this should be able to run as some other target user too, that user would need to be added to the parenthesized run-as list.NOPASSWD:: skips the password re-prompt for exactly this alias, which is what makes it usable in scripts and automation, at the cost of removing that speed bump, which is why scoping the alias tightly matters more here than on a rule that still requires a password.
Safe editing and testing
Always edit sudoers with visudo (or visudo -f /etc/sudoers.d/ops-restart for a drop in file under /etc/sudoers.d/, the cleaner approach since it keeps custom rules out of the main file and out of the way of package updates), never with a plain text editor directly on /etc/sudoers: visudo locks the file against concurrent edits and, critically, refuses to save a syntactically broken file, which is the single biggest protection against locking every sudo user out of sudo entirely with one typo. visudo -c (or visudo -c -f <path> for a drop in file) checks syntax without opening an editor at all, which is the right tool for a script or a pre-deploy check rather than an interactive edit. Test with the actual target user, not root: sudo -l -U alice (as root) lists exactly what sudo believes alice is authorized to run, without executing anything, and sudo -u alice -l from alice's own session confirms what she sees; only after both agree should you consider the change live.
Worked example (executed)
groupadd ops
cat > /tmp/ops-sudoers <<'EOF'
Cmnd_Alias OPS_CMDS = /usr/bin/systemctl restart httpd, /usr/sbin/ethtool
%ops ALL=(root) NOPASSWD: OPS_CMDS
EOF
visudo -c -f /tmp/ops-sudoers
Real output confirming the syntax is valid before it is ever placed where sudo would actually read it:
/tmp/ops-sudoers: parsed OK
Only after that check passes would this be moved into place with visudo -f /etc/sudoers.d/ops-restart (which re-runs the same validation on save), never copied in directly with cp or a text editor, since neither of those re-validates syntax before the file takes effect.
Trade-offs and pitfalls
The most common mistake is writing the alias against a bare command name (systemctl instead of /usr/bin/systemctl), which either fails to match at all (if sudoers strictly requires the absolute path, which it does) or, worse on a misconfigured or old sudo build, matches more loosely than intended; always use the full path exactly as which systemctl (or command -v) reports it on the actual target host, since that path can differ between distributions. A second mistake is granting ethtool with no argument restriction when the actual intent was "let ops check NIC stats," since unrestricted ethtool can also change link settings, offload features, and other configuration that a "just checking" ticket did not intend to authorize; if the real requirement is read only, the sudoers entry should say so explicitly rather than relying on ops members' good judgment never to run a mutating flag. Third, never test a new sudoers rule by immediately logging out of your only privileged session: keep a second root session open, exactly as with PAM changes, until sudo -l -U <testuser> and an actual test invocation both confirm the rule behaves as intended.
Describe how you would use strace to diagnose a program that fails with 'Permission denied' when opening a config file. Include the strace invocation, how to capture output across forks, how to filter to file-related syscalls, and what specific return values you'll look for in the trace.
Sample Answer
Direct answer
Run the failing program under strace, filter to the file-related syscalls, and read the specific syscall's
return value and its error number (errno), since a "Permission denied" surfacing at the application level is
really the LAST in a sequence of file-related syscalls the kernel may have rejected, and strace shows
exactly which call, on exactly which path, with exactly which errno, the actual ground truth for what the
kernel refused and why.
Structured elaboration
Basic invocation:
strace -f -e trace=open,openat,stat,fstat,access -o trace.out <command>
-f (follow forks) matters for anything that forks or execs a child before the actual open() happens,
shell wrappers and many daemons do exactly this, without -f you'd only see the parent's syscalls and miss
the real one entirely. -e trace=<syscalls> filters the noise so you're reading the relevant handful of
lines instead of thousands of unrelated syscalls. -o writes to a file rather than interleaving with the
program's own output, keeping the trace clean to search afterward.
Capturing output across forks: add -ff alongside -f to write one trace file per process
(trace.out.<pid>) instead of interleaving every forked process into one file, useful when many children
are involved and you want to isolate one specifically. -tt adds microsecond timestamps to every line, so
you can correlate exact syscall timing against application-log timestamps.
What to look for: a line such as openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied). The arguments show the EXACT path actually resolved (catching, for example, a
symlink resolving somewhere unexpected, or a relative path resolving against a different working directory
than assumed), and the return value plus errno pair is the ground truth for what the kernel rejected. If the
trace instead shows the open succeeding but a LATER read or write failing, that points away from basic
file permissions entirely (toward something like a quota, a stale NFS handle, or a security-module check
that gates specific operations differently from the open itself), so read the whole relevant sequence, not
just the first hit.
Attaching to an already-running process instead of launching fresh: strace -f -p <pid> -e trace=open,openat,access, for a failure buried deep inside an already-running long-lived service rather
than at simple startup.
Worked example
A non-root user's Python process attempts open("/etc/myapp/config.conf") on a file chmod'd to 000 (no
permission bits for anyone). The trace shows:
openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied)
and Python's own traceback confirms PermissionError: [Errno 13] Permission denied, the same EACCES
(errno 13) the kernel returned, simply re-surfaced through the language runtime. This trace line is what
actually PROVES the cause is the Unix permission bits themselves, rather than, for example, the path not
existing at all (which would show ENOENT instead) or a mandatory-access-control denial such as SELinux
(which produces the identical EACCES at this syscall level, and only the audit log, not strace,
distinguishes the two).
Trade-offs and pitfalls
strace adds real overhead, each traced syscall round-trips through the kernel's ptrace mechanism (the
same facility strace itself is built on: the kernel stops the traced process at every syscall entry and
exit so strace can inspect and log it before letting it continue), so attaching it broadly (especially
-f across many threads or children) to a busy, latency-sensitive production process can itself cause
visible latency or timeouts. Scope the syscall filter as tightly as possible and detach as soon as the one
relevant event has been captured, rather than leaving a broad trace running. A genuinely intermittent
permission failure (works most of the time, fails occasionally) is much better caught by attaching with
-p and a narrow filter left running briefly than by trying to reproduce it fresh each time, since a
"works on retry" bug is often a race (a config file briefly holding the wrong permissions during a deploy,
for example) that a single fresh invocation won't reliably hit.
Design a permissions model for /srv/projects where multiple teams need write access to their own project directories but must not access other teams' data. Use POSIX ACLs and default ACLs to show how you would set this up and how to ensure new files inherit desired ACLs.
Sample Answer
Direct answer
Give each team's directory to a dedicated group, set the directory itself to deny "other" access entirely, add a default ACL (Access Control List, the extended per-user/per-group permission entries beyond the classic owner/group/other bits) on each team directory so new files and subdirectories automatically inherit the team's group permissions, and never rely on umask alone to keep teams apart, since umask is a per-process default, not an enforced boundary.
Design
- Base directory, traversal only.
/srv/projectsitself getschmod 751: owner (root) full access, group read+execute (list and traverse), other execute only (traverse but not list). This means nobody canls /srv/projectsand enumerate every team's name unless they already know it, without blocking legitimate traversal into a subdirectory they do have rights to. - One group per team, one directory per team, setgid plus a matching group owner. For team A:
groupadd teamA,mkdir /srv/projects/teamA,chown root:teamA /srv/projects/teamA,chmod 2770 /srv/projects/teamA. The2sets the setgid bit on the directory, so every file and subdirectory created inside automatically inherits the groupteamAregardless of the creating user's primary group;770gives the owner and group full access and gives other nothing. - Default ACL for the group, so inheritance is explicit and covers ACL entries, not just the base group bit.
setfacl -d -m g:teamA:rwx /srv/projects/teamAandsetfacl -d -m o::--- /srv/projects/teamA. The setgid bit alone already gets you group inheritance for the classic Unix group bit; the default ACL matters once you need more than one principal to inherit access, for example a cross team auditor group that should get read only access to every team's directory without being the owning group. - Repeat per team, each with its own group, directory, and default ACL. Isolation between teamA and teamB comes entirely from teamB never being named in teamA's ACL or group, combined with
other::---at both the base directory (traversal denied without a matching entry) and inside each team directory (no residual access for anyone not explicitly granted it).
Worked example (executed, including a real cross team access attempt)
Set this up for real, created a user in each team, and actually tried to cross the boundary rather than just inspecting the ACL:
groupadd teamA; groupadd teamB
useradd -m -g teamA alice
useradd -m -g teamB bob
chmod 751 /srv/projects
mkdir /srv/projects/teamA
chown root:teamA /srv/projects/teamA
chmod 2770 /srv/projects/teamA
setfacl -d -m g:teamA:rwx /srv/projects/teamA
setfacl -d -m o::--- /srv/projects/teamA
Resulting ACL (getfacl):
user::rwx
group::rwx
other::---
default:user::rwx
default:group::rwx
default:group:teamA:rwx
default:mask::rwx
default:other::---
Alice (teamA) creates a file, and it correctly inherits the team's group:
-rw-rw----+ 1 alice teamA 0 Sep 24 00:24 /srv/projects/teamA/notes.txt
Bob (teamB, not a member of teamA) attempting to touch teamA's directory at all:
$ ls /srv/projects/teamA
ls: cannot open directory '/srv/projects/teamA': Permission denied
$ cat /srv/projects/teamA/notes.txt
cat: /srv/projects/teamA/notes.txt: Permission denied
Both denied, exactly as designed. Alice, meanwhile, reads and writes her own team's file without issue.
To show the actual value a default ACL adds beyond the setgid bit alone, I added a third, cross-cutting audit group with read only access via ACL, not group ownership:
groupadd audit; useradd -m -g audit carol
setfacl -m g:audit:r-x /srv/projects/teamA
setfacl -d -m g:audit:r-x /srv/projects/teamA
A brand new file created afterward by alice automatically picked up the audit entry (getfacl on the new file showed group:audit:r-x #effective:r-- without anyone touching that file's permissions by hand), and carol, who is in audit but not teamA, could read it (cat succeeded) but not write it (echo >> ... was denied). That is what the default ACL buys you that a plain setgid directory cannot: inherited access for a principal other than the one owning group, applied automatically to every new file without a script running behind every touch.
Trade-offs and pitfalls
The most common failure mode is setting the default ACL but leaving the base directory's own access ACL open to other, so the inheritance rule is correct but the door was never actually closed; always set other::--- on the directory itself, not only in the default entries. Second, ACLs are invisible to tools that only look at the classic nine permission bits (ls -l shows a + suffix as the only hint an ACL exists at all), so document that getfacl is the source of truth here, not ls -l, or a future admin will "simplify" the permissions and quietly remove someone's access. Third, the ACL mask (the entry that caps the maximum effective permission for every named user and group entry) can silently reduce a permission you thought you granted; always check the #effective: annotation getfacl prints when the mask is restricting something, not just the nominal entry.
You are responsible for hardening SSH across a fleet. Provide an actionable plan covering /etc/ssh/sshd_config changes, centralized key or certificate management, rate limiting/brute force protection, use of bastion hosts, monitoring and alerting for suspicious access, and safe rollout steps to avoid locking out administrators.
Sample Answer
Direct answer
Hardening SSH fleet-wide means treating it as a single, deliberate rollout, not a checklist you paste into every host and hope for the best: lock down sshd_config to remove the weak defaults, move key management off individual hosts and onto something centralized and rotatable, add rate limiting so brute-force attempts can't run unbounded, funnel access through bastion hosts so there's one narrow, monitored front door instead of many, alert on suspicious access patterns, and roll every change out in a way that guarantees you can't lock yourself out of a fleet you can no longer reach.
Structured elaboration
sshd_config changes (at least eight concrete ones worth making):
PermitRootLogin no(no direct root login over SSH at all; usesudoafter logging in as a named user).PasswordAuthentication no(key-based or certificate-based auth only; passwords are brute-forceable, keys practically aren't).PubkeyAuthentication yes(explicit, though it's the default on modern OpenSSH; being explicit avoids surprises across distro defaults).AllowUsersorAllowGroupsrestricting SSH to a named allowlist of accounts or groups, so a stray local account with a weak password (if password auth were ever accidentally re-enabled) still can't SSH in.MaxAuthTries 3(caps how many authentication attempts a single connection gets before the server drops it, slowing down automated guessing).ClientAliveInterval 300paired withClientAliveCountMax 2(server sends a keepalive probe every 5 minutes and drops the session after 2 missed responses, so a stale or hijacked idle session doesn't linger indefinitely).LoginGraceTime 30(a connection that doesn't complete authentication within 30 seconds is dropped, limiting how long a slow-drip attack can hold a connection slot open).X11Forwarding noandAllowTcpForwarding nounless a specific host genuinely needs them, since both expand what a compromised SSH session can be used for.- Restricting
Ciphers,MACs(Message Authentication Codes, which verify that each packet
wasn't tampered with in transit), andKexAlgorithms(Key Exchange algorithms, which negotiate
the shared session key when a connection is first established) to modern, non-deprecated options
only (removes legacy algorithms a downgrade attack could target).
Changing the default port (22) is a genuine trade-off, not a free win: it cuts down the sheer volume of automated scanning noise hitting your logs, which has real operational value (smaller haystack for real signal), but it is security through obscurity, not a security control on its own: a targeted attacker doing a full port scan finds it in seconds. Do it for noise reduction if you want it, never present it as a substitute for the actual auth hardening above.
Centralized key or certificate management: distributing and rotating individual authorized_keys files across a fleet by hand doesn't scale and makes revocation slow (you have to touch every host). An SSH certificate authority (issuing short-lived signed certificates instead of long-lived static keys) lets you revoke access by simply not renewing a certificate rather than hunting down and removing a key from every host it was copied to, and it gives you a single place to audit who was issued access and when.
Rate limiting and brute-force protection: something watching auth logs and temporarily banning source IPs after repeated failures (fail2ban or an equivalent) adds a layer beyond MaxAuthTries, which only limits attempts within a single connection, not across many connection attempts from the same source.
Bastion hosts: route all SSH access through a small number of hardened, heavily monitored jump hosts rather than exposing every fleet member directly to the internet or even to the whole internal network; this shrinks the actual attack surface to those few hosts and gives you one place to centralize logging and session recording.
Monitoring and alerting: ship sshd auth logs (successful and failed logins, key fingerprints used) to centralized logging, and alert on patterns like repeated failures from one source, logins at unusual hours for that account, or logins from a new geography for a user who normally connects from one office.
Safe rollout, so you don't lock out your own administrators:
- Never deploy
PasswordAuthentication noto a fleet until you've confirmed every administrator who needs access already has a working key deployed and tested. - Always test the config file for syntax errors before applying:
sshd -tvalidatessshd_configwithout touching the running daemon. - Reload (
systemctl reload sshd), don't restart, where possible: a reload applies new config to the existing daemon without dropping already-established connections, so if something is subtly wrong you still have your current session to fix it from. - Roll out through configuration management in waves (a canary group of a few non-critical hosts first, verified, then the rest), never as a single fleet-wide push, and keep a documented, tested out-of-band access path (console access, a cloud provider's serial console, or a bastion path outside the change) as a safety net for the rollout itself.
Worked example
For a 200-host fleet, the rollout order matters as much as the content: first, issue everyone SSH certificates from a central certificate authority and confirm every admin can authenticate with the new certificate while password auth is STILL enabled as a fallback. Second, push the full sshd_config hardening set above to a 5-host canary group via configuration management, with PasswordAuthentication left yes for one more cycle even there, and confirm certificate-based login works cleanly and sshd -t passes on every canary host. Third, once the canary group has been stable for a monitoring cycle (confirming no locked-out sessions, no unexpected auth-log errors), flip PasswordAuthentication no on the canary group specifically, verify again, then roll the complete config (hardening plus password auth disabled) to the rest of the fleet in batches, keeping the documented out-of-band console access path live and tested throughout in case any batch goes wrong.
Trade-offs & pitfalls
- Disabling password authentication before every administrator has a verified, working key is the single most common way this kind of hardening locks a team out of their own fleet; sequencing that verification first is not optional.
- A restart (versus a reload) of
sshdmid-change, on a host you're connected to only via SSH, can drop your own session before you've confirmed the new config actually works, leaving you locked out with no way back in except out-of-band access. - Renaming the SSH port reduces log noise but provides no real protection against a targeted attacker; treat it as a nice-to-have, not a control you'd cite in a security review.
Audit for setuid and setgid programs in your infrastructure. Provide a command to find them and describe steps to minimize risk (e.g., removal, replacing with capabilities, using sudo wrappers). Also describe how to test that removing a setuid binary does not break critical functionality.
Sample Answer
Direct answer
Find every setuid and setgid binary with a single find command, then for each one ask whether it
genuinely needs to run with another user or group's privileges or whether that's a historical accident, and
prefer removing the privilege-escalation mechanism entirely, via Linux capabilities (a narrower, per-binary
grant of just the one specific kernel privilege actually needed) rather than leaving a broad setuid-root
binary in place because "it's always been there".
Structured elaboration
Finding them: find / -xdev -type f \( -perm -4000 -o -perm -2000 \) -exec ls -la {} \; (permission bit
4000 is the setuid bit, meaning the binary runs with the FILE OWNER's privileges rather than the invoking
user's, most dangerous when the owner is root; 2000 is the setgid equivalent for the group). -xdev stays
within one filesystem so the scan doesn't wander into a mounted network share. Run this against a known-good
baseline image and diff future scans against it, the actionable signal is usually "what setuid binary
appeared that wasn't in the base image", not re-auditing the same handful of standard binaries (passwd,
sudo, mount) every time.
Minimizing risk, in order of preference:
- Remove it entirely if the functionality isn't used at all, the least privileged code running is always
the safest state. - Replace the privilege mechanism with a Linux capability: instead of a binary being setuid-root (full root
for its entire runtime), grant it just the one specific operation it needs, for examplesetcap cap_net_bind_service=+ep /usr/bin/myserverlets a process bind to a privileged port (below 1024) without
being setuid-root at all, dramatically narrowing what a vulnerability in that binary could actually do. - Replace ad-hoc setuid scripts or wrappers with
sudorules scoped to the exact command and arguments
needed, centralizing the grant in one auditable, loggable place (everysudoinvocation is logged by
default) instead of a standing setuid bit that grants privilege silently on every invocation with no
audit trail. - Where you genuinely cannot remove or replace it (some vendor binaries ship setuid-root by design), at
minimum restrict who can even execute it via filesystem permissions or group membership, and track it
explicitly in the baseline diff so any accidental regression (it becoming world-executable) is caught.
Testing that removal doesn't break critical functionality:
- Strip the privilege on a single non-production canary first, not fleet-wide, this is fundamentally a "did
anything actually depend on this" question best answered incrementally. - Before removing, build a list of actual callers: an auditd watch rule on the binary's path
(-a always,exit -F path=<binary> -F perm=x -k setuid_audit),strace/ltraceon a representative
workload, or grepping scripts, cron entries, and application code for direct invocations. - After changing it on the canary, run the full relevant test suite, and leave monitoring in place (the same
auditd watch) for a full realistic usage cycle, covering weekly or monthly jobs, not just daily traffic,
before declaring it safe, an infrequent-but-critical caller (a monthly batch job, an annual certificate
renewal) is exactly what a short canary window misses. - Roll out gradually with an easy revert path (the removed binary or capability set kept ready to restore)
rather than a single fleet-wide change.
Worked example
A security audit finds /usr/bin/mount and /usr/bin/umount are setuid-root, standard and expected, since
mounting filesystems is inherently privileged, but also finds a home-grown /usr/local/bin/backup-agent is
setuid-root, absent from the base image, added ad hoc by a previous engineer so a cron job running as an
unprivileged user could read files owned by other users during backups. Tracing a real backup run confirms
the ONLY privileged operation the binary actually performs is bypassing file read permission checks, so the
fix is setcap cap_dac_read_search=+ep /usr/local/bin/backup-agent, followed by removing the setuid bit
entirely and validating a full backup cycle succeeds identically on a canary host before fleet rollout.
Trade-offs and pitfalls
Capabilities are a real security improvement but not a magic fix, a capability like CAP_SYS_ADMIN (a
notoriously broad, catch-all capability that in practice grants nearly as much power as full root for many
purposes) can be almost as dangerous as setuid-root if chosen too broadly, so the audit must name the
SPECIFIC capability needed, not just swap "setuid" for "some capability" reflexively. sudo wrappers shift
privilege to invocation time and rely on the sudoers rule being scoped tightly, an overly broad NOPASSWD: ALL entry recreates the exact same risk as a setuid-root binary, just relocated into a config file instead
of a permission bit.
Unlock Full Question Bank
Get access to all 17 Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.