Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
Describe the special permission bits: setuid, setgid and the sticky bit. Explain common uses (e.g., sudo setuid bit, shared temporary directories sticky bit), and show a find command to locate files with setuid or setgid on a system.
Sample Answer
The three special permission bits
Beyond the normal read/write/execute bits, Linux permissions carry three extra bits that change how a file or directory behaves:
| Bit | Set on | Effect | Symbolic marker | Octal prefix |
|---|---|---|---|---|
| setuid | executable files | the program runs with the file owner's UID, not the invoking user's UID | s in the owner-execute slot (-rwsr-xr-x) | 4 |
| setgid | executable files | the program runs with the file's group GID, not the invoking user's primary group | s in the group-execute slot (-rwxr-sr-x) | 2 |
| setgid | directories | new files/subdirectories created inside inherit the directory's group, not the creating user's primary group | s in the group-execute slot on a directory | 2 |
| sticky | directories | any user can create files, but only the file's owner (or root) can delete or rename files inside, even though the directory itself may be world-writable | t in the others-execute slot (drwxrwxrwt) | 1 |
Common uses
- setuid on
sudo/passwd:/usr/bin/sudoand/usr/bin/passwdcarry the setuid bit and are owned byroot, which is exactly why an ordinary user invokingpasswdcan update a system file (/etc/shadow) they otherwise have no write access to: the process briefly runs asrootfor the duration of that one operation. - setgid on shared project directories: setting setgid on a team's shared directory means every new file created inside automatically belongs to the team's group, so collaborators don't end up creating files nobody else on the team can write to.
- sticky bit on
/tmp:/tmpis world-writable (mode1777, i.e.drwxrwxrwt) so any process can drop scratch files there, but the sticky bit stops one user from deleting or renaming another user's files in that shared space, which would otherwise be possible on any ordinary world-writable directory.
Locating setuid/setgid files
find / -xdev -perm -4000 -type f # setuid files
find / -xdev -perm -2000 -type f # setgid files
-perm -4000 matches files where the setuid bit is set (regardless of the other permission bits); -xdev keeps the search from wandering into other mounted filesystems (network mounts, /proc, containers), which both speeds up the search and avoids false positives from filesystems you don't intend to audit. Verified: on a stock Ubuntu 24.04 container, find / -xdev -perm -4000 -type f 2>/dev/null returned a short, expected list including /usr/bin/chfn, /usr/bin/chsh, /usr/bin/gpasswd, /usr/bin/mount, and /usr/bin/newgrp, all legitimate setuid-root utilities that ship with the base OS.
Why a setuid-root binary is a real security risk, not just a mechanism
The chmod octal syntax (chmod 4755 file sets rwxr-xr-x plus setuid, the leading 4) is mechanically simple, but the risk is what that bit does at runtime: any bug in a setuid-root program (a buffer overflow, an unsanitized environment variable, a path a user can influence) is a bug that executes with full root privileges instead of the invoking user's own limited privileges. This is the standard privilege-escalation pattern: an attacker who can get an already-setuid-root binary to misbehave (rather than needing to run their own code as root, which they don't have permission to do) inherits root through the binary's own elevated identity. This is precisely why the list a fresh install produces is short and every entry is a well-audited, purpose-built binary (passwd, mount, su), and why finding an unexpected setuid-root binary on a host (in /tmp, in a user's home directory, or on any binary that isn't a stock OS utility) is treated as a serious incident, not routine noise: it's exactly the artifact a successful privilege-escalation attack leaves behind, and it's also why file-integrity monitoring tools baseline the setuid/setgid inventory and alert on any new addition to it.
Trade-offs and pitfalls
- Setuid has no effect on scripts (shell scripts, Python scripts) on Linux; the kernel silently ignores the setuid bit on a script's shebang line for security reasons (a classic old race-condition exploit), so
chmod 4755 myscript.shlooks like it worked but grants nothing. Setuid only works on compiled binaries. - The sticky bit and setuid/setgid share the same octal namespace (
1,2,4) and can be combined (chmod 3770sets both setgid and sticky), which is easy to get wrong if you're not deliberately adding the digits. - Removing an unexpectedly-found setuid bit (
chmod u-s file) is safe and non-destructive, but the right first step on an unexplained setuid-root binary is to investigate how it got there (compare against the package manager's manifest, e.g.dpkg -Vorrpm -V) before assuming a simple chmod fix closes the incident.
Write a one-line bash command to search under /var/log for files modified in the last 24 hours that contain the string ERROR (case-insensitive), and print the filename and matching lines. Explain the flags you used and how your command handles binary or compressed logs.
Sample Answer
Direct answer
find /var/log -type f -mtime -1 -print0 | while IFS= read -r -d '' f; do
case "$f" in
*.gz) zgrep -Hi 'error' -- "$f" ;;
*.bz2) bzgrep -Hi 'error' -- "$f" ;;
*.xz) xzgrep -Hi 'error' -- "$f" ;;
*) grep -HIi 'error' -- "$f" ;;
esac
done
find -mtime -1 selects regular files modified within the last 24 hours. Feeding paths through
-print0 / read -r -d '' (NUL-separated) rather than plain newlines is what keeps this
correct on log paths containing spaces. The case dispatches compressed files to the matching
z/bz/xz-prefixed grep variant (which decompresses on the fly) and everything else to plain
grep -HIi: -H always prints the filename even when only one file is scanned, -i makes the
match case-insensitive, and -I (capital i) tells grep to treat a binary file as if it had no
matching data at all, which is what silently skips binary files like core dumps instead of
spewing Binary file matches or garbled terminal output.
Structured elaboration
Why not a single grep -r
grep -r 'ERROR' /var/log alone does not filter by modification time, and it does not
decompress .gz/.bz2/.xz rotated logs, since compressed bytes never contain the literal
string ERROR, they contain compressed bytes. Both the recency filter and the
compression-awareness require find plus dispatch logic, not a single grep invocation.
Flag-by-flag
| Flag | Effect |
|---|---|
find ... -type f | Only regular files, skip directories and device nodes |
find ... -mtime -1 | Modified less than 1*24h ago |
find ... -print0 | NUL-separated output, safe for any filename |
IFS= | Clears the shell's field separator for this one read, so leading/trailing whitespace in a filename is kept instead of being trimmed |
read -r -d '' | Reads NUL-separated records, -r stops backslashes from being interpreted |
grep -H | Force the filename prefix on every match line |
grep -i | Case-insensitive match |
grep -I | Skip binary files instead of scanning them |
grep -- | Marks end of options, so a filename starting with - is never parsed as a flag |
Binary and compressed logs
- Compressed:
.gzfiles needzgrep,.bz2needbzgrep,.xzneedxzgrep. These are thin
wrappers that decompress to a pipe and rungrepover the stream; they accept the same-Hi
flags asgrepitself.- Binary: a rotated log directory sometimes contains a stray binary artifact (a core dump, an
old journal file).grep -Itreats such a file "as if it did not contain matching data",
meaning it is silently skipped rather than scanned byte by byte or printed to the terminal as
garbage. Without-I, plaingrepstill tries to detect binary content heuristically and
printsbinary file <name> matchesinstead of the line, which is noisy but not wrong;-I
is the explicit, reliable way to just skip such files.
- Binary: a rotated log directory sometimes contains a stray binary artifact (a core dump, an
Worked example
Verified in a container with a realistic mixed /var/log-style directory:
$ ls -la testdir
app.log # modified now, contains "An Error occurred..."
app.log.1.gz # modified now, gzip of a line containing "fatal ERROR"
nomatch.log # modified now, no match
old.log # modified 3 days ago, contains "ERROR" but too old
core.bin # modified now, raw bytes including the substring ERROR
$ find testdir -type f -mtime -1 -print0 | while IFS= read -r -d '' f; do
case "$f" in
*.gz) zgrep -Hi 'error' -- "$f" ;;
*) grep -HIi 'error' -- "$f" ;;
esac
done
testdir/app.log:2026-09-23 10:00:00 An Error occurred connecting to db
testdir/app.log.1.gz:2026-09-23 09:00:00 fatal ERROR: disk full
old.log is correctly excluded by -mtime -1 despite containing a match, nomatch.log is
correctly excluded by content, and core.bin is correctly excluded by -I even though it
literally contains the bytes ERROR, all confirmed by the exact two lines printed above and
nothing else.
Trade-offs and pitfalls
-mtime -1uses whole-day buckets under the hood; if you need a precise rolling 24-hour
window rather than "changed since the last midnight-aligned day boundary" quirk somefind
implementations have, use-newermt '24 hours ago'(GNU find) for an exact cutoff instead.- Running this as root against
/var/logwill read files owned by other services; on a
multi-tenant host, consider whether the on-call engineer running it is authorized to read
every log file that matches, some may contain sensitive data. - A directory with thousands of matching files will spawn a
zgrep/grepprocess per file
sequentially; for a genuinely large/var/log, considerxargs -0 -P4to parallelize, at the
cost of interleaved output that is harder to read live. - This one-liner does not descend into
journald's binary journal files at all; those need
journalctlwith its own time and grep-equivalent (-g) flags, a plainfind/greppass over
/var/logwill never see anything logged only to the journal.
Provide a sed one-liner to replace all occurrences of the string foo with bar in place for all .conf files under /etc, making a backup with extension .bak. Explain the sed flags you used and how to test changes before mass-apply.
Sample Answer
Direct answer
find /etc -type f -name '*.conf' -exec sed -i.bak 's/foo/bar/g' {} +
-i.bak edits each file in place while first saving the original as <file>.bak in the same
directory, so nothing is destroyed if the substitution turns out wrong. s/foo/bar/g is the
standard substitute command: replace every (g = global, not just the first) occurrence of
foo with bar on each line. find ... -exec ... {} + batches as many matched files as fit on
one command line into a single sed invocation (far cheaper than -exec ... {} \;, which forks
a new sed process per file).
Structured elaboration
Flag by flag
| Piece | Meaning |
|---|---|
find /etc -type f -name '*.conf' | Only regular files, recursively, whose name ends in .conf |
sed -i.bak | In-place edit; .bak is the backup suffix, applied to the original filename |
s/foo/bar/g | Substitute foo with bar, g for every match per line, not just the first |
-exec ... {} + | Pass matched files in batches as arguments to one sed call, not one call per file |
Why -i.bak specifically, not bare -i
Bare sed -i 's/foo/bar/g' file on GNU sed edits with no backup at all; on BSD/macOS sed, -i
requires a suffix argument (even -i '' for none) or it errors out or, worse, consumes the
next argument as the suffix by mistake. Writing -i.bak (no space) is the portable, safe form
that always leaves a recovery copy and works the same way with both sed implementations'
sane behaviors.
Testing before mass-apply
Never run the in-place version first against a directory tree you have not previewed:
- Preview matches without editing anything:
find /etc -type f -name '*.conf' -exec grep -l 'foo' {} +lists exactly which files containfooat all. - Preview the resulting text without writing it:
sed -n 's/foo/bar/gp' fileprints only the
lines that would change, post-substitution, without touching the file (-nsuppresses normal
output,pexplicitly prints matched-and-substituted lines). - Run the real in-place command only after step 1 and 2 look right, then diff a sample:
diff file.bak fileon one or two files to confirm the actual change matches what step 2 predicted. - Clean up backups once confirmed:
find /etc -name '*.conf.bak' -delete, or keep them if the
change touches anything you might need to roll back on a live service.
Worked example
Verified in a container against a small /etc-style tree:
$ cat a.conf
listen=foo
name=foobar
other=baz
$ sed -n 's/foo/bar/gp' a.conf # dry run preview, no file touched
listen=bar
name=barbar
$ find . -type f -name '*.conf' -exec sed -i.bak 's/foo/bar/g' {} +
$ cat a.conf
listen=bar
name=barbar
other=baz
$ diff a.conf.bak a.conf
1,2c1,2
< listen=foo
< name=foobar
---
> listen=bar
> name=barbar
A sibling c.txt (not matching *.conf) containing unrelated=foo was confirmed untouched,
proving the -name '*.conf' filter correctly scoped the change to configuration files only.
Trade-offs and pitfalls
s/foo/bar/gmatchesfooanywhere in a line, including inside a longer token
(myfoobar123). If the intent is a whole-word or whole-value replacement, use a
word-boundary anchor (s/\bfoo\b/bar/gwith GNU sed's\b, ors/\<foo\>/bar/gportably) or
anchor to the key (s/^foo=/bar=/) so you do not corrupt unrelated identifiers that merely
contain the substring.-i.bakwrites the backup next to the original; on a read-only or size-constrained/etc
partition (uncommon but real on some embedded or hardened images) doubling every touched file
temporarily can matter. Sweep and remove.bakfiles once verified.- A restart of the affected service is a separate step this command does not do; changing a
.conffile does not make a running daemon re-read it, plan the reload (systemctl reload <service>or equivalent) as part of the change, not as an afterthought. - Running this as a single blanket sweep across every
*.confunder/etcconflates unrelated
services; a safer production pattern scopes thefindto a specific subdirectory
(/etc/myapp/) rather than the entire/etctree, unless the change genuinely is meant to be
system-wide.
A host is showing high memory usage. Describe how you would use the free command and other tools to interpret memory and swap usage. What indicators tell you that swapping is harming performance versus being benign? Include commands and a brief interpretation guide.
Sample Answer
Direct answer
Start with free -h to get the shape of the problem (total memory, how much is genuinely used versus reclaimable cache, and whether swap has been touched at all), then use vmstat 1 to watch swap-in/swap-out activity over several seconds. A single non-zero swap value in free is not an incident. Swapping is only hurting you when the swap-in/swap-out columns in vmstat stay non-zero for sustained periods while CPU time shifts into wait-for-I/O (wa), because that means the kernel is actively paging memory that a running process needs right now, not just quietly parking cold pages.
Reading the tools
free -h columns, and the two traps in them:
usedis memory actively allocated by processes and the kernel (excluding reclaimable cache).buff/cacheis the page cache and buffer cache: memory Linux is using opportunistically to speed up disk I/O. It is reclaimable on demand, so it is not "wasted" memory even though it shows up as a large number.availableis the number that matters for "am I about to run out of memory": an estimate of how much memory could be given to a new process right now without swapping, after accounting for reclaimable cache. A host withusedat 90% butavailablestill comfortably above zero is healthy.- The
Swap:row showsused(swpd): how much has ever been paged out and not yet reclaimed. Non-zero is normal on almost every long-running Linux host.
vmstat 1 (sampling every second) adds the time dimension free lacks:
si/so(swap-in / swap-out) in KB/s: this is the number that separates benign from harmful. A one-time blip when a rarely-used background job's cold pages got swapped out is fine. Sustained non-zerosiandsoacross many samples means pages are being evicted and then immediately needed again: the kernel is thrashing.wain the CPU section: percentage of CPU time spent waiting on I/O. Risingwaalongside non-zerosi/sois the strongest signal that swapping is actually costing you latency, because every page fault that has to go to disk stalls the requesting thread.
Supporting tools:
top/htop: per-processRES(resident memory) versus a swap column, and system-wide%wa.cat /proc/meminfo:MemAvailable,Active,Inactive,Dirtyfor a finer-grained snapshot thanfreegives.sar -W 1 5(from thesysstatpackage, if installed) to log swap activity over a window, useful for building a trend rather than a single sample.ps aux --sort=-%memorsmem -tkto find which process holds the largest resident set, andgrep VmSwap /proc/<pid>/statusto find which process is actually the one sitting in swap.
Worked example
On a lightly loaded host, free -h might read:
total used free shared buff/cache available
Mem: 16Gi 841Mi 14Gi 23Mi 614Mi 14Gi
Swap: 16Gi 0B 16Gi
Here used is small, available is nearly the whole box, and swap is untouched: nothing to investigate.
Now take a scenario where vmstat 1 10 shows si around 1200 and so around 800 (KB/s) on every sample, with wa sitting at 30-35%. Do the arithmetic on so: at a 4 KiB page size, 800 KB/s of swap-out is roughly 800 / 4 = 200 pages evicted per second. If each of those pages later has to be faulted back in from swap at a conservative 2 ms of added latency (typical for a busy block device under contention, not a bare SSD spec sheet number), that is 200 x 2 ms = 400 ms of extra latency being injected into the workload every second, on top of whatever the application would normally take. That is not a rounding error: it is enough to turn a snappy service into one users describe as "randomly slow," and it recurs every second the pattern holds, unlike a one-off swap-out that never gets touched again.
Trade-offs and pitfalls
- The most common misdiagnosis is alarm over the
usedorbuff/cachenumbers infreeinstead ofavailable. A host can look "90% full" and be completely healthy. - Swap size is not a fix for a host that is genuinely undersized for its workload. Adding more swap just delays the point where things get slow; it does not remove the underlying memory pressure, and once you are thrashing, more swap space means the thrashing has more room to get worse, not that it stops.
- There is a real trade-off in removing swap entirely on a memory-constrained host: with swap present, the system degrades gracefully (slow) under pressure; with no swap, the kernel's OOM killer (out-of-memory killer, the mechanism that force-terminates a process, chosen by an
oom_scoreheuristic, when the system cannot free enough memory another way) steps in sooner and kills something outright. Neither is automatically "better": a service that would rather fail fast and restart than serve degraded latency may prefer no swap; a service where any downtime is worse than slowness wants swap as a shock absorber. - Treat this as one sample point, not a verdict: always take several
vmstatsamples over time before deciding a host is thrashing, since a single second ofsi/soactivity coinciding with, say, a cron job waking up is not the same signal as a sustained pattern.
Describe how to ensure correct system time and timezone on a fleet of Linux servers using timedatectl and chrony. Show commands to set the timezone to UTC, enable and start chronyd/chrony, and verify that the system clock is synchronized. Explain why accurate time is important for distributed systems and log correlation.
Sample Answer
Direct answer
timedatectl sets and reports the system's timezone and overall clock status; chrony (the chronyd daemon, current default on RHEL/Fedora and Ubuntu alike, having replaced the older ntpd) is what actually keeps the clock synchronized against remote time sources over NTP (Network Time Protocol). Accurate, synchronized time matters on a fleet for one concrete reason above all others: correlating log lines and trace spans across multiple hosts only means anything if "5ms apart" reflects real elapsed time rather than clock drift between the two machines being compared.
Setting the timezone
timedatectl set-timezone UTC
timedatectl
Standardizing every server's timezone to UTC specifically (rather than each host's local timezone) is standard practice precisely because it removes a whole category of correlation error: comparing timestamps across hosts in different local timezones (or the same timezone during a daylight saving transition) requires converting before comparing, which is exactly the kind of manual step that gets skipped under incident pressure and produces a wrong conclusion about ordering of events.
Enabling and starting chrony
The service name differs by distribution, confirmed against real package contents rather than assumed:
- Debian/Ubuntu: the package installs
chrony.service.systemctl enable --now chrony. - RHEL/Fedora/Rocky: the package installs
chronyd.service.systemctl enable --now chronyd.
Getting this wrong (using the Debian service name on a RHEL host or vice versa) fails immediately with "unit not found," which is a fast, obvious failure, but worth knowing ahead of time rather than discovering it live during an incident.
Verifying synchronization
timedatectl
chronyc tracking
chronyc sources -v
timedatectl's own output includes an NTP service: active (or System clock synchronized: yes) line as a quick top level check. chronyc tracking gives the detail that actually matters: System time (the current offset from the reference source), Last offset, and RMS offset (a running measure of typical error), which tell you not just "is it synced" but "how good is the sync." chronyc sources -v lists every configured time source with a symbol indicating which one chrony currently trusts as its best reference (*) versus ones it is aware of but not currently using (+, or excluded ones marked differently), useful for confirming the fleet is actually reaching its intended internal or external NTP sources rather than silently falling back to something else or free running with no sync at all.
Why accurate time matters for distributed systems and log correlation specifically
- Ordering events across hosts. A distributed trace spanning three services on three hosts is only interpretable if their clocks agree closely; a host running even a few hundred milliseconds fast can make a downstream span appear to start before the upstream request that triggered it, actively misleading whoever is debugging a slow request path.
- Certificate and token validation. TLS certificate validity windows and time bound authentication tokens (session tokens, Kerberos tickets, many OIDC/JWT based flows (OIDC: OpenID Connect, a login-federation protocol; JWT: JSON Web Token, the signed, time-stamped credential those flows issue)) are checked against the local clock; a host with a clock meaningfully wrong can reject valid credentials or accept ones that should have expired, a security relevant failure mode, not merely a cosmetic logging inconvenience.
- Log aggregation and alerting thresholds. Centralized logging (ELK, Graylog, or similar) generally timestamps by the log line's own embedded time, not arrival time; a drifted host silently skews every event it contributes into the aggregate timeline, and a rate based alert (X errors within Y minutes) can misfire or fail to fire at all if the contributing hosts disagree about what time it currently is.
Worked example (verified service naming; live sync not reproducible in this sandbox)
timedatectl set-timezone UTC
systemctl enable --now chrony # Debian/Ubuntu
systemctl enable --now chronyd # RHEL/Fedora/Rocky
chronyc tracking
chronyc sources -v
The two service names above were confirmed by inspecting the actual installed package contents rather than assumed: the Ubuntu chrony package installs chrony.service (along with chrony-wait.service and chrony-dnssrv@.service), while the Rocky Linux chrony package installs chronyd.service (along with chronyd-restricted.service). Actually reaching and synchronizing against a real time server is not something this sandboxed container environment can demonstrate live (it has no route to real NTP infrastructure and, separately, timedatectl itself requires a running systemd/dbus session (systemd being the init system and service manager most current Linux distributions use to start and supervise services) that a plain container does not provide), so no fabricated chronyc tracking numbers are presented as if captured; the service names and command syntax above are real and verified, and any specific offset or stratum value from a live host would need to come from an actual running fleet, not be invented here to look convincing.
Trade-offs and pitfalls
The most common mistake is assuming synchronization is fine because the service is simply running, without ever checking chronyc tracking's actual offset: a service can be active while still meaningfully out of sync during its first few minutes after boot, or if every configured source is unreachable due to a firewall rule blocking outbound UDP 123. Second, mixing timezones across a fleet (some hosts local time, some UTC) is a self-inflicted correlation problem that costs nothing to avoid up front and a great deal of debugging time to untangle after the fact, once years of logs exist in inconsistent zones. Third, do not assume the RHEL service name works on Debian or vice versa when writing fleet wide automation; branch on distribution family explicitly rather than hardcoding one service name everywhere.
Unlock Full Question Bank
Get access to all 41 Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.