Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
Demonstrate how to use journalctl to: 1) show logs for the previous boot, 2) follow logs in real-time for the unit nginx.service, and 3) filter messages by priority and timestamp. Explain the difference between persistent and volatile journals and how to enable persistent journaling.
Sample Answer
Direct answer
journalctl is the query tool for the systemd journal (systemd-journald, the logging service that replaced plain text /var/log files with an indexed, structured binary log on most current distributions). Use -b to select a boot, -f -u <unit> to follow one service's logs live, and -p/--since/--until to filter by priority and time. Whether that data survives a reboot depends entirely on whether persistent journaling is enabled, and on a fresh install of several distributions it is not.
The three asks
1. Logs for the previous boot. journalctl -b -1 shows the boot immediately before the current one (-b -2 the one before that, and so on); journalctl --list-boots shows every boot the journal actually has recorded, each with an index and a time range, which is the safer first step since -1 only works if persistent journaling was already on during that earlier boot.
2. Follow logs live for one unit. journalctl -u nginx.service -f: -u filters to a single systemd unit, -f behaves like tail -f, streaming new entries as they arrive. Add -n 50 to also show the last 50 existing lines before it starts following, since -f alone starts from "now" with no history.
3. Filter by priority and timestamp. journalctl -p err --since "2026-09-01 00:00:00" --until "2026-09-01 06:00:00": -p accepts either a numeric syslog priority (0 emerg through 7 debug) or its name (err, warning, crit, and so on), and by default means "this priority or more severe," so -p err includes crit, alert, and emerg too, not only exact matches. --since/--until accept both absolute timestamps and relative ones like --since "1 hour ago".
Persistent versus volatile journals
By default on many distributions the journal lives only in /run/log/journal, which is a tmpfs (RAM backed filesystem) and is wiped on every reboot: that is the "volatile" journal. A "persistent" journal lives in /var/log/journal on disk and survives reboots, which is what makes journalctl -b -1 actually useful days after the fact instead of only immediately after a crash.
To enable it: mkdir -p /etc/systemd/journald.conf.d, create a drop in with Storage=persistent under [Journal], then restart the journal service (systemctl restart systemd-journald). Journald also switches to persistent mode on its own the moment /var/log/journal exists, even without touching the config: mkdir -p /var/log/journal followed by systemd-tmpfiles --create --prefix /var/log/journal (the real systemd utility that applies tmpfiles.d-style directory and permission rules; there is no separate systemd-tmpfsutils binary) and a restart is enough, because journald detects the directory's existence and switches to persistent mode without a config change at all. The config drop in is for making that explicit and permanent regardless of whether the directory happens to exist. Confirm which mode is currently active with journalctl --disk-usage, which reports 0 or near it when everything is still volatile and about to disappear on reboot.
Worked example
# previous boot's logs
journalctl -b -1
# see what boots the journal actually has on disk
journalctl --list-boots
# follow nginx live, with 50 lines of history first
journalctl -u nginx.service -n 50 -f
# everything at error severity or worse in a specific window
journalctl -p err --since "2026-09-01 00:00:00" --until "2026-09-01 06:00:00"
# enable persistent journaling
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal
systemctl restart systemd-journald
journalctl --disk-usage
I ran the journalctl --version and flag inventory (journalctl --help) against a real systemd 255 install to confirm the exact flag names above (-b, -u, -f, -p, -g, --since, --until, -n/--lines) rather than trusting memory, since a wrong flag name here is exactly the kind of small error that looks plausible and is not. This sandbox has no running systemd-journald (containers do not run systemd as PID 1, the very first process a Linux system starts at boot and normally responsible for bringing up and supervising every other process, here), so journalctl -k correctly reports "No journal files were found" rather than showing fabricated log lines: any example log content in an answer like this should be labeled illustrative, because inventing plausible looking denial or error messages as if captured would be worse than saying nothing.
Trade-offs and pitfalls
Forgetting to enable persistent journaling is the single most common mistake here: a host that only ever logs to the volatile, RAM backed journal loses everything the moment it reboots, including the exact evidence you need to diagnose why it rebooted. A second mistake is treating -p err as "only errors" and being confused when crit and emerg lines show up too; the priority filter is a floor of severity, not an exact match, by design. Persistent journaling does have a real cost worth naming: unbounded disk growth if you never set SystemMaxUse= in journald.conf, so enabling persistence should come with an explicit size cap, not just flipping the storage mode on and walking away.
Explain Unix file permissions for files and directories. Describe the meaning of rwx for user, group, and others, explain numeric octal notation such as 0755 and 0640, and show how chmod u+x differs from chmod 755 with an example scenario for each.
Sample Answer
Direct answer
Every file and directory carries three permission triads (owner, group, others), each holding read (r), write (w), and execute (x). Octal notation packs each triad into one digit (r=4, w=2, x=1, summed), so 0755 means owner has read/write/execute, group and others have read/execute only. chmod u+x is a symbolic, additive edit that touches only the bit you name and leaves everything else on the file untouched; chmod 755 is an absolute overwrite that sets the whole permission set to exactly that value, wiping out anything else that was there before.
Structured elaboration
What rwx means, and it means something different on a directory than on a file:
| Bit | On a file | On a directory |
|---|---|---|
| r | read the file's contents | list the directory's entries (ls) |
| w | modify or truncate the file's contents | create, rename, or delete entries inside it |
| x | execute the file as a program or script | cd into it, or access anything inside by exact path (needed even to cat a file inside it) |
The three triads always appear in the same order: user (owner), group, others. rwxr-x--- reads left to right as owner=rwx, group=r-x, others=(nothing).
Octal notation is just base-8 shorthand for the same three triads:
- 4 = read, 2 = write, 1 = execute; add the ones you want per triad.
0755= owner 7 (4+2+1=rwx), group 5 (4+0+1=r-x), others 5 (r-x). Typical for an executable script others should be able to run but not modify.0640= owner 6 (rw-), group 4 (r--), others 0 (---). Typical for a config file with secrets: the owning service can edit it, its group can read it, nobody else can touch it.- The leading
0is the optional special-bits digit (setuid/setgid/sticky); omitting it is the same as0.
Symbolic (u+x) versus absolute (755):
- Symbolic operators are
u/g/o/a(who) combined with+/-/=(add, remove, set) andr/w/x(what).chmod u+xadds execute for the owner ONLY, and only if you ask for+; it never resets anything else. - Numeric mode always specifies the complete permission set for all three triads at once.
chmod 755 filesets that file to exactlyrwxr-xr-xno matter what it was before, silently dropping anything wider (like a group-write bit someone had intentionally set) or narrower.
Worked example
Scenario for chmod u+x (targeted, additive): a deploy script owned alice:devops is checked out fresh from git and lands as rw-r--r-- (0644): readable by everyone, executable by no one. You only want alice (the owner) to be able to run it directly, and nothing else about it should change. chmod u+x deploy.sh adds exactly the owner-execute bit, giving rwxr--r-- (0744); the group and others triads are untouched, still r-- and r--.
Scenario for chmod 755 (absolute, replace-the-whole-thing): that same config directory has drifted over time, say app.conf owned alice:devops is currently rwxrwxrwx (0777) because someone ran a careless chmod -R 777 while debugging a permission error and never reverted it. Running chmod 755 app.conf does not "add" anything, it rewrites the entire permission set in one step to exactly rwxr-xr-x: owner keeps full access, group and others are knocked down to read/execute with write removed. That closes the world-writable hole immediately, but it is a blunt instrument: if that file legitimately needed group-write for a shared deploy group, 755 silently removes it too, because absolute mode has no memory of what was there before and no way to say "just fix the others bit."
Trade-offs & pitfalls
- Reach for symbolic mode (
u+x,g-w) when you want a targeted, reversible edit and don't want to know or restate the rest of the permission bits; reach for octal when you want a known, auditable end state (config management tools like Ansible'smode: '0640'prefer octal for exactly this reproducibility). - A common mistake is using
chmod 777"to fix a permission error" during debugging and never reverting it: it silently grants world write, which on a config file or script is a real security hole, not just untidiness. - Numeric mode does not preserve special bits (setuid 4000, setgid 2000, sticky 1000) unless you include that leading digit;
chmod 755 fileon a setuid binary strips setuid. - Directory execute is the bit people forget: a directory with
rbut notxlets you list filenames but not stat or open anything inside, which produces confusing "permission denied" errors on files that look perfectly permissioned themselves.
You suspect a host can't resolve external domains. Walk through how you'd isolate whether the problem is DNS configuration, the resolver itself, or something else entirely, and how you'd interpret what you find at each step.
Sample Answer
Direct answer
The fastest way to isolate whether a "can't resolve external domains" symptom is genuinely DNS, versus something else entirely (routing, a firewall, an application-level issue), is to bypass DNS altogether first: try reaching something by raw IP address. If that works, the problem is specifically in name resolution, not connectivity; if it also fails, DNS was never the actual problem, and the investigation moves to routing, firewall rules, or the application layer instead. From there, work inward through the resolver configuration, then the resolution mechanism itself.
Structured elaboration
Step 1: rule DNS in or out entirely, by removing it from the equation. A bare TCP-level check to a known-good public IP, bypassing name lookup completely, tells you immediately whether basic network connectivity out of this host works at all: nc -zv 1.1.1.1 443 (or curl -v telnet://1.1.1.1:443) opens a raw TCP connection with no DNS lookup and no TLS negotiation involved. Reaching for curl -v https://<ip> directly is a trap worth knowing about: many production services today sit behind a shared, SNI-routed IP (a CDN or reverse proxy fronting many hostnames on one address), so connecting straight to that IP over HTTPS without also presenting the right hostname produces a TLS handshake failure that looks exactly like a connectivity problem but is actually just the server rejecting an unrecognized SNI value, nothing to do with DNS, routing, or a firewall at all. If you specifically need to test HTTPS to a named service while forcing a particular IP, curl -v --resolve <domain>:443:<ip> https://<domain> is the correct form, since it forces resolution to that IP while still sending the real domain in the TLS handshake and the Host header, so a failure there is actually meaningful. If even the bare TCP check fails, DNS isn't the culprit and the investigation redirects toward routing (ip route, traceroute toward that IP), a firewall blocking outbound traffic generally, or a network path issue upstream of this host, not toward anything DNS-specific.
Step 2: if direct-IP connectivity works, confirm what resolver configuration this host is actually using. resolvectl status (on hosts running systemd-resolved, the resolver management service that ships with most modern distributions) shows the effective nameservers currently in use, which is not always the same as what's naively expected from reading /etc/resolv.conf directly, since systemd-resolved often manages that file itself and the visible contents can be a stub pointing back at the local resolver service rather than the real upstream nameserver. Older or minimal systems without systemd-resolved read /etc/resolv.conf directly as the source of truth for nameservers.
Step 3: test raw DNS resolution directly, bypassing any local caching layer. dig @<nameserver-ip> example.com queries a specific nameserver directly, letting you determine whether the configured nameserver itself is reachable and answering correctly, independent of anything cached or misconfigured locally on this host. Comparing that against dig example.com (using the host's normal configured resolution path) tells you whether the problem is upstream (the configured nameserver itself failing to answer) or local (something wrong in how this host is routing its own queries to that nameserver, such as a local caching resolver stuck or misconfigured).
Step 4: confirm the resolution order and any local overrides. /etc/nsswitch.conf's hosts: line determines the order name resolution actually checks (commonly files dns, meaning /etc/hosts is consulted before DNS at all); a stale or incorrect entry in /etc/hosts for the domain in question, or for the nameserver's own hostname if it's referenced by name rather than IP, can shadow DNS entirely and produce a misleading symptom that looks like a DNS failure but is actually a local override taking precedence. getent ahosts <domain> resolves through the full configured chain (respecting nsswitch.conf order, unlike dig, which talks to a nameserver directly and skips local overrides entirely), which is useful specifically for confirming whether the application-visible resolution path matches what dig alone would suggest.
Step 5: if a nameserver in the chain appears reachable by IP but queries to it hang or get no response, check whether traffic is actually reaching it. A firewall or network policy blocking outbound traffic on port 53 (the standard DNS port, using UDP for lookups and sometimes TCP for larger or specific queries) produces exactly this symptom: the nameserver's IP is reachable for other things, but DNS queries specifically never get an answer. tcpdump -i any port 53 while retrying a lookup shows whether the query is even leaving the host, and whether any response comes back at all, which tells you definitively whether this is a local firewall problem, a silent network drop somewhere upstream, or a nameserver that received the query and simply isn't answering.
Interpreting each step's result as you go, rather than running the whole checklist blindly: a failure at step 1 means stop, this was never a DNS problem. A pass at step 1 with a failure resolving via the configured nameserver in step 3 means the nameserver itself, or the path to it, is the problem, not this host's configuration. A mismatch between dig's direct answer and getent's full-chain answer in step 4 means a local override (most often /etc/hosts or resolver misconfiguration) is intercepting resolution before DNS is even consulted.
Worked example
A host can't reach api.example.com. nc -zv 1.1.1.1 443 (a stable, dedicated public IP, bypassing DNS entirely) succeeds immediately, confirming basic connectivity out of the host is fine and this is specifically a name-resolution problem, not a routing or firewall issue. resolvectl status shows the effective nameserver configured is 10.0.0.2, an internal resolver. dig @10.0.0.2 api.example.com times out with no response at all, while dig @8.8.8.8 api.example.com (a public nameserver, reached directly) succeeds, isolating the problem specifically to reaching or getting an answer from 10.0.0.2. tcpdump -i any port 53 and host 10.0.0.2 while retrying the failing lookup shows the query leaving the host but no response ever arriving back, which, combined with ping 10.0.0.2 succeeding fine (so the host itself is reachable for other traffic), points at either the internal resolver service being down or overloaded rather than the network path, or a firewall change that specifically blocks port 53 to that resolver while leaving general connectivity to its IP untouched. That combination of evidence (reachable by ping, unreachable specifically on port 53, works fine against an entirely different nameserver) narrows the fix to that specific resolver's own health or its port-53 firewall rule, not this host's own configuration.
Trade-offs & pitfalls
- Jumping straight to
digor resolver-config inspection without first ruling DNS in or out via a direct-IP test risks spending time deep in resolver configuration on a problem that was actually a routing or firewall issue with nothing to do with DNS at all. digtalking directly to a nameserver andgetent/application-level lookups going through the fullnsswitch.confchain can legitimately disagree; treating them as interchangeable checks misses local override problems like a stale/etc/hostsentry.- Testing raw connectivity with
curl -v https://<ip>directly (instead of a bare TCP check or--resolve) can produce a TLS handshake failure purely because of SNI-based routing on a shared IP, which is easy to misread as "so the network path is broken" when the path is actually fine and only the TLS handshake's hostname failed to match; a bare TCP check (nc -zv) avoids that false signal entirely. systemd-resolvedenvironments can make/etc/resolv.confmisleading if read naively, since it may point at a local stub resolver rather than showing the real upstream nameserver directly;resolvectl statusis the more reliable source of truth on those hosts.
Explain the difference between SIGTERM and SIGKILL. As an SRE, how would you design a service (or its wrapper) to perform graceful termination on SIGTERM? Describe what to do in the service code and in the systemd unit to ensure graceful shutdown and proper cleanup.
Sample Answer
Direct answer
SIGTERM (signal 15) is a polite, catchable request to terminate that gives a process the chance to clean
up before exiting; SIGKILL (signal 9) is an immediate, kernel-enforced termination that a process cannot
catch, block, or run any of its own code in response to. Designing for graceful shutdown means trapping
SIGTERM to stop accepting new work, finish or checkpoint what's in flight within a bounded time, then exit
cleanly, and treating SIGKILL as only the enforced backstop for a process that fails to do that in time.
In the service code
Install a handler for SIGTERM (and typically SIGINT too, so local development with Ctrl-C behaves the same
way as a real stop) that: marks the process as draining so a load balancer's health check starts failing it
and routing stops; lets in-flight work finish naturally up to a bounded grace period; explicitly closes
resources (database connections, file handles) rather than relying on process exit to reclaim them; and
exits with code 0 to signal an intentional, clean stop rather than a crash.
In the systemd unit
systemd is the service manager most modern Linux distributions use to start, stop, and
supervise long-running services; each service it manages is described by a unit, a small
configuration file like the one below. The following directives control how systemd stops this
particular unit.
KillSignal=SIGTERM(the default, but naming it explicitly documents intent) is what systemd sends
first onsystemctl stop.TimeoutStopSec=<N>should comfortably exceed the worst-case graceful-shutdown time (the longest
in-flight request or job duration, plus margin), since systemd escalates unconditionally to SIGKILL once
this elapses, silently truncating an otherwise-working shutdown that just needed a bit longer.ExecStop=lets the unit run an explicit stop command instead of relying purely on a raw signal, useful
when a service has its own correct way to request a graceful drain (many servers expose a dedicated
admin command for exactly this).
The broader signal taxonomy
- SIGHUP (signal 1, historically "hangup" when a controlling terminal closed) is commonly repurposed by
long-running daemons as a "reload configuration without restarting" signal, a convention distinct from
termination that's worth knowing when a service seems to ignore what looks like a stop request. - SIGQUIT (signal 3) behaves like SIGINT but additionally requests a core dump by default, useful for
capturing debug state from a hung process; like SIGTERM and SIGINT it is catchable and can be ignored. - SIGSTOP (signal 19) pauses a process until a later SIGCONT resumes it and, like SIGKILL, cannot be
caught, blocked, or ignored; it underlies job control and container/orchestration freeze operations.
Worked example
A REST API service: on SIGTERM it sets a draining flag its health-check endpoint reads, so the load
balancer stops sending new requests within one health-check interval (say 5 seconds); a 25-second window
then lets existing requests finish; the process then calls its HTTP server's own graceful-shutdown method
and exits 0. The systemd unit sets TimeoutStopSec=30 (5 seconds of drain visibility plus 25 seconds of
drain time), so systemd's own SIGKILL backstop should never fire under normal operation, only if something
is genuinely stuck.
Trade-offs and pitfalls
A TimeoutStopSec set too short truncates legitimate slow shutdowns into forced kills (dropped requests,
lost work); set too long, a genuinely hung process blocks deploys and reboots for that entire window. A
common wrong turn is doing expensive cleanup (flushing large buffers to a slow disk, closing thousands of
connections one at a time) synchronously inside the signal handler itself; a handler should just set a flag
or trigger a shutdown routine on the main thread, since handlers run in a restricted context where calling
most library functions or doing I/O directly can be unsafe or deadlock.
The /proc filesystem contains runtime state about processes and the kernel. For a PID you suspect of leaking resources, list which /proc files you would inspect (for example cmdline, environ, status, fd, io, limits, smaps) and explain what each file reveals and how you would use it to diagnose the problem.
Sample Answer
The diagnostic approach
/proc/<pid>/ is a virtual filesystem the kernel generates on the fly, one directory per running process, exposing that process's live state as if it were a set of ordinary files. For a process suspected of leaking resources (memory, file descriptors, or something else growing unbounded), the useful files answer different questions, so the right approach is to check several of them together rather than any single one in isolation.
What each file reveals
cmdline: the exact command line the process was launched with (arguments are null-byte separated, sotr '\0' ' ' < /proc/<pid>/cmdlineprints it readably). First step: confirm you're diagnosing the process you think you are, not a same-named process launched with different arguments.environ: the process's environment variables at launch (also null-separated). Useful when a leak is configuration-dependent, e.g. a cache size or connection-pool limit set via an env var that's larger than intended.status: a human-readable summary including the process state, memory figures (VmRSSfor resident memory actually in physical RAM,VmSizefor total virtual address space), thread count, and signal masks. This is usually the fastest first check for "is memory actually growing," by samplingVmRSSover time.fd: a directory of symlinks, one per open file descriptor, each pointing at what that descriptor references (a file, socket, pipe, or device). Counting entries (ls /proc/<pid>/fd | wc -l) over time reveals a file-descriptor leak (sockets or files opened and never closed); the symlink targets tell you what kind of resource is piling up.io: cumulative bytes and syscalls the process has read and written (rchar,wchar,read_bytes,write_bytes,syscr,syscw). Confirmed present and populated on a live process in testing; useful for spotting a process that's pathologically read- or write-heavy, e.g. re-reading the same file in a loop instead of caching it.limits: the effective resource limits (fromulimit, the shell built-in that sets limits for a process and whatever it launches; PAM; or the unit file of systemd, the init system and service manager most current Linux distributions use to start and supervise services) actually in force for this specific process, includingMax open files,Max processes, andMax resident set. Confirms whether an observed leak has room to keep growing before hitting a hard ceiling, or whether the process is already near a limit that could explain crashes rather than pure growth.smaps: a detailed, per-memory-mapping breakdown (far more granular than the singleVmRSSfigure instatus), showing how much resident memory each individual shared library, heap segment, or memory-mapped file is contributing. This is the file to reach for oncestatushas confirmed memory is growing but you need to know which mapping is responsible, e.g. distinguishing a leaking heap from an ever-growing memory-mapped cache file.
Worked example: chasing a suspected file-descriptor leak
PID=12345
watch -n 5 'ls /proc/'"$PID"'/fd | wc -l'
If the count climbs steadily rather than plateauing under steady load, list what's accumulating:
ls -l /proc/$PID/fd | awk '{print $NF}' | sort | uniq -c | sort -rn | head
Symlink targets repeating heavily (many entries pointing at socket:[...] of the same remote address, or many /var/log/app.log duplicates) point at exactly what kind of handle the code is failing to close, which turns "there's a leak" into "there's a leak in the code path that opens sockets to service X" without needing a debugger attached.
Trade-offs and pitfalls
/procis virtual and in-memory: nothing you read from it is stored on disk, and repeatedly polling it (especiallysmaps, which is expensive for the kernel to generate for a process with many mappings) at a very high frequency has a real, if usually small, CPU cost; poll on the order of seconds, not in a tight loop.VmSize(virtual memory size) is frequently far larger than actual physical memory pressure and is a poor leak indicator on its own, since it includes reserved-but-never-touched address space;VmRSS(resident set size) is the figure that tracks physical memory the process is actually using.- Reading another user's
/proc/<pid>/environor/proc/<pid>/fdrequires appropriate privilege (root, or being the same user that owns the process); attempting it as an unrelated unprivileged user returns a permission error, which is a deliberate protection against leaking another user's secrets (credentials are a common thing to find inenviron).
Unlock Full Question Bank
Get access to all Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.