Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
Describe how you would use strace to diagnose a program that fails with 'Permission denied' when opening a config file. Include the strace invocation, how to capture output across forks, how to filter to file-related syscalls, and what specific return values you'll look for in the trace.
Sample Answer
Direct answer
Run the failing program under strace, filter to the file-related syscalls, and read the specific syscall's
return value and its error number (errno), since a "Permission denied" surfacing at the application level is
really the LAST in a sequence of file-related syscalls the kernel may have rejected, and strace shows
exactly which call, on exactly which path, with exactly which errno, the actual ground truth for what the
kernel refused and why.
Structured elaboration
Basic invocation:
strace -f -e trace=open,openat,stat,fstat,access -o trace.out <command>
-f (follow forks) matters for anything that forks or execs a child before the actual open() happens,
shell wrappers and many daemons do exactly this, without -f you'd only see the parent's syscalls and miss
the real one entirely. -e trace=<syscalls> filters the noise so you're reading the relevant handful of
lines instead of thousands of unrelated syscalls. -o writes to a file rather than interleaving with the
program's own output, keeping the trace clean to search afterward.
Capturing output across forks: add -ff alongside -f to write one trace file per process
(trace.out.<pid>) instead of interleaving every forked process into one file, useful when many children
are involved and you want to isolate one specifically. -tt adds microsecond timestamps to every line, so
you can correlate exact syscall timing against application-log timestamps.
What to look for: a line such as openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied). The arguments show the EXACT path actually resolved (catching, for example, a
symlink resolving somewhere unexpected, or a relative path resolving against a different working directory
than assumed), and the return value plus errno pair is the ground truth for what the kernel rejected. If the
trace instead shows the open succeeding but a LATER read or write failing, that points away from basic
file permissions entirely (toward something like a quota, a stale NFS handle, or a security-module check
that gates specific operations differently from the open itself), so read the whole relevant sequence, not
just the first hit.
Attaching to an already-running process instead of launching fresh: strace -f -p <pid> -e trace=open,openat,access, for a failure buried deep inside an already-running long-lived service rather
than at simple startup.
Worked example
A non-root user's Python process attempts open("/etc/myapp/config.conf") on a file chmod'd to 000 (no
permission bits for anyone). The trace shows:
openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied)
and Python's own traceback confirms PermissionError: [Errno 13] Permission denied, the same EACCES
(errno 13) the kernel returned, simply re-surfaced through the language runtime. This trace line is what
actually PROVES the cause is the Unix permission bits themselves, rather than, for example, the path not
existing at all (which would show ENOENT instead) or a mandatory-access-control denial such as SELinux
(which produces the identical EACCES at this syscall level, and only the audit log, not strace,
distinguishes the two).
Trade-offs and pitfalls
strace adds real overhead, each traced syscall round-trips through the kernel's ptrace mechanism (the
same facility strace itself is built on: the kernel stops the traced process at every syscall entry and
exit so strace can inspect and log it before letting it continue), so attaching it broadly (especially
-f across many threads or children) to a busy, latency-sensitive production process can itself cause
visible latency or timeouts. Scope the syscall filter as tightly as possible and detach as soon as the one
relevant event has been captured, rather than leaving a broad trace running. A genuinely intermittent
permission failure (works most of the time, fails occasionally) is much better caught by attaching with
-p and a narrow filter left running briefly than by trying to reproduce it fresh each time, since a
"works on retry" bug is often a race (a config file briefly holding the wrong permissions during a deploy,
for example) that a single fresh invocation won't reliably hit.
Explain basic Linux firewall concepts: netfilter hooks, the iptables chains INPUT/OUTPUT/FORWARD, default policies, and how firewalld maps zones to rules. Show commands to list current rules and how to temporarily allow a port for debugging.
Sample Answer
Direct answer
Netfilter is the packet-filtering framework built into the Linux kernel; it defines five fixed points (hooks) a packet passes through on its way in, out, or through the box, and both iptables and firewalld are just user-space tools for placing rules at those hooks, not separate filtering engines of their own. iptables's filter table exposes three of those hooks as chains you actually write rules against day to day: INPUT (packets destined for this host), OUTPUT (packets originating from this host), and FORWARD (packets passing through this host to somewhere else, relevant on a router or NAT gateway). firewalld sits a layer above either iptables or the newer nftables backend and lets you think in terms of named zones (trust levels assigned to interfaces or sources) instead of raw chains and rules directly.
The hooks, the chains, and how firewalld maps onto them
flowchart TD
A[Packet arrives on NIC] --> B[PREROUTING hook]
B --> C{For this host,<br/>or passing through?}
C -->|for this host| D[INPUT hook]
D --> E[Local process]
C -->|passing through| F[FORWARD hook]
E -->|reply sent| G[OUTPUT hook]
F --> H[POSTROUTING hook]
G --> H
H --> I[Packet leaves NIC]
- The kernel's netfilter framework defines all five hooks shown above; the
iptablesfiltertable (the default table, used for allow/deny decisions) only attaches toINPUT,OUTPUT, andFORWARD. Thenattable attaches toPREROUTINGandPOSTROUTINGinstead (DNAT rewrites the destination inPREROUTING, before the routing decision; SNAT andMASQUERADErewrite the source inPOSTROUTING, after it;natalso has anOUTPUTchain for packets the host itself generates), which is why NAT and filtering rules live in conceptually different tables even though they ride the same underlying hooks. - Each chain has a default policy (
ACCEPT,DROP, orREJECT) applied when no rule in the chain matches the packet:iptables -P INPUT DROPsets theINPUTchain's fallback to silently drop anything not explicitly allowed, which is the standard posture for a host you want locked down by default. firewalldmaps interfaces or source addresses to zones (public,internal,trusted,drop, and others), each zone carrying its own set of allowed services and ports; adding an interface to thetrustedzone effectively opens everything on it, whiledropsilently discards everything with no reply at all. Underneath,firewalldtranslates zone and service configuration into actualnftablesrules (the default backend on current Fedora/RHEL) oriptablesrules on older systems; you are still ultimately configuring the same netfilter hooks, just through a higher-level vocabulary aimed at "which network is this interface on" rather than "which chain and rule."
Listing rules and temporarily allowing a port for debugging
Listing current rules:
iptables -L -n -v
(-n avoids slow DNS lookups on every address, -v adds packet and byte counters per rule, useful for confirming a rule is actually being hit at all.)
With firewalld:
firewall-cmd --list-all
shows the active zone's services, ports, and any rich rules in one summary.
Temporarily allowing a port for debugging (not surviving a reboot or firewall reload):
iptables -I INPUT -p tcp --dport 8080 -j ACCEPT
(-I inserts at the top of the chain, so it is evaluated before any earlier DROP rule that would otherwise match first.)
firewall-cmd --add-port=8080/tcp
(without --permanent, this only affects the current runtime configuration and reverts on the next firewall-cmd --reload or service restart, which is exactly the point for a debugging-only change you do not want to accidentally leave in place.)
Trade-offs and pitfalls
- A rule added with plain
iptables -I/-A(without also saving it through the distribution's persistence mechanism, such asiptables-savepiped into a startup script, ornetfilter-persistent) is gone on the next reboot; this is a feature when you are debugging, and a trap if you meant it to be permanent and forgot the second step. - Inserting with
-I(top of chain) versus appending with-A(bottom of chain) matters a great deal on a chain that already has aDROPrule partway through: append a newACCEPTrule after an existingDROPfor the same traffic and it will never be reached, sinceiptablesevaluates rules in order and stops at the first match. - Mixing manual
iptablesrules withfirewalldrunning on the same host is a common source of confusing, hard-to-reproduce behavior, sincefirewalldperiodically reconciles the actual ruleset against its own configuration and can silently remove rules it does not recognize as its own; pick one management layer for a given host and stay with it.
Audit for setuid and setgid programs in your infrastructure. Provide a command to find them and describe steps to minimize risk (e.g., removal, replacing with capabilities, using sudo wrappers). Also describe how to test that removing a setuid binary does not break critical functionality.
Sample Answer
Direct answer
Find every setuid and setgid binary with a single find command, then for each one ask whether it
genuinely needs to run with another user or group's privileges or whether that's a historical accident, and
prefer removing the privilege-escalation mechanism entirely, via Linux capabilities (a narrower, per-binary
grant of just the one specific kernel privilege actually needed) rather than leaving a broad setuid-root
binary in place because "it's always been there".
Structured elaboration
Finding them: find / -xdev -type f \( -perm -4000 -o -perm -2000 \) -exec ls -la {} \; (permission bit
4000 is the setuid bit, meaning the binary runs with the FILE OWNER's privileges rather than the invoking
user's, most dangerous when the owner is root; 2000 is the setgid equivalent for the group). -xdev stays
within one filesystem so the scan doesn't wander into a mounted network share. Run this against a known-good
baseline image and diff future scans against it, the actionable signal is usually "what setuid binary
appeared that wasn't in the base image", not re-auditing the same handful of standard binaries (passwd,
sudo, mount) every time.
Minimizing risk, in order of preference:
- Remove it entirely if the functionality isn't used at all, the least privileged code running is always
the safest state. - Replace the privilege mechanism with a Linux capability: instead of a binary being setuid-root (full root
for its entire runtime), grant it just the one specific operation it needs, for examplesetcap cap_net_bind_service=+ep /usr/bin/myserverlets a process bind to a privileged port (below 1024) without
being setuid-root at all, dramatically narrowing what a vulnerability in that binary could actually do. - Replace ad-hoc setuid scripts or wrappers with
sudorules scoped to the exact command and arguments
needed, centralizing the grant in one auditable, loggable place (everysudoinvocation is logged by
default) instead of a standing setuid bit that grants privilege silently on every invocation with no
audit trail. - Where you genuinely cannot remove or replace it (some vendor binaries ship setuid-root by design), at
minimum restrict who can even execute it via filesystem permissions or group membership, and track it
explicitly in the baseline diff so any accidental regression (it becoming world-executable) is caught.
Testing that removal doesn't break critical functionality:
- Strip the privilege on a single non-production canary first, not fleet-wide, this is fundamentally a "did
anything actually depend on this" question best answered incrementally. - Before removing, build a list of actual callers: an auditd watch rule on the binary's path
(-a always,exit -F path=<binary> -F perm=x -k setuid_audit),strace/ltraceon a representative
workload, or grepping scripts, cron entries, and application code for direct invocations. - After changing it on the canary, run the full relevant test suite, and leave monitoring in place (the same
auditd watch) for a full realistic usage cycle, covering weekly or monthly jobs, not just daily traffic,
before declaring it safe, an infrequent-but-critical caller (a monthly batch job, an annual certificate
renewal) is exactly what a short canary window misses. - Roll out gradually with an easy revert path (the removed binary or capability set kept ready to restore)
rather than a single fleet-wide change.
Worked example
A security audit finds /usr/bin/mount and /usr/bin/umount are setuid-root, standard and expected, since
mounting filesystems is inherently privileged, but also finds a home-grown /usr/local/bin/backup-agent is
setuid-root, absent from the base image, added ad hoc by a previous engineer so a cron job running as an
unprivileged user could read files owned by other users during backups. Tracing a real backup run confirms
the ONLY privileged operation the binary actually performs is bypassing file read permission checks, so the
fix is setcap cap_dac_read_search=+ep /usr/local/bin/backup-agent, followed by removing the setuid bit
entirely and validating a full backup cycle succeeds identically on a canary host before fleet rollout.
Trade-offs and pitfalls
Capabilities are a real security improvement but not a magic fix, a capability like CAP_SYS_ADMIN (a
notoriously broad, catch-all capability that in practice grants nearly as much power as full root for many
purposes) can be almost as dangerous as setuid-root if chosen too broadly, so the audit must name the
SPECIFIC capability needed, not just swap "setuid" for "some capability" reflexively. sudo wrappers shift
privilege to invocation time and rely on the sudoers rule being scoped tightly, an overly broad NOPASSWD: ALL entry recreates the exact same risk as a setuid-root binary, just relocated into a config file instead
of a permission bit.
Describe how SSH public key authentication works end-to-end. Include steps to generate a key pair, install the public key on a remote host, secure the private key, and how ssh-agent and agent forwarding work. Mention common pitfalls during setup.
Sample Answer
The end-to-end flow
SSH (Secure Shell) public-key authentication relies on asymmetric cryptography: a key pair where the private key never leaves the machine it was generated on, and the public key can be shared freely because it's mathematically infeasible to derive the private key from it. Authentication works by the server challenging the client to prove possession of the private key, without the private key ever being transmitted.
-
Generate a key pair on the client machine:
bashssh-keygen -t ed25519 -C "alice@laptop"-t ed25519selects a modern elliptic-curve algorithm (shorter keys, fast, and currently considered at least as strong as RSA at commonly used key sizes; RSA with-t rsa -b 4096remains a fine, more universally compatible fallback for older systems that don't support ed25519).-Cattaches a comment (typically an identifying label, not a secret) to make it easy to tell keys apart later. This produces two files by default:~/.ssh/id_ed25519(the private key) and~/.ssh/id_ed25519.pub(the public key). -
Install the public key on the remote host, appending it to that account's authorized-keys list:
bashssh-copy-id -i ~/.ssh/id_ed25519.pub user@remote-hostssh-copy-idconnects (using whatever auth method currently works, typically a password for this one-time setup step), appends the public key's contents to~/.ssh/authorized_keyson the remote account, and fixes the permissions on~/.sshand the file if needed. Doing it by hand is equivalent:cat id_ed25519.pub | ssh user@remote-host 'mkdir -p ~/.ssh && cat >> ~/.ssh/authorized_keys'. -
Authentication itself: when the client connects, the server looks up the offered public key in
authorized_keys, and if it's listed, issues a cryptographic challenge; the client's SSH agent (or thesshclient directly) signs it with the matching private key, and the server verifies that signature against the public key it already has on file. The private key itself is never sent over the network at any point.
Securing the private key
- File permissions matter and are enforced by the SSH client itself:
~/.sshshould be700and the private key file600;sshwill refuse to use a private key file with looser permissions (a deliberate protection against an accidentally world-readable key), reporting a "bad permissions" error rather than silently using an insecure key. - Generate the key with a passphrase (
ssh-keygenprompts for one interactively) so that anyone who copies the private key file itself still can't use it without also knowing the passphrase; an unencrypted private key is a single point of failure if the disk it's stored on is ever compromised or backed up somewhere less secure. - Never copy a private key to a remote or shared host "to make things easier"; the whole point of public-key auth is that only the public half needs to travel, and a private key that exists in two places doubles the surface for it to leak.
ssh-agent and agent forwarding
Typing a passphrase on every single connection is impractical, so ssh-agent holds a decrypted copy of the private key in memory for the duration of a session (ssh-add ~/.ssh/id_ed25519 loads it once, after which subsequent ssh connections use the agent-held key transparently without re-prompting). Agent forwarding (ssh -A user@jump-host) extends this further: it lets a private key held by your agent on your laptop be used to authenticate from a second hop (a jump/bastion host) onward to a third host, without the private key itself, or even a copy of it, ever being placed on the jump host; the jump host only relays signing requests back to your agent over the forwarded socket.
Common pitfalls during setup
- Wrong permissions on
~/.sshorauthorized_keyson the remote end (should also be700/600) is one of the most common causes of "I added my key and it still asks for a password," becausesshdon the server side enforces the same strictness and silently falls back to other auth methods rather than erroring clearly. - Agent forwarding (
ssh -A) trusts the jump host's root user (and anyone else with sufficient privilege there) not to abuse the forwarded agent socket while your session is active; on a jump host you don't fully control, this is a real exposure, and a safer alternative for exactly the "hop through a bastion" use case isProxyJump(ssh -J bastion target), which tunnels the connection through the bastion without exposing your agent socket to it at all. - Mixing up which key is which after generating several:
ssh -v(or-vvvfor more detail) shows exactly which key the client is offering and why the server accepted or rejected it, which is the fastest way to debug "why isn't this working" instead of guessing.
Trade-offs and pitfalls (setup-wide)
- A passphrase-protected key held in
ssh-agentis a reasonable balance between security and usability, but an agent with a loaded key is itself a resource: if your laptop is compromised while the agent is running and unlocked, an attacker can use the loaded key (via the agent) without ever needing the passphrase, since the agent has already done the decryption. ProxyJumpis generally the safer default over agent forwarding for bastion-hop scenarios specifically because it never extends trust to the intermediate host's ability to request signatures from your key.
Describe how LUKS/dm-crypt encryption interacts with LVM and the filesystem layer. Explain the safe sequence and caveats to resize an encrypted LVM volume and its contained filesystem, how to rekey a LUKS device with minimal downtime, and emergency recovery steps if the LUKS header becomes corrupted.
Sample Answer
Direct answer
Layer LUKS (Linux Unified Key Setup, the standard on-disk format for the kernel's dm-crypt block encryption)
directly on the raw partition or disk, and put LVM (Logical Volume Manager, the layer that lets you resize,
snapshot, and combine block devices without repartitioning) on top of the single decrypted mapping, unless
different logical volumes genuinely need different keys or separate access boundaries, in which case put
LVM below and encrypt each logical volume individually and accept multiple unlock prompts. Resizing,
rekeying, and recovery all follow directly from getting that layering right.
How the layers interact
- LUKS-under-LVM (recommended default): one LUKS container wraps the raw partition; LVM's physical
volume sits on the single decrypted/dev/mapper/<name>device; one passphrase or key unlocks everything
underneath. Simpler operations, one thing to unlock at boot. - LVM-under-LUKS: LVM creates logical volumes directly on the raw partition, and each logical volume is
individually wrapped in its own LUKS container. More keys to manage and back up, but supports genuinely
separate access boundaries per volume (for example, different tenants or different sensitivity levels on
the same physical disk).
Resizing an encrypted LVM volume, bottom to top
The sequence matters strictly bottom to top, each layer only sees the space the layer below reports:
- Grow the underlying block device itself (a cloud provider's disk-resize API, or a physical disk swap).
- Grow the partition if one is in play (
growpart <device> <partition-number>). cryptsetup resize <mapped-name>to grow the LUKS mapping into the newly available space. Both header
formats support doing this online, while the mapping stays active and mounted, noluksClose/luksOpen
cycle is required for either one. The real difference is elsewhere: LUKS1's header never stores a size
at all, so resizing it is a pure device-mapper table operation that needs no passphrase, while LUKS2
persists the new size into its own authenticated metadata, socryptsetup resizeon a LUKS2 device
requires a passphrase to write that update even though the mapping itself is never torn down. (Verified
directly: growing an active, opened LUKS1 mapping withcryptsetup resizesucceeded immediately with
no passphrase prompt and no close/reopen; the equivalent LUKS2 resize on an active mapping needed a
passphrase before it would proceed, but neither one required closing the mapping.)pvresizeon the mapped device, so LVM's physical volume sees the larger space.lvextend -l +100%FREE <volume-group>/<logical-volume>to grow the logical volume into the newly
available extents.resize2fs(ext4) orxfs_growfs(XFS) to grow the filesystem to fill the logical volume.
Doingpvresizebeforecryptsetup resizeis a common, cleanly diagnosable mistake: LVM will report zero
new physical extents available, because the crypt mapping is still reporting its old size to everything
above it.
Rekeying with minimal downtime
cryptsetup luksAddKey <device> adds a new passphrase into a free keyslot (LUKS2 supports up to 32
keyslots depending on header parameters, LUKS1 exactly 8) alongside the existing one, without touching the
actual data encryption key at all, only the passphrase-wrapped copy of it changes. Verify the new
passphrase unlocks the device, then cryptsetup luksRemoveKey (or luksKillSlot) the old one. Because the
underlying master key never changes, this completes in effectively constant time regardless of volume size,
with zero downtime since the device stays mounted throughout. Contrast this with cryptsetup reencrypt
(LUKS2's online re-encryption, which genuinely rewrites the master key and re-encrypts every block), a much
heavier operation that LUKS2 supports resuming from a checkpoint if interrupted, precisely because a crash
mid-reencrypt without that resume metadata would be dangerous.
Emergency recovery from a corrupted LUKS header
The header holds every keyslot's metadata; if it's destroyed, the master key may be unreachable through any
passphrase, by design, that unrecoverability is the point of full-disk encryption, not a defect. The only
reliable recovery path is a header backup taken proactively: cryptsetup luksHeaderBackup --header-backup-file <file> <device> immediately after luksFormat and after every luksAddKey/
luksKillSlot change. If a backup exists, cryptsetup luksHeaderRestore --header-backup-file <file> <device> restores it and normal unlock resumes. Without a backup, cryptsetup repair can attempt limited
recovery of a LUKS1 header, and cryptsetup luksDump shows whatever keyslot metadata survives; if even one
valid keyslot remains, the data is recoverable, if none do, it genuinely is not.
Worked example
A database host runs LUKS directly on a partition with LVM on top, needing to grow from 500GiB (gibibyte,
2^30 bytes) to 750GiB after the cloud disk itself was resized: growpart /dev/nvme1n1 1, then
cryptsetup resize luks_data, then pvresize /dev/mapper/luks_data, then lvextend -l +100%FREE /dev/vgdata/lvdata, then resize2fs /dev/vgdata/lvdata. Skipping the cryptsetup resize step and running
pvresize directly after growpart is the concrete, checkable symptom of doing this out of order: pvresize
reports the physical volume is unchanged in size, because the crypt mapping in between still reports the old
500GiB.
Trade-offs and pitfalls
LUKS-under-LVM's single unlock is operationally simpler (one boot-time passphrase, or a network-bound
unlock via a TPM, trusted platform module, backed mechanism) but means anyone holding that one key can see
every logical volume underneath; LVM-under-LUKS supports per-volume keys at the cost of managing and backing
up several header backups instead of one. The most common mistake in "rotation" is running luksAddKey and
stopping there: an old, potentially compromised passphrase left in its keyslot after adding a new one has
not actually been rotated, only supplemented, luksRemoveKey on the old slot is the step that completes it.
Unlock Full Question Bank
Get access to all 17 Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.