Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
Describe how you would use strace to diagnose a program that fails with 'Permission denied' when opening a config file. Include the strace invocation, how to capture output across forks, how to filter to file-related syscalls, and what specific return values you'll look for in the trace.
Sample Answer
Direct answer
Run the failing program under strace, filter to the file-related syscalls, and read the specific syscall's
return value and its error number (errno), since a "Permission denied" surfacing at the application level is
really the LAST in a sequence of file-related syscalls the kernel may have rejected, and strace shows
exactly which call, on exactly which path, with exactly which errno, the actual ground truth for what the
kernel refused and why.
Structured elaboration
Basic invocation:
strace -f -e trace=open,openat,stat,fstat,access -o trace.out <command>
-f (follow forks) matters for anything that forks or execs a child before the actual open() happens,
shell wrappers and many daemons do exactly this, without -f you'd only see the parent's syscalls and miss
the real one entirely. -e trace=<syscalls> filters the noise so you're reading the relevant handful of
lines instead of thousands of unrelated syscalls. -o writes to a file rather than interleaving with the
program's own output, keeping the trace clean to search afterward.
Capturing output across forks: add -ff alongside -f to write one trace file per process
(trace.out.<pid>) instead of interleaving every forked process into one file, useful when many children
are involved and you want to isolate one specifically. -tt adds microsecond timestamps to every line, so
you can correlate exact syscall timing against application-log timestamps.
What to look for: a line such as openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied). The arguments show the EXACT path actually resolved (catching, for example, a
symlink resolving somewhere unexpected, or a relative path resolving against a different working directory
than assumed), and the return value plus errno pair is the ground truth for what the kernel rejected. If the
trace instead shows the open succeeding but a LATER read or write failing, that points away from basic
file permissions entirely (toward something like a quota, a stale NFS handle, or a security-module check
that gates specific operations differently from the open itself), so read the whole relevant sequence, not
just the first hit.
Attaching to an already-running process instead of launching fresh: strace -f -p <pid> -e trace=open,openat,access, for a failure buried deep inside an already-running long-lived service rather
than at simple startup.
Worked example
A non-root user's Python process attempts open("/etc/myapp/config.conf") on a file chmod'd to 000 (no
permission bits for anyone). The trace shows:
openat(AT_FDCWD, "/etc/myapp/config.conf", O_RDONLY|O_CLOEXEC) = -1 EACCES (Permission denied)
and Python's own traceback confirms PermissionError: [Errno 13] Permission denied, the same EACCES
(errno 13) the kernel returned, simply re-surfaced through the language runtime. This trace line is what
actually PROVES the cause is the Unix permission bits themselves, rather than, for example, the path not
existing at all (which would show ENOENT instead) or a mandatory-access-control denial such as SELinux
(which produces the identical EACCES at this syscall level, and only the audit log, not strace,
distinguishes the two).
Trade-offs and pitfalls
strace adds real overhead, each traced syscall round-trips through the kernel's ptrace mechanism (the
same facility strace itself is built on: the kernel stops the traced process at every syscall entry and
exit so strace can inspect and log it before letting it continue), so attaching it broadly (especially
-f across many threads or children) to a busy, latency-sensitive production process can itself cause
visible latency or timeouts. Scope the syscall filter as tightly as possible and detach as soon as the one
relevant event has been captured, rather than leaving a broad trace running. A genuinely intermittent
permission failure (works most of the time, fails occasionally) is much better caught by attaching with
-p and a narrow filter left running briefly than by trying to reproduce it fresh each time, since a
"works on retry" bug is often a race (a config file briefly holding the wrong permissions during a deploy,
for example) that a single fresh invocation won't reliably hit.
You suspect a host can't resolve external domains. Walk through how you'd isolate whether the problem is DNS configuration, the resolver itself, or something else entirely, and how you'd interpret what you find at each step.
Sample Answer
Direct answer
The fastest way to isolate whether a "can't resolve external domains" symptom is genuinely DNS, versus something else entirely (routing, a firewall, an application-level issue), is to bypass DNS altogether first: try reaching something by raw IP address. If that works, the problem is specifically in name resolution, not connectivity; if it also fails, DNS was never the actual problem, and the investigation moves to routing, firewall rules, or the application layer instead. From there, work inward through the resolver configuration, then the resolution mechanism itself.
Structured elaboration
Step 1: rule DNS in or out entirely, by removing it from the equation. A bare TCP-level check to a known-good public IP, bypassing name lookup completely, tells you immediately whether basic network connectivity out of this host works at all: nc -zv 1.1.1.1 443 (or curl -v telnet://1.1.1.1:443) opens a raw TCP connection with no DNS lookup and no TLS negotiation involved. Reaching for curl -v https://<ip> directly is a trap worth knowing about: many production services today sit behind a shared, SNI-routed IP (a CDN or reverse proxy fronting many hostnames on one address), so connecting straight to that IP over HTTPS without also presenting the right hostname produces a TLS handshake failure that looks exactly like a connectivity problem but is actually just the server rejecting an unrecognized SNI value, nothing to do with DNS, routing, or a firewall at all. If you specifically need to test HTTPS to a named service while forcing a particular IP, curl -v --resolve <domain>:443:<ip> https://<domain> is the correct form, since it forces resolution to that IP while still sending the real domain in the TLS handshake and the Host header, so a failure there is actually meaningful. If even the bare TCP check fails, DNS isn't the culprit and the investigation redirects toward routing (ip route, traceroute toward that IP), a firewall blocking outbound traffic generally, or a network path issue upstream of this host, not toward anything DNS-specific.
Step 2: if direct-IP connectivity works, confirm what resolver configuration this host is actually using. resolvectl status (on hosts running systemd-resolved, the resolver management service that ships with most modern distributions) shows the effective nameservers currently in use, which is not always the same as what's naively expected from reading /etc/resolv.conf directly, since systemd-resolved often manages that file itself and the visible contents can be a stub pointing back at the local resolver service rather than the real upstream nameserver. Older or minimal systems without systemd-resolved read /etc/resolv.conf directly as the source of truth for nameservers.
Step 3: test raw DNS resolution directly, bypassing any local caching layer. dig @<nameserver-ip> example.com queries a specific nameserver directly, letting you determine whether the configured nameserver itself is reachable and answering correctly, independent of anything cached or misconfigured locally on this host. Comparing that against dig example.com (using the host's normal configured resolution path) tells you whether the problem is upstream (the configured nameserver itself failing to answer) or local (something wrong in how this host is routing its own queries to that nameserver, such as a local caching resolver stuck or misconfigured).
Step 4: confirm the resolution order and any local overrides. /etc/nsswitch.conf's hosts: line determines the order name resolution actually checks (commonly files dns, meaning /etc/hosts is consulted before DNS at all); a stale or incorrect entry in /etc/hosts for the domain in question, or for the nameserver's own hostname if it's referenced by name rather than IP, can shadow DNS entirely and produce a misleading symptom that looks like a DNS failure but is actually a local override taking precedence. getent ahosts <domain> resolves through the full configured chain (respecting nsswitch.conf order, unlike dig, which talks to a nameserver directly and skips local overrides entirely), which is useful specifically for confirming whether the application-visible resolution path matches what dig alone would suggest.
Step 5: if a nameserver in the chain appears reachable by IP but queries to it hang or get no response, check whether traffic is actually reaching it. A firewall or network policy blocking outbound traffic on port 53 (the standard DNS port, using UDP for lookups and sometimes TCP for larger or specific queries) produces exactly this symptom: the nameserver's IP is reachable for other things, but DNS queries specifically never get an answer. tcpdump -i any port 53 while retrying a lookup shows whether the query is even leaving the host, and whether any response comes back at all, which tells you definitively whether this is a local firewall problem, a silent network drop somewhere upstream, or a nameserver that received the query and simply isn't answering.
Interpreting each step's result as you go, rather than running the whole checklist blindly: a failure at step 1 means stop, this was never a DNS problem. A pass at step 1 with a failure resolving via the configured nameserver in step 3 means the nameserver itself, or the path to it, is the problem, not this host's configuration. A mismatch between dig's direct answer and getent's full-chain answer in step 4 means a local override (most often /etc/hosts or resolver misconfiguration) is intercepting resolution before DNS is even consulted.
Worked example
A host can't reach api.example.com. nc -zv 1.1.1.1 443 (a stable, dedicated public IP, bypassing DNS entirely) succeeds immediately, confirming basic connectivity out of the host is fine and this is specifically a name-resolution problem, not a routing or firewall issue. resolvectl status shows the effective nameserver configured is 10.0.0.2, an internal resolver. dig @10.0.0.2 api.example.com times out with no response at all, while dig @8.8.8.8 api.example.com (a public nameserver, reached directly) succeeds, isolating the problem specifically to reaching or getting an answer from 10.0.0.2. tcpdump -i any port 53 and host 10.0.0.2 while retrying the failing lookup shows the query leaving the host but no response ever arriving back, which, combined with ping 10.0.0.2 succeeding fine (so the host itself is reachable for other traffic), points at either the internal resolver service being down or overloaded rather than the network path, or a firewall change that specifically blocks port 53 to that resolver while leaving general connectivity to its IP untouched. That combination of evidence (reachable by ping, unreachable specifically on port 53, works fine against an entirely different nameserver) narrows the fix to that specific resolver's own health or its port-53 firewall rule, not this host's own configuration.
Trade-offs & pitfalls
- Jumping straight to
digor resolver-config inspection without first ruling DNS in or out via a direct-IP test risks spending time deep in resolver configuration on a problem that was actually a routing or firewall issue with nothing to do with DNS at all. digtalking directly to a nameserver andgetent/application-level lookups going through the fullnsswitch.confchain can legitimately disagree; treating them as interchangeable checks misses local override problems like a stale/etc/hostsentry.- Testing raw connectivity with
curl -v https://<ip>directly (instead of a bare TCP check or--resolve) can produce a TLS handshake failure purely because of SNI-based routing on a shared IP, which is easy to misread as "so the network path is broken" when the path is actually fine and only the TLS handshake's hostname failed to match; a bare TCP check (nc -zv) avoids that false signal entirely. systemd-resolvedenvironments can make/etc/resolv.confmisleading if read naively, since it may point at a local stub resolver rather than showing the real upstream nameserver directly;resolvectl statusis the more reliable source of truth on those hosts.
When investigating an issue on a Linux server, which system log files would you check first and why? Mention at least three log files under /var/log and what type of information each contains.
Sample Answer
Where to look first, and why
When investigating an issue on a Linux server, the first log to check is whichever one is closest to the layer where the symptom is showing up, because logs under /var/log are organized by subsystem, not chronologically merged, and reading the wrong one wastes time. A practical default order: check the application's own log first (it usually has the most specific, actionable detail), then the system/kernel-level logs if the application log shows nothing unusual or the symptom looks infrastructure-related (a crash, an OOM kill, a hardware issue), then authentication logs if the issue involves access or a login failure.
Three (and a few more) log files under /var/log, and what each contains
/var/log/syslog(Debian/Ubuntu) or/var/log/messages(RHEL/CentOS-family): the general-purpose system log, aggregating messages from the kernel and most system daemons that haven't been routed elsewhere. This is the right first stop for "something went wrong and I don't yet know which subsystem," since it's the broadest catch-all./var/log/auth.log(Debian/Ubuntu) or/var/log/secure(RHEL/CentOS-family): authentication and authorization events: SSH logins (successful and failed),sudoinvocations,su, and PAM module activity. This is the first place to look for a suspected brute-force attempt, an unexpected login, or a permission-related failure./var/log/kern.log: kernel-ring-buffer messages specifically (a subset of what also appears in syslog on many distributions), including hardware errors, driver messages, and out-of-memory (OOM) killer activity. Relevant when a process died with no application-level explanation at all, since the OOM killer terminates processes silently from the application's own point of view./var/log/<application-name>/: most non-trivial services (web servers, databases, container runtimes) write their own dedicated logs under a subdirectory here (e.g./var/log/nginx/error.log), which is usually far more specific and actionable than anything in the generic system logs.
Worked example
A web request is returning HTTP 502 errors. The investigation order: (1) /var/log/nginx/error.log (or the equivalent for whatever reverse proxy is in front), since a 502 specifically means the proxy couldn't get a valid response from the upstream, and the proxy's own log usually names which upstream and why (connection refused, timeout); (2) if the upstream application itself crashed, its own application log (or journalctl -u <service> on a systemd host, where systemd is the init system and service manager most modern Linux distributions use to start and supervise processes, and journalctl reads the systemd journal, its structured log store) for a stack trace or crash reason; (3) /var/log/syslog or /var/log/kern.log only if neither of the above shows anything, to check whether the process was killed by something outside its own control, such as the OOM killer.
Trade-offs and pitfalls
- On any modern systemd-based distribution, some of these plain-text files may be thin or entirely absent by default, because
journald(systemd's own logging system) is the primary sink and only forwards to/var/log/syslog-style files ifrsyslogor a similar bridge is installed and configured; don't assume/var/log/syslogexists or is current on every host without checking. - Log rotation (via
logrotate) means the file you're looking at may only cover a recent window; older entries live in numbered or compressed siblings (syslog.1,syslog.2.gz), andzgrep/zcatare needed to search the compressed ones without decompressing them by hand first. - Different distributions use different filenames for conceptually the same log (
auth.logvssecure,syslogvsmessages); scripts and runbooks written against one distribution family will silently find nothing useful (not an error, just an empty or missing file) if pointed at the other family without adjustment.
Explain Linux routing tables (main, local, and others), reverse path filtering (rp_filter), and how policy-based routing (ip rule) interacts with connection tracking and netfilter. Provide an example ip rule plus ip route that routes traffic from a specific source via a dedicated table and note potential pitfalls.
Sample Answer
Direct answer
Linux keeps more than one routing table (main, local, and any number of custom numbered tables), and ip rule decides which table to actually consult for a given packet, based on criteria like source address, before the kernel ever looks at that table's routes. Reverse path filtering (rp_filter) is a separate, independent anti-spoofing check: it drops an incoming packet if the interface it arrived on is not the one the kernel would use to route a reply back to that packet's source address, which is exactly the kind of check that starts silently dropping legitimate traffic the moment you introduce policy routing with asymmetric paths.
The pieces and how they interact
localtable: automatically maintained by the kernel, holds routes for the host's own addresses and broadcast/local traffic; you essentially never edit this by hand.maintable: the table mostip routecommands operate on by default, and what a plainip route showdisplays.- Custom tables: created implicitly just by adding a route with
table <name-or-number>; they only take effect if something actually directs traffic to consult them, which is exactly whatip ruleis for. ip rule: an ordered list of rules, each matching on criteria like source address, firewall mark (fwmark, a tag a netfilter rule can attach to a packet), or incoming interface, that says "for traffic matching this, look in this table," evaluated in priority order (lower priority number is checked first). Without anip ruleentry pointing at it, a custom table is invisible to the routing decision no matter how many routes you add to it.rp_filter: a per-interface sysctl (net.ipv4.conf.<iface>.rp_filter), with three meaningful values:0(off, no check),1(strict, RFC 3704 style: the packet is only accepted if the best route back to its source uses the same interface it arrived on),2(loose: accepted as long as any route exists back to the source, through any interface). Strict mode assumes symmetric routing (traffic to and from a given source always uses the same interface); policy routing very often deliberately creates the opposite, asymmetric routing, which is precisely the setup strictrp_filterwas designed to distrust. One thing this setting is not is purely per-interface in effect: the kernel's own documentation states the check actually applied to a given interface uses "the max value from conf/{all,interface}/rp_filter", so the effective mode is whichever ofconf.all.rp_filterand that interface's own value is numerically higher, not just whatever you set on the interface in isolation. Since loose (2) is the highest of the three values, raising a single interface to loose always takes effect regardless of whatconf.allis set to; but lowering a single interface to0to turn the check off entirely does nothing ifconf.all.rp_filteris still1or2, since the higher, more restrictive value still wins for that interface. This is the actual answer to "loose versus disabling it entirely": loose is the safer per-interface fix precisely because it is not subject to being silently overridden by the global setting the way fully disabling it on one interface is.
Worked example
Routing traffic from a specific source address out through a dedicated interface and gateway, regardless of what the main table would otherwise choose:
ip route add default via 192.168.2.1 dev eth1 table 200
ip rule add from 192.168.1.100 table 200 priority 500
Reading this back with ip rule show after adding it:
0: from all lookup local
500: from 192.168.1.100 lookup 200
32766: from all lookup main
32767: from all lookup default
Traffic sourced from 192.168.1.100 gets matched by the priority-500 rule and routed via table 200's default gateway through eth1, while every other host on the box continues using the ordinary main table (priority 32766) unaffected, since rules are evaluated in ascending priority order and the first match wins.
The pitfall this setup walks straight into
If net.ipv4.conf.eth1.rp_filter is left at the strict default (1) that many distributions ship, and the return traffic for connections initiated from 192.168.1.100 happens to come back on a different interface than eth1 (a very normal outcome of this exact kind of source-based, asymmetric policy routing), the kernel will silently drop that return traffic as looking spoofed, because the best route back to that source, from the kernel's perspective, does not match the interface it actually arrived on. The fix is setting that interface's rp_filter to loose (2) rather than strict, explicitly acknowledging that asymmetric routing is intentional here, not a sign of a spoofed packet; because the kernel takes the max of conf.all.rp_filter and the interface's own value, bumping the interface to 2 reliably takes effect no matter what conf.all is set to, whereas trying to fix this by setting the interface to 0 instead would silently fail to do anything if conf.all.rp_filter is still 1.
Trade-offs and pitfalls
- A custom table with correct routes in it that nothing ever gets routed to is a very common state to end up in after adding routes but forgetting the matching
ip rule; always verify withip rule showandip route get <address>together, not just by checking that the routes exist in the table. - Rule priority collisions or an unintentionally broad
from allcustom rule inserted above thelocaltable's priority-0 rule can break local traffic on the box entirely (even to itself), which is a serious enough failure mode that testing policy routing changes on a host you only have remote access to deserves real caution, ideally with a rollback plan (a scheduledatjob to revert, for example) in case the change locks you out. - None of this is persistent by default; like plain
ip route, these rules and routes vanish on reboot unless the distribution's network configuration layer (Netplan,systemd-networkd, or a boot-time script) is set up to reapply them.
Implement (or describe well-structured pseudocode for) a robust Python program that runs as a systemd service and detects when a cloud persistent disk has been resized by the provider. The program should safely resize partitions (if present), grow LVM PVs, extend LVs, and grow ext4/XFS filesystems with idempotency, logging, error handling and a dry-run mode. Describe how you would test this tool before fleet rollout.
Sample Answer
Direct answer
Structure the tool as three clearly separated stages, each independently idempotent: detect (does the
underlying disk report more space than the partition/LVM/filesystem currently uses), act (grow only the
layers that need it, skipping any layer already at full size), and verify (confirm the filesystem now
reflects the new size). Run it as a periodic systemd (the Linux service manager) service with a dry-run mode, and treat a partition
tool's "nothing to grow" response as a normal, expected outcome, not an error.
Structured elaboration
Detection: for each block device, compare the partition's current end against the disk's actual size
(lsblk --bytes, blockdev --getsize64), and separately compare the filesystem's reported size against the
partition or logical volume it lives on. Both comparisons are read-only and safe to run on every tick of a
timer.
Acting, layer by layer, only where needed:
growpart <disk> <partition-number>if the partition doesn't already reach the disk's new end. This
tool's own exit code carries the idempotency signal: it exits 0 withCHANGEDwhen it grows the
partition, and exits 1 with aNOCHANGEmessage when the partition is already at maximum size, a
distinct, checkable outcome that the wrapper must treat as success, not failure.- If the partition sits on an LVM (Logical Volume Manager, a layer that pools one or more physical
disks or partitions, the physical volumes, into flexible logical volumes that can be resized
without repartitioning) physical volume,pvresizeto let LVM see the new space, thenlvextend -l +100%FREEon the target logical volume. resize2fs(ext4) orxfs_growfs(XFS, which must be mounted to grow, unlike ext4) to grow the
filesystem into the newly available space. Both tools already no-op cleanly (resize2fsprints "Nothing
to do!" and exits 0) when the filesystem is already at full size, so calling them unconditionally on
every tick is itself safely idempotent.
Testing before fleet rollout (the question's explicit ask):
- Unit-test the decision logic (should I act on this layer or not) with
subprocesscalls mocked out, so
the control flow is verified without touching any real block device. - Integration-test on a disposable host or container with a loop-device-backed disk: run with
--dry-run
first and confirm it reports the correct planned actions with zero actual changes, then run for real and
confirm sizes withdf/lsblkbefore and after, then run it again immediately and confirm the second run
is a true no-op, the idempotency property a timer firing every few minutes depends on. - Canary rollout to one low-risk host first, watching the journal for a full cycle and confirming the
filesystem stays healthy (a read-onlyfsck -ncheck), before a gradual, percentage-based fleet rollout
with an easy kill switch (a systemd drop-in disabling the timer). - Failure-injection tests: a
growpartNOCHANGE path, a missing or corrupt LVM physical volume, and an
unrecognized filesystem type, confirming the tool logs and exits cleanly in each case rather than
crash-looping under systemd, which is also whyRestart=on-failureshould carry a backoff, not fire
unconditionally.
Worked example
#!/usr/bin/env python3
"""disk-autogrow.py: idempotent, dry-run-capable disk/partition/filesystem grower.
Intended to run as a systemd service (oneshot, triggered by a timer) on cloud hosts
where the provider can resize the underlying disk without a reboot.
"""
import argparse
import logging
import subprocess
import sys
log = logging.getLogger("disk-autogrow")
def run(cmd, dry_run):
log.info("would run: %s", " ".join(cmd)) if dry_run else log.info("running: %s", " ".join(cmd))
if dry_run:
return 0, ""
proc = subprocess.run(cmd, capture_output=True, text=True)
return proc.returncode, (proc.stdout + proc.stderr)
def grow_partition(disk: str, partnum: int, dry_run: bool) -> bool:
"""Returns True if a change was made (or would be, in dry-run)."""
code, out = run(["growpart", disk, str(partnum)], dry_run)
if dry_run:
return True # can't know without running; caller logs intent only
if code == 0:
log.info("partition grown: %s", out.strip())
return True
if "NOCHANGE" in out:
log.info("partition already at full size, nothing to do")
return False
log.error("growpart failed unexpectedly: %s", out.strip())
raise RuntimeError(f"growpart failed on {disk}{partnum}: {out.strip()}")
def grow_filesystem(fstype: str, target: str, dry_run: bool) -> None:
cmd = ["resize2fs", target] if fstype == "ext4" else ["xfs_growfs", target]
code, out = run(cmd, dry_run)
if dry_run:
return
if code != 0:
raise RuntimeError(f"filesystem grow failed on {target}: {out.strip()}")
log.info("filesystem grow result: %s", out.strip())
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--disk", required=True)
ap.add_argument("--partnum", type=int, required=True)
ap.add_argument("--fstype", choices=["ext4", "xfs"], required=True)
ap.add_argument("--target", required=True, help="mounted filesystem path or device")
ap.add_argument("--dry-run", action="store_true")
args = ap.parse_args()
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
try:
grow_partition(args.disk, args.partnum, args.dry_run)
grow_filesystem(args.fstype, args.target, args.dry_run)
except RuntimeError as e:
log.error("aborting: %s", e)
return 1
return 0
if __name__ == "__main__":
sys.exit(main())
Matching systemd units:
# /etc/systemd/system/disk-autogrow.service
[Unit]
Description=Detect and grow a resized cloud disk's partition and filesystem
[Service]
Type=oneshot
ExecStart=/usr/local/bin/disk-autogrow.py --disk /dev/nvme1n1 --partnum 1 --fstype ext4 --target /data
# /etc/systemd/system/disk-autogrow.timer
[Unit]
Description=Run disk-autogrow.service periodically
[Timer]
OnUnitActiveSec=5min
Persistent=true
[Install]
WantedBy=timers.target
Type=oneshot tells systemd this service is expected to run once to completion and then exit,
not stay running in the background, which is exactly right for a check-and-grow script a timer
triggers repeatedly rather than a long-lived daemon. Persistent=true on the timer tells systemd
to remember the last time this timer actually fired, across reboots, and if the host was off or
asleep past a scheduled run, to fire it once immediately on boot to catch up instead of silently
skipping that cycle.
On a real cloud disk that has genuinely grown, this sequence (growpart then resize2fs) reliably reports
the change and grows the filesystem to fill it; run a second time with no further disk resize, growpart
reports NOCHANGE and resize2fs reports Nothing to do!, both exiting 0, exactly the idempotent
behavior a timer firing every five minutes needs.
Trade-offs and pitfalls
Treating growpart's exit code 1 as a generic failure (rather than checking specifically for its
NOCHANGE message) is the single most common bug in scripts like this, it turns the normal, expected "no
work to do" case into an alert-worthy failure on every single tick after the first successful grow. XFS
specifically requires the filesystem to be mounted to grow it (xfs_growfs operates on the mount point,
not the raw device), unlike resize2fs, which can run on an unmounted ext4 filesystem too, a real
per-filesystem-type difference the tool must account for rather than assuming one code path covers both.
Unlock Full Question Bank
Get access to all Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.