Linux System Administration Questions
Operating and maintaining Linux and Unix systems: filesystems and permissions, process and service management, package management and updates, shell and scripting, storage, and network configuration on the host. Covers day-to-day administration, troubleshooting, monitoring and tuning host resource usage (CPU, memory, swap, and I/O), and hardening of Linux servers that underpin most infrastructure. The Linux operator's core skill set.
Write a Bash function that creates a secure temporary directory, sets permissions so only the owner can access it, and registers a trap to remove it on EXIT. Provide the function and an example of how to use it in a script safely handling spaces in filenames.
Sample Answer
Direct answer
create_secure_tmpdir() {
local prefix="${1:-app}"
SECURE_TMPDIR="$(mktemp -d "${TMPDIR:-/tmp}/${prefix}.XXXXXXXX")"
chmod 700 "$SECURE_TMPDIR"
}
cleanup() {
local status=$?
[[ -n "${SECURE_TMPDIR:-}" && -d "$SECURE_TMPDIR" ]] && rm -rf -- "$SECURE_TMPDIR"
exit "$status"
}
main() {
create_secure_tmpdir "myapp"
trap cleanup EXIT
local target="$SECURE_TMPDIR/a file with spaces.txt"
printf 'hello world\n' > "$target"
}
main "$@"
mktemp -d creates the directory atomically with a random suffix (no race between "check if a
name is free" and "create it", which is what makes it secure: a predictable name in /tmp
that another user could pre-create or symlink is a classic local privilege-escalation vector).
chmod 700 then restricts it to the owner. The trap cleanup EXIT removes it whenever the
script exits, on success, on error, or on a signal, so a crash mid-script never leaves a
world-creatable temp directory lying around.
Structured elaboration
Why the function sets a global, not a local returned via command substitution
The tempting first draft calls mktemp, chmod, and trap all inside one function and then
captures its output with workdir="$(create_secure_tmpdir)". That is broken, and the breakage
is subtle enough to ship: $(...) runs the function in a subshell. The subshell exits the
instant the function returns (right after it prints the path), which fires the trap ... EXIT
that was set inside that subshell, deleting the directory before the parent script ever
receives the path back. The fix is to never route the path through command substitution:
create_secure_tmpdir sets a global variable directly, and trap is registered afterward in
the same shell that will eventually exit for real (the top-level script), not in a throwaway
subshell.
Handling spaces in filenames correctly
- Quote every expansion of the path:
"$SECURE_TMPDIR","$target". An unquoted$SECURE_TMPDIR
undergoes word splitting and would be treated as multiple arguments if the prefix orTMPDIR
contains spaces. - If you later need to iterate over files inside the directory, never do
for f in $(ls "$dir"); usefind "$dir" -print0 | while IFS= read -r -d '' f; do ... done, which uses a
NUL byte as the record separator so spaces (and even newlines) inside filenames cannot be
mistaken for separators. - Always terminate option lists with
--before a path argument (rm -rf -- "$dir") so a
filename that happens to start with-is never parsed as a flag.
Permissions
700 (rwx------) means only the owning user can read, write, or traverse into the directory;
group and others get nothing. mktemp -d on Linux already creates directories 700 by default,
so the explicit chmod 700 here is defense against a non-default umask (the mask that strips
permission bits off every newly created file or directory before any explicit chmod runs;
e.g. a umask of 022 would otherwise leave a fresh directory group- and world-readable) on some other
distribution or in some other user's shell, and it documents the intent instead of relying on an
implicit default.
Worked example
Run against the corrected function on a live container (Ubuntu 24.04):
$ bash myscript.sh
$ ls /tmp
(directory is already gone: cleanup fired on script exit)
Inside main, before exit, the directory and its file are confirmed present and correctly
permissioned:
$ stat -c '%a' "$SECURE_TMPDIR"
700
$ cat "$SECURE_TMPDIR/a file with spaces.txt"
hello world
By contrast, the broken "function sets its own trap and returns the path via $(...)" version,
run in the same environment, prints the path but the directory is already deleted by the time
the caller's next line runs, because the subshell's EXIT trap already fired.
Trade-offs and pitfalls
mktempwithout-dcreates a file; for a directory you must pass-d, and forgetting it
is a common copy-paste error when adapting a temp-file recipe.- Do not rely on
/tmpbeing cleaned automatically (some distributions runtmpfiles.dsweeps
on a schedule, others do not); a script must clean up after itself rather than assuming the OS
will. - A
trap ... EXITset inside a function is scoped to whichever shell process actually runs
that function. In a plain (non-subshell) function call this is the main script's shell, which
is correct; the moment you wrap the call in$(...), backticks, a pipeline, or&
(background), you have created a subshell and the trap silently stops protecting the resource
the caller thinks it owns. - If the script can be interrupted by a signal other than normal exit (Ctrl-C,
kill),trap cleanup EXITstill fires in bash because EXIT is triggered on any script termination path,
but if you also trapINT/TERMexplicitly for other cleanup, make sure that handler does not
itselfexitwithout eventually reaching the shared cleanup logic, or the temp directory leaks.
You suspect a host can't resolve external domains. Walk through how you'd isolate whether the problem is DNS configuration, the resolver itself, or something else entirely, and how you'd interpret what you find at each step.
Sample Answer
Direct answer
The fastest way to isolate whether a "can't resolve external domains" symptom is genuinely DNS, versus something else entirely (routing, a firewall, an application-level issue), is to bypass DNS altogether first: try reaching something by raw IP address. If that works, the problem is specifically in name resolution, not connectivity; if it also fails, DNS was never the actual problem, and the investigation moves to routing, firewall rules, or the application layer instead. From there, work inward through the resolver configuration, then the resolution mechanism itself.
Structured elaboration
Step 1: rule DNS in or out entirely, by removing it from the equation. A bare TCP-level check to a known-good public IP, bypassing name lookup completely, tells you immediately whether basic network connectivity out of this host works at all: nc -zv 1.1.1.1 443 (or curl -v telnet://1.1.1.1:443) opens a raw TCP connection with no DNS lookup and no TLS negotiation involved. Reaching for curl -v https://<ip> directly is a trap worth knowing about: many production services today sit behind a shared, SNI-routed IP (a CDN or reverse proxy fronting many hostnames on one address), so connecting straight to that IP over HTTPS without also presenting the right hostname produces a TLS handshake failure that looks exactly like a connectivity problem but is actually just the server rejecting an unrecognized SNI value, nothing to do with DNS, routing, or a firewall at all. If you specifically need to test HTTPS to a named service while forcing a particular IP, curl -v --resolve <domain>:443:<ip> https://<domain> is the correct form, since it forces resolution to that IP while still sending the real domain in the TLS handshake and the Host header, so a failure there is actually meaningful. If even the bare TCP check fails, DNS isn't the culprit and the investigation redirects toward routing (ip route, traceroute toward that IP), a firewall blocking outbound traffic generally, or a network path issue upstream of this host, not toward anything DNS-specific.
Step 2: if direct-IP connectivity works, confirm what resolver configuration this host is actually using. resolvectl status (on hosts running systemd-resolved, the resolver management service that ships with most modern distributions) shows the effective nameservers currently in use, which is not always the same as what's naively expected from reading /etc/resolv.conf directly, since systemd-resolved often manages that file itself and the visible contents can be a stub pointing back at the local resolver service rather than the real upstream nameserver. Older or minimal systems without systemd-resolved read /etc/resolv.conf directly as the source of truth for nameservers.
Step 3: test raw DNS resolution directly, bypassing any local caching layer. dig @<nameserver-ip> example.com queries a specific nameserver directly, letting you determine whether the configured nameserver itself is reachable and answering correctly, independent of anything cached or misconfigured locally on this host. Comparing that against dig example.com (using the host's normal configured resolution path) tells you whether the problem is upstream (the configured nameserver itself failing to answer) or local (something wrong in how this host is routing its own queries to that nameserver, such as a local caching resolver stuck or misconfigured).
Step 4: confirm the resolution order and any local overrides. /etc/nsswitch.conf's hosts: line determines the order name resolution actually checks (commonly files dns, meaning /etc/hosts is consulted before DNS at all); a stale or incorrect entry in /etc/hosts for the domain in question, or for the nameserver's own hostname if it's referenced by name rather than IP, can shadow DNS entirely and produce a misleading symptom that looks like a DNS failure but is actually a local override taking precedence. getent ahosts <domain> resolves through the full configured chain (respecting nsswitch.conf order, unlike dig, which talks to a nameserver directly and skips local overrides entirely), which is useful specifically for confirming whether the application-visible resolution path matches what dig alone would suggest.
Step 5: if a nameserver in the chain appears reachable by IP but queries to it hang or get no response, check whether traffic is actually reaching it. A firewall or network policy blocking outbound traffic on port 53 (the standard DNS port, using UDP for lookups and sometimes TCP for larger or specific queries) produces exactly this symptom: the nameserver's IP is reachable for other things, but DNS queries specifically never get an answer. tcpdump -i any port 53 while retrying a lookup shows whether the query is even leaving the host, and whether any response comes back at all, which tells you definitively whether this is a local firewall problem, a silent network drop somewhere upstream, or a nameserver that received the query and simply isn't answering.
Interpreting each step's result as you go, rather than running the whole checklist blindly: a failure at step 1 means stop, this was never a DNS problem. A pass at step 1 with a failure resolving via the configured nameserver in step 3 means the nameserver itself, or the path to it, is the problem, not this host's configuration. A mismatch between dig's direct answer and getent's full-chain answer in step 4 means a local override (most often /etc/hosts or resolver misconfiguration) is intercepting resolution before DNS is even consulted.
Worked example
A host can't reach api.example.com. nc -zv 1.1.1.1 443 (a stable, dedicated public IP, bypassing DNS entirely) succeeds immediately, confirming basic connectivity out of the host is fine and this is specifically a name-resolution problem, not a routing or firewall issue. resolvectl status shows the effective nameserver configured is 10.0.0.2, an internal resolver. dig @10.0.0.2 api.example.com times out with no response at all, while dig @8.8.8.8 api.example.com (a public nameserver, reached directly) succeeds, isolating the problem specifically to reaching or getting an answer from 10.0.0.2. tcpdump -i any port 53 and host 10.0.0.2 while retrying the failing lookup shows the query leaving the host but no response ever arriving back, which, combined with ping 10.0.0.2 succeeding fine (so the host itself is reachable for other traffic), points at either the internal resolver service being down or overloaded rather than the network path, or a firewall change that specifically blocks port 53 to that resolver while leaving general connectivity to its IP untouched. That combination of evidence (reachable by ping, unreachable specifically on port 53, works fine against an entirely different nameserver) narrows the fix to that specific resolver's own health or its port-53 firewall rule, not this host's own configuration.
Trade-offs & pitfalls
- Jumping straight to
digor resolver-config inspection without first ruling DNS in or out via a direct-IP test risks spending time deep in resolver configuration on a problem that was actually a routing or firewall issue with nothing to do with DNS at all. digtalking directly to a nameserver andgetent/application-level lookups going through the fullnsswitch.confchain can legitimately disagree; treating them as interchangeable checks misses local override problems like a stale/etc/hostsentry.- Testing raw connectivity with
curl -v https://<ip>directly (instead of a bare TCP check or--resolve) can produce a TLS handshake failure purely because of SNI-based routing on a shared IP, which is easy to misread as "so the network path is broken" when the path is actually fine and only the TLS handshake's hostname failed to match; a bare TCP check (nc -zv) avoids that false signal entirely. systemd-resolvedenvironments can make/etc/resolv.confmisleading if read naively, since it may point at a local stub resolver rather than showing the real upstream nameserver directly;resolvectl statusis the more reliable source of truth on those hosts.
You are responsible for hardening SSH across a fleet. Provide an actionable plan covering /etc/ssh/sshd_config changes, centralized key or certificate management, rate limiting/brute force protection, use of bastion hosts, monitoring and alerting for suspicious access, and safe rollout steps to avoid locking out administrators.
Sample Answer
Direct answer
Hardening SSH fleet-wide means treating it as a single, deliberate rollout, not a checklist you paste into every host and hope for the best: lock down sshd_config to remove the weak defaults, move key management off individual hosts and onto something centralized and rotatable, add rate limiting so brute-force attempts can't run unbounded, funnel access through bastion hosts so there's one narrow, monitored front door instead of many, alert on suspicious access patterns, and roll every change out in a way that guarantees you can't lock yourself out of a fleet you can no longer reach.
Structured elaboration
sshd_config changes (at least eight concrete ones worth making):
PermitRootLogin no(no direct root login over SSH at all; usesudoafter logging in as a named user).PasswordAuthentication no(key-based or certificate-based auth only; passwords are brute-forceable, keys practically aren't).PubkeyAuthentication yes(explicit, though it's the default on modern OpenSSH; being explicit avoids surprises across distro defaults).AllowUsersorAllowGroupsrestricting SSH to a named allowlist of accounts or groups, so a stray local account with a weak password (if password auth were ever accidentally re-enabled) still can't SSH in.MaxAuthTries 3(caps how many authentication attempts a single connection gets before the server drops it, slowing down automated guessing).ClientAliveInterval 300paired withClientAliveCountMax 2(server sends a keepalive probe every 5 minutes and drops the session after 2 missed responses, so a stale or hijacked idle session doesn't linger indefinitely).LoginGraceTime 30(a connection that doesn't complete authentication within 30 seconds is dropped, limiting how long a slow-drip attack can hold a connection slot open).X11Forwarding noandAllowTcpForwarding nounless a specific host genuinely needs them, since both expand what a compromised SSH session can be used for.- Restricting
Ciphers,MACs(Message Authentication Codes, which verify that each packet
wasn't tampered with in transit), andKexAlgorithms(Key Exchange algorithms, which negotiate
the shared session key when a connection is first established) to modern, non-deprecated options
only (removes legacy algorithms a downgrade attack could target).
Changing the default port (22) is a genuine trade-off, not a free win: it cuts down the sheer volume of automated scanning noise hitting your logs, which has real operational value (smaller haystack for real signal), but it is security through obscurity, not a security control on its own: a targeted attacker doing a full port scan finds it in seconds. Do it for noise reduction if you want it, never present it as a substitute for the actual auth hardening above.
Centralized key or certificate management: distributing and rotating individual authorized_keys files across a fleet by hand doesn't scale and makes revocation slow (you have to touch every host). An SSH certificate authority (issuing short-lived signed certificates instead of long-lived static keys) lets you revoke access by simply not renewing a certificate rather than hunting down and removing a key from every host it was copied to, and it gives you a single place to audit who was issued access and when.
Rate limiting and brute-force protection: something watching auth logs and temporarily banning source IPs after repeated failures (fail2ban or an equivalent) adds a layer beyond MaxAuthTries, which only limits attempts within a single connection, not across many connection attempts from the same source.
Bastion hosts: route all SSH access through a small number of hardened, heavily monitored jump hosts rather than exposing every fleet member directly to the internet or even to the whole internal network; this shrinks the actual attack surface to those few hosts and gives you one place to centralize logging and session recording.
Monitoring and alerting: ship sshd auth logs (successful and failed logins, key fingerprints used) to centralized logging, and alert on patterns like repeated failures from one source, logins at unusual hours for that account, or logins from a new geography for a user who normally connects from one office.
Safe rollout, so you don't lock out your own administrators:
- Never deploy
PasswordAuthentication noto a fleet until you've confirmed every administrator who needs access already has a working key deployed and tested. - Always test the config file for syntax errors before applying:
sshd -tvalidatessshd_configwithout touching the running daemon. - Reload (
systemctl reload sshd), don't restart, where possible: a reload applies new config to the existing daemon without dropping already-established connections, so if something is subtly wrong you still have your current session to fix it from. - Roll out through configuration management in waves (a canary group of a few non-critical hosts first, verified, then the rest), never as a single fleet-wide push, and keep a documented, tested out-of-band access path (console access, a cloud provider's serial console, or a bastion path outside the change) as a safety net for the rollout itself.
Worked example
For a 200-host fleet, the rollout order matters as much as the content: first, issue everyone SSH certificates from a central certificate authority and confirm every admin can authenticate with the new certificate while password auth is STILL enabled as a fallback. Second, push the full sshd_config hardening set above to a 5-host canary group via configuration management, with PasswordAuthentication left yes for one more cycle even there, and confirm certificate-based login works cleanly and sshd -t passes on every canary host. Third, once the canary group has been stable for a monitoring cycle (confirming no locked-out sessions, no unexpected auth-log errors), flip PasswordAuthentication no on the canary group specifically, verify again, then roll the complete config (hardening plus password auth disabled) to the rest of the fleet in batches, keeping the documented out-of-band console access path live and tested throughout in case any batch goes wrong.
Trade-offs & pitfalls
- Disabling password authentication before every administrator has a verified, working key is the single most common way this kind of hardening locks a team out of their own fleet; sequencing that verification first is not optional.
- A restart (versus a reload) of
sshdmid-change, on a host you're connected to only via SSH, can drop your own session before you've confirmed the new config actually works, leaving you locked out with no way back in except out-of-band access. - Renaming the SSH port reduces log noise but provides no real protection against a targeted attacker; treat it as a nice-to-have, not a control you'd cite in a security review.
Compare inspecting logs under /var/log with using journalctl on a modern systemd-based host. When would you prefer one over the other, and how do you extract logs for a single service across multiple boots and time ranges? Provide example journalctl commands.
Sample Answer
Two different logging models
Plain-text logs under /var/log are files written by individual daemons (or by rsyslog/syslog-ng aggregating messages from several daemons) using whatever format each program chooses; they persist as ordinary files, survive reboots by definition, and are searched with generic text tools (grep, awk, zgrep on rotated copies). journalctl queries journald, systemd's structured logging daemon, which stores entries (optionally) in a binary, indexed format with rich metadata attached to every entry (unit name, PID, UID, boot ID, priority level, and more) captured automatically at write time, not just whatever text the program chose to print.
When to prefer one over the other
Prefer journalctl when the question is about a specific systemd-managed unit (which service, across how many restarts, in what time window), because the structured metadata makes exactly that kind of filtering fast and precise without needing to know or guess a log filename or format. Prefer plain /var/log files (or an application's own dedicated log) when the software isn't managed by systemd at all, when you need a log that will survive journald's storage limits or a service like rsyslog shipping copies off-host for long-term retention (journald's in-memory or on-disk ring buffer is not unlimited and can be configured to rotate aggressively), or when a tool downstream of the logs (a log-shipping agent, a legacy parser) only understands plain text files, not the journal's binary format.
Extracting a single service's logs across boots and time ranges
journalctl -u nginx.service # everything logged for this unit, across all retained boots
journalctl -u nginx.service -b -1 # only the previous boot
journalctl -u nginx.service --since "2026-09-20" --until "2026-09-21 06:00"
journalctl -u nginx.service -f # follow new entries live, like tail -f
journalctl -u nginx.service -p err # only entries at error priority or worse
journalctl --list-boots # see which boot IDs/indices are available at all
-u <unit> filters to exactly one systemd unit regardless of which boot logged it, which is the core advantage over grepping plain files: a plain-text approach to "this service's logs across the last three boots" would require locating and concatenating however many rotated files happen to still exist, whereas journalctl -u <unit> reads across boot boundaries transparently as long as journald has retained that history. --since/--until accept natural time expressions ("1 hour ago", "yesterday", explicit timestamps), which is considerably more convenient than reconstructing a time window from log-line timestamps by hand.
Trade-offs and pitfalls
journalctl's retained history depends entirely onjournald's configured storage limits (SystemMaxUse=, time-based retention, or purely in-memory with no persistence at all on some minimal configurations); a host configured for volatile-only journal storage loses all journal history on reboot, which plain files written to persistent storage would not.journalctl -u <unit>only shows what that unit logged tostdout/stderr(captured automatically by systemd) or explicitly sent via the journal API; if the same application also writes its own separate log file (common with databases and web servers), the two views can disagree, and neither alone is guaranteed to be the complete picture.- The journal's binary format means it cannot be searched with plain
grepdirectly on the underlying files; every query must go throughjournalctl(or a library that understands the format), which is a real workflow difference for anyone used to piping/var/logfiles into arbitrary text tools.
A server fails to boot after a kernel update. GRUB starts but the boot halts with 'unable to find root' and the system drops to an initramfs shell. As the Systems Administrator, explain step-by-step how you would recover the system using a rescue/live USB: mount partitions, chroot into the installed system, regenerate initramfs or fix /etc/fstab, and reinstall GRUB if necessary. Include concrete commands.
Sample Answer
Direct answer
"Unable to find root" dropping to an initramfs shell almost always means the boot loader found and loaded the kernel fine, but the early boot environment cannot locate or mount the real root filesystem, most often because the initramfs (the temporary root filesystem loaded into memory that contains just enough drivers and tools to find and mount the real one) is missing the driver for whatever changed, or /etc/fstab/the kernel's root parameter is pointing at a UUID or device that no longer matches after the kernel update. The fix is always the same shape: boot rescue media, mount the real filesystems, chroot into them so you are running the installed system's own tools against its own files, then regenerate whatever is stale (initramfs, GRUB config, or occasionally the boot loader itself) before rebooting.
Step by step recovery
- Boot a rescue or live USB matching the installed distribution's family (a Debian live image to rescue a Debian/Ubuntu system, a Rocky/RHEL rescue image for that family), since the chroot step later needs compatible tools and library versions.
- Identify the real root partition from the rescue environment:
lsblk -fshows every partition with its filesystem type and label; the previously failing UUID (visible in the initramfs error message, or in the rescue system's own view of/etc/fstabon the mounted partition) should match one of them. - Mount the real filesystem tree under a temporary mountpoint:
mount /dev/sda2 /mnt, then mount anything that was a separate partition (/bootis commonly separate:mount /dev/sda1 /mnt/boot). - Bind mount the virtual kernel filesystems the chroot needs to behave like a real running system, not just a static file tree:
for d in dev proc sys run; do mount --bind /$d /mnt/$d; done. chrootinto the installed system:chroot /mnt /bin/bash. From this point every command runs using the installed system's own binaries and libraries, which matters because a rescue environment'sgrub2-installormkinitrdcan be a different version than what the installed system expects.- Diagnose and fix the actual cause, which is usually one of:
- A stale or missing initramfs for the newly installed kernel, most often because a driver needed to find the root device (an NVMe, RAID, or LVM driver, where LVM, Logical Volume Manager, is a layer that lets storage be resized, snapshotted, and organized into logical volumes independently of the underlying physical partitions) is not included in it. Regenerate it:
update-initramfs -u -k allon Debian/Ubuntu, ordracut -f --regenerate-allon RHEL/Fedora, which rebuilds the initramfs for every installed kernel using the currently installed set of modules and rules. - A
/etc/fstabentry (or the kernel command line'sroot=parameter in GRUB's config) referencing a UUID that no longer exists, for example after a disk was reformatted or an LVM volume renamed. Fix/etc/fstabdirectly with the correct current UUID fromblkid, and if the kernel command line itself is wrong, that lives in GRUB's generated config, fixed by regenerating it (next step) after correcting whatever GRUB reads to build theroot=line, typically/etc/default/grub'sGRUB_CMDLINE_LINUXor the installed kernel's own defaults. - Occasionally the boot loader's own installation is damaged (rare after a kernel update alone, more common after a disk swap or partition table change): reinstall it with
grub-install /dev/sda(Debian/Ubuntu) orgrub2-install /dev/sda(RHEL/Fedora), pointed at the whole disk device, not a partition.
- A stale or missing initramfs for the newly installed kernel, most often because a driver needed to find the root device (an NVMe, RAID, or LVM driver, where LVM, Logical Volume Manager, is a layer that lets storage be resized, snapshotted, and organized into logical volumes independently of the underlying physical partitions) is not included in it. Regenerate it:
- Regenerate the boot loader's configuration so it actually reflects any fix above:
update-grub(Debian/Ubuntu, which wrapsgrub-mkconfig -o /boot/grub/grub.cfg) orgrub2-mkconfig -o /boot/grub2/grub.cfg(RHEL/Fedora BIOS;/boot/efi/EFI/<distro>/grub.cfgon UEFI systems, so confirm which boot mode is in use with[ -d /sys/firmware/efi ]before targeting the config path). - Exit the chroot, unmount everything in reverse order, and reboot:
exit, thenumount -R /mnt, then reboot without the rescue media attached.
Worked example (representative command sequence)
# from the rescue/live environment
lsblk -f
mount /dev/sda2 /mnt
mount /dev/sda1 /mnt/boot
for d in dev proc sys run; do mount --bind /$d /mnt/$d; done
chroot /mnt /bin/bash
# inside the chroot, Debian/Ubuntu family
update-initramfs -u -k all
update-grub
# inside the chroot, RHEL/Fedora family, if that is the installed distro instead
dracut -f --regenerate-all
grub2-mkconfig -o /boot/grub2/grub.cfg
exit
umount -R /mnt
reboot
This procedure is a well established, standard recovery sequence rather than something to execute end to end in a sandboxed container: a real boot loader install and a real initramfs regeneration require an actual bootable disk and firmware environment (BIOS or UEFI), neither of which exists inside this environment's containers, so nothing here claims to have been run and captured as live output; every command above is the exact, current syntax for each package (verified as the currently documented options for update-initramfs, dracut, grub-install/grub2-install, and update-grub/grub2-mkconfig), not folklore.
Trade-offs and pitfalls
The most common mistake is regenerating GRUB's config without first fixing the underlying cause (a missing driver in initramfs, or a stale UUID): update-grub/grub2-mkconfig only rewrites the boot menu from whatever is currently true about the system, it does not fix a wrong fstab entry or rebuild a driver-incomplete initramfs on its own. A second mistake is running grub-install against a partition (/dev/sda1) instead of the whole disk device (/dev/sda) on a BIOS/MBR system, which installs the boot loader in the wrong location entirely. Third, forgetting the bind mounts for /dev, /proc, /sys before chrooting means tools like dracut or grub2-mkconfig inside the chroot cannot see real devices or the running kernel's state, and will either fail outright or silently generate an incorrect configuration based on an incomplete view of the system.
Unlock Full Question Bank
Get access to all Linux System Administration interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.