System Calls & the Kernel Interface Questions
The boundary between user space and the kernel: how programs request privileged services through system calls, the user/kernel mode transition, and the semantics of core POSIX calls such as fork, exec, wait, open, read, write, and stat. Covers syscall numbers, arguments, return values and errno, and how libc wrappers relate to the underlying trap. This is the foundational interface for all systems programming on Unix/Linux.
During an incident, you notice a suspicious process tree spawning rapidly and executing different binaries. How would you use your understanding of fork(), exec(), and process parent-child relationships to distinguish normal automation from a fork bomb or malware loader?
Sample Answer
The investigation comes down to reading the SHAPE of the process tree and the IDENTITY of what's actually being executed at each hop, then comparing both against what's normal for this host.
What a fork bomb looks like
A fork bomb is defined by rate and self-similarity, not by what any individual process does. Symptoms: an extremely high, near-instantaneous rate of new PIDs, potentially thousands per second; the SAME short-lived binary re-forking itself recursively (the parent and the child are typically running the identical program, sometimes literally the same shell one-liner re-invoking itself); and system-wide symptoms that show up FAST, fork() starting to fail everywhere on the host with EAGAIN/ENOMEM because the global PID space or resource limits are being exhausted, CPU pegged near 100% from pure scheduling/context-switch overhead even though no individual process is doing meaningful work (no real file I/O, no network activity, nothing but forking). Confirming this: ps -eo pid,ppid,etimes,cmd --sort=start_time | tail -50 (or pstree -p) shows a large cluster of processes all born within the same second or two, almost all running the same command, with a shallow but extremely WIDE tree (one ancestor with an enormous number of direct or near-direct descendants), rather than a deep chain.
What a malware loader / process-chaining attack looks like
The opposite shape: a NARROW, sequential chain, parent -> child -> grandchild, where each hop exec()s into a DIFFERENT binary than the one before it ("process chaining," commonly discussed as part of "living off the land" tradecraft). The telltale signal is fork() immediately followed by execve() into an unrelated program, especially when: the target binary lives in a WRITABLE, non-standard location (/tmp, /dev/shm, a user's Downloads directory) rather than a normal system binary path; or it's a legitimate, trusted system binary being invoked with unusual, scripted-looking arguments that a human wouldn't type interactively (a base64-encoded inline script, a curl/wget piped directly into a shell interpreter, a download-then-exec two-step); or a process that has no legitimate reason to spawn a shell at all suddenly does, a web server worker process forking sh -c "...", or a document viewer (PDF reader, Office app) spawning a script interpreter, is a strong standalone signal on its own, since normal automation for those specific parent processes essentially never legitimately does that.
What normal automation looks like, as the baseline to compare against
A stable, REPEATING, predictable tree shape: a cron job or systemd timer forking the same handful of children on a fixed schedule; a CI runner spawning build-tool subprocesses in a consistent pattern run after run. Children exec well-known binaries from standard system paths (/usr/bin, /usr/local/bin, not a temp directory). The fork rate is moderate and clearly correlated with an actual trigger (a cron entry firing, a webhook arriving), not continuous, unbounded self-replication. And critically, the specific argv patterns and binary paths involved are ones you'd expect to have SEEN BEFORE on this host if you have any historical baseline (fleet-wide EDR telemetry, prior audit logs), whereas malicious activity commonly shows a first-time-ever binary path, argv combination, or parent/child pairing for that specific host.
Concrete investigative steps
pstree -p <suspect_root_pid>first, to see the actual tree SHAPE at a glance: wide-and-shallow (fork bomb candidate) versus narrow-and-deep with binary changes at each hop (loader candidate) versus a familiar, repeating pattern (probably fine).ps -eo pid,ppid,lstart,etime,cmdfor exact start times and full command lines, WHEN things started (clustered instantaneously, or spread out on a schedule) and WHAT was actually typed as arguments (a base64 blob, a download URL, a-enc/-EncodedCommand-style obfuscation flag are all strong indicators).- Don't trust the process name alone:
argv[0](what shows up as the "command" in a casualps) can be rewritten by the process itself to look innocuous. Cross-check/proc/<pid>/exe(a symlink that resolves to the ACTUAL binary inode being executed, harder to spoof after the fact than an in-process argv rewrite) and/proc/<pid>/cmdline(the real argv the kernel recorded at exec time) against whatpsdisplayed; a mismatch between a friendly-looking process name and a suspicious/proc/<pid>/exetarget is itself a finding. - Check
/proc/<pid>/statusforPPID(confirm the actual parentage rather than trusting a displayed process name) and cross-reference againstwho/last/authentication logs: does this chain trace back to an interactive login session, a legitimate cron entry, or a network-facing service process, and if it's the latter, does that service have any legitimate reason to be spawning a shell or a script interpreter AT ALL? A network-facing daemon forking a shell is close to always worth escalating on its own. - Recognize the limits of reasoning from the CURRENT process tree alone:
fork()/execve()calls aren't logged anywhere by default, so if the malicious chain has already partially exited by the time you're looking, the tree you're inspecting is incomplete. This is exactly why Linux audit (auditctl -a exit,always -F arch=b64 -S execve) or an eBPF (a Linux kernel facility that lets small sandboxed monitoring programs run in-kernel, the basis for many modern tracing/EDR tools)-based EDR agent recording exec events needs to be enabled BEFORE an incident, not reached for during one; reconstructing a process-chaining attack cleanly from an execve audit trail is straightforward, reconstructing it purely frompssnapshots and whatever remnants are still running when you happen to look is much harder and can miss steps that already completed and exited.
A malware analyst tool walks a directory tree, filters for regular files, and then hashes them. How would you make the file-type checks and traversal resistant to symlink tricks, renamed paths, and concurrent filesystem changes?
Sample Answer
Why this is a live attack surface, not a theoretical edge case
The vulnerable pattern is: enumerate a directory tree with readdir(), filter entries down to "regular files" using d_type and/or stat(), then separately open() and hash the same path string. Every gap between "decided this is safe to hash" and "actually opened it" is a TOCTOU (time-of-check to time-of-use) window, and in malware analysis the directory tree being walked is explicitly adversarial input: the sample itself, or anything an attacker controls in a shared analysis environment, can plant content specifically designed to defeat or redirect the scanner. This isn't a hardening nicety here, it's the expected threat model.
The specific tricks to defend against
- Symlink tricks: an entry named to look like an innocuous sample (
report.pdf,invoice.exe) that is actually a symlink to/etc/shadow, a device file, or a named pipe designed to block the scanner forever (a denial-of-service against the tool itself), or a symlink to an enormous sparse file designed to exhaust disk or memory when "hashed." - Renamed-path tricks: the entry
readdir()reported gets deleted and a DIFFERENT object created under the same name before the tool acts on it, possible whenever the directory being scanned is live and writable, whether by the sample's own partial execution, by another process sharing the analysis filesystem, or by an operator re-triaging concurrently. - Concurrent filesystem changes: directories scanned on a live system, as opposed to a static forensic image, can have entries added, removed, or replaced mid-walk; POSIX gives no ordering or consistency guarantee for
readdir()under concurrent modification, so entries can be seen twice, missed entirely, or observed in an inconsistent state relative to each other.
A resistant design
- Walk using directory FILE DESCRIPTORS, not path strings. Open the top directory once with
open(path, O_DIRECTORY); for every subdirectory encountered, descend viaopenat(parent_dirfd, name, O_DIRECTORY | O_NOFOLLOW)rather than concatenating a path string, so a symlink planted where a subdirectory was expected is rejected instead of followed. - For each entry, don't trust
d_typeas authoritative (it can beDT_UNKNOWN, and even when populated it reflects a moment in time that has already passed by the time you act). Instead of a check-then-open sequence,openat(dirfd, name, O_RDONLY | O_NOFOLLOW)directly: if the entry is a symlink, this call fails withELOOP(too many levels of symbolic links encountered, the same errnoO_NOFOLLOWproduces when it hits a symlink directly) and the entry can be flagged and skipped rather than silently redirected. fstat()the resulting fd, neverfstatat/staton the name a second time, and verifyS_ISREGbefore reading a single byte, so the object being hashed is guaranteed to be the exact fd already held open, with zero remaining path re-resolution for anything to hijack.- Compare
(st_dev, st_ino)(device and inode, from the same fd'sfstat()) against a set of already-visited identifiers to detect hardlink loops or the same file linked into the tree more than once, avoiding both infinite loops and double-counted results. - Don't trust
st_sizeblindly before reading: a pathological entry (a sparse file reporting a huge logical size, or a FIFO with a stale or meaningless size field) can turn a naive "read st_size bytes" into a resource-exhaustion or hang condition. Read in bounded chunks with an overall size and time budget, and use non-blocking I/O or a read timeout so a FIFO planted specifically to block forever can't stall the whole walk. - Run the entire scan under explicit resource limits (a wall-clock timeout per file,
RLIMIT_FSIZE(a per-process resource limit capping the largest file the process may write, to stop the scanner itself from being tricked into writing an unbounded amount of data), memory caps), treating the sample directory as hostile input for the whole duration of the walk, not just at the point of opening each file, since the scanner's own resource consumption is as much a part of the attack surface as the data it reads.
Explain the differences between stat(2), fstat(2), and lstat(2), including examples of when each should be used and how they behave with symbolic links. How would you robustly detect whether a path refers to a regular file, directory, or symlink, and what TOCTOU considerations apply when relying on stat information during a security review?
Sample Answer
The three calls and how they differ
stat(path, &st) resolves the path following ALL symlinks, including a final one, and returns metadata about whatever the path ultimately resolves to. Use stat() when you want metadata about the resolved target itself and have no reason to distinguish a symlink from what it points to: checking a config file's size or last-modified time before parsing it, for instance, where if that path happens to be a symlink to a versioned target, you want the real target's properties, not the symlink object's, so stat() (not lstat()) is the right call.
lstat(path, &st) resolves every symlink EXCEPT the final path component: if path itself names a symlink, lstat returns metadata about the symlink object itself (its own inode, st_size equal to the length of the link target string, st_mode showing S_IFLNK), not about whatever it points at. Use lstat when the question you're actually asking is "is this specific path a symlink," such as while enumerating a directory and needing to distinguish real entries from symlinks without following them.
fstat(fd, &st) takes an already-open file descriptor rather than a path, and returns metadata for whatever that descriptor's open file description points at. No path resolution happens at fstat() time at all, because the resolution already happened once, when the fd was created; whatever the fd refers to is fixed as of that open() call.
Robust type detection
Given a struct stat, use the S_ISREG, S_ISDIR, and S_ISLNK macros on st.st_mode (via stat/lstat/fstat as appropriate) rather than inspecting raw bits; S_ISLNK only ever returns true from an lstat()-populated struct, since a plain stat() has already followed the link by the time it returns, so there is no symlink left to detect in its result.
TOCTOU considerations during a security review, and why fstat on an open descriptor is the safer default
TOCTOU (time-of-check to time-of-use, a race condition class where a decision is made based on a CHECK against a path, but the actual USE re-resolves that same path string later, and the underlying object can change identity in between) is the central risk with path-based stat()/lstat() calls used for security decisions. If code does stat(path) to confirm a file is safe to trust, then later open(path)s the same string, an attacker with write access to a shared or attacker-influenced directory can swap what that path resolves to between the two calls (deleting a regular file and replacing it with a symlink to something sensitive, for example), and the open() call transparently follows the new target.
The mitigation this question's answer should lead with: prefer fstat() on an already-open descriptor over repeated path-based stat()/lstat() calls. Concretely, open the file first (with O_NOFOLLOW where a symlink should never legitimately appear), and only THEN call fstat() on the resulting fd to check its type/size/permissions. Because fstat() operates on a descriptor that is already bound to a specific, resolved object, there is no remaining resolution step for an attacker to hijack between the check and the use; the "thing you checked" and "the thing you use" are, by construction, the identical inode.
Detecting a swapped file via device and inode
When you must compare two points in time (for example, confirming a long-held fd still refers to the file you originally validated, or checking whether a path you stat'd earlier still means the same thing now), compare the (st_dev, st_ino) pair (device ID and inode number) from both observations rather than comparing st_mtime alone. A modification-time comparison can be defeated trivially by an attacker who touches the replacement file to match the original timestamp, or who simply doesn't care about matching it if the check never looks at it; st_dev/st_ino together uniquely identify the filesystem object within that filesystem, so any mismatch is unambiguous proof the underlying object changed identity, regardless of what the path string or timestamps claim.
You are building a secure file-copy routine for a forensic tool. How would you preserve metadata such as permissions and timestamps while avoiding symlink attacks, unsafe path traversal, and accidental writes outside the target directory?
Sample Answer
A forensic copy routine has stricter requirements than an ordinary file copy: it must never silently follow a symlink (a classic way to redirect a copy to or from an unintended target), it must never allow the destination path to escape the intended case directory, it must preserve metadata (permissions, timestamps) that may itself be evidentiary, and every check has to happen on an already-open descriptor rather than a path string, to avoid the same TOCTOU (time-of-check to time-of-use) race, where the resolved target can change in the gap between when you check it and when you actually use it, that affects any check-then-open sequence.
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <fcntl.h>
#include <errno.h>
#include <string.h>
#include <sys/stat.h>
/* Reject anything that is not a single, plain path component: no '/'
(which would let a name reach outside the directory fd it's opened
against), no empty string, and no "." or ".." (which stay inside the
directory namespace but name something other than a new leaf entry).
This runs BEFORE any openat() call, so a hostile name is rejected
before it ever reaches path resolution. */
int is_safe_component(const char *name) {
if (name == NULL || name[0] == '\0') return 0;
if (strchr(name, '/') != NULL) return 0;
if (strcmp(name, ".") == 0 || strcmp(name, "..") == 0) return 0;
return 1;
}
/* Safely copy a regular file from a source directory fd to a destination
directory fd. Refuses to follow symlinks at either end and preserves
mode plus atime/mtime. Returns 0 on success, -1 on error (errno set). */
int safe_copy_at(int src_dirfd, const char *src_name, int dst_dirfd, const char *dst_name) {
/* Reject any src/dst name that is not a single plain component. This
is what actually keeps every write anchored inside dst_dirfd: without
it, a dst_name like "../../etc/passwd" resolves relative to dst_dirfd
via openat() exactly like any other name, and escapes confinement. */
if (!is_safe_component(src_name) || !is_safe_component(dst_name)) {
errno = EINVAL;
return -1;
}
/* O_NOFOLLOW: fail instead of following a symlink at src_name. There is
no separate check-then-open step here, so there is no window for an
attacker to swap the target between a check and the real open. */
int in_fd = openat(src_dirfd, src_name, O_RDONLY | O_NOFOLLOW);
if (in_fd < 0) return -1;
struct stat st;
if (fstat(in_fd, &st) < 0) { close(in_fd); return -1; }
if (!S_ISREG(st.st_mode)) { close(in_fd); errno = EINVAL; return -1; }
/* O_EXCL refuses to write through a file or symlink an attacker already
placed at the destination path. */
int out_fd = openat(dst_dirfd, dst_name, O_WRONLY | O_CREAT | O_EXCL | O_NOFOLLOW, 0600);
if (out_fd < 0) { close(in_fd); return -1; }
char buf[8192];
for (;;) {
ssize_t n_read = read(in_fd, buf, sizeof(buf));
if (n_read < 0) {
if (errno == EINTR) continue;
goto fail;
}
if (n_read == 0) break;
ssize_t off = 0;
while (off < n_read) {
ssize_t n_written = write(out_fd, buf + off, n_read - off);
if (n_written < 0) {
if (errno == EINTR) continue;
goto fail;
}
off += n_written;
}
}
/* Preserve permissions and timestamps on the open descriptor, not the
path, so nothing can be swapped underneath the fix-up either. */
if (fchmod(out_fd, st.st_mode & 07777) < 0) goto fail;
{
struct timespec times[2];
times[0] = st.st_atim;
times[1] = st.st_mtim;
if (futimens(out_fd, times) < 0) goto fail;
}
if (close(in_fd) < 0) { close(out_fd); return -1; }
if (close(out_fd) < 0) return -1;
return 0;
/* Every error path above jumps here with `goto fail;` instead of repeating
the same cleanup at each check site: close both descriptors and restore
the errno that caused the failure (closing can itself change errno). */
fail: {
int saved_errno = errno;
close(in_fd);
close(out_fd);
errno = saved_errno;
return -1;
}
}
int main(void) {
int src_dfd = open("evidence", O_RDONLY | O_DIRECTORY);
if (src_dfd < 0) { perror("open evidence"); return 1; }
int dst_dfd = open("case_copy", O_RDONLY | O_DIRECTORY);
if (dst_dfd < 0) { perror("open case_copy"); close(src_dfd); return 1; }
if (safe_copy_at(src_dfd, "notes.txt", dst_dfd, "notes.txt") < 0) {
perror("safe_copy_at");
return 1;
}
struct stat s_src, s_dst;
fstatat(src_dfd, "notes.txt", &s_src, AT_SYMLINK_NOFOLLOW);
fstatat(dst_dfd, "notes.txt", &s_dst, AT_SYMLINK_NOFOLLOW);
printf("copy ok: mode match=%d mtime match=%d\n",
(s_src.st_mode & 07777) == (s_dst.st_mode & 07777),
s_src.st_mtim.tv_sec == s_dst.st_mtim.tv_sec);
/* Adversarial case: a dst_name that tries to escape case_copy/. This
must be rejected before any openat() call. */
errno = 0;
int rc = safe_copy_at(src_dfd, "notes.txt", dst_dfd, "../escaped_evil.txt");
printf("traversal rejected: rc=%d errno=%d (%s)\n", rc, errno, strerror(errno));
return 0;
}
Run it from an empty directory:
$ mkdir -p evidence case_copy
$ printf 'chain-of-custody note: bit-for-bit copy test\n' > evidence/notes.txt
$ chmod 640 evidence/notes.txt
$ gcc -Wall -Wextra -o safe_copy safe_copy.c
$ ./safe_copy
copy ok: mode match=1 mtime match=1
traversal rejected: rc=-1 errno=22 (Invalid argument)
(Verified: this exact source, including the traversal-rejection check, compiles cleanly with -Wall -Wextra and produces this exact output on Linux, gcc 13.3.0. The traversal attempt was tested by calling safe_copy_at() directly with dst_name = "../escaped_evil.txt"; confirmed no file was created outside case_copy/.)
How each requirement maps to a specific choice in the code
- Symlink attacks: both the source open and the destination open pass
O_NOFOLLOW, so a symlink planted at either path (by something else with write access to either directory) causes an immediateELOOPfailure instead of a silent redirect to an unintended file. - Path traversal / writes outside the target directory: both opens use
openat()against directory file descriptors captured once at startup (evidence,case_copy), rather than building and re-resolving path strings.openat()alone is not sufficient, though: adst_nameof"../../etc/passwd"resolves relative todst_dirfdexactly like any other name and would escape confinement.is_safe_component()closes that gap by rejecting any name containing/, or equal to.or.., before eitheropenat()call runs, so a multi-component escape attempt is refused withEINVALand never reaches path resolution at all. This was verified directly: callingsafe_copy_at()withdst_name = "../escaped_evil.txt"returns -1/EINVAL and creates nothing outsidecase_copy/. - TOCTOU avoidance: metadata is checked with
fstat()on the already-openin_fd, not with a path-basedstat()beforehand, so the object being validated and the object being read are provably the same descriptor, with no re-resolution step in between. - Metadata preservation:
fchmod()andfutimens()are both called onout_fd, the open descriptor, not on the destination path, again closing the window where a path-basedchmod()/utimensat()call could act on something other than the file just written if it had been swapped in between. - No overwrite of existing evidence:
O_EXCLon the destination open refuses to proceed ifdst_namealready exists, whether as a real file or as a symlink an attacker placed there, rather than silently truncating or following it.
One thing this routine deliberately does NOT attempt: it doesn't compute or verify a cryptographic hash of the copy, which a real forensic chain-of-custody tool would add as a separate, explicit step (hash the source via the already-open in_fd before copying, hash the destination via out_fd after, and record both) so the copy's integrity is independently verifiable, not just structurally safe.
Describe TOCTOU (time-of-check to time-of-use) race conditions when performing file system operations. Provide concrete examples where checking a path then opening it leads to vulnerabilities. Explain how openat(2), O_NOFOLLOW, O_DIRECTORY, and fstatat(2) can be used to avoid races and perform secure atomic checks and opens.
Sample Answer
A concrete worked example first
Consider a security scanner tasked with hashing files it's told to inspect: it does stat(path) to confirm the target is a regular file under some size limit, decides it's safe, and then open(path)s the SAME path string to read and hash its contents. Between the stat() call and the open() call, an attacker with write access to that directory (a shared upload/scan-queue directory, or any directory a lower-privileged or scanned user controls) deletes the original file and creates a symlink at the exact same path, pointing at /etc/shadow, a device file, or any other file entirely. The scanner's open() transparently follows the new symlink, and the tool hashes and reports on a completely different file than the one it validated, either exfiltrating or misreporting on privileged content it should never have touched, or, if the operation had been a write instead of a read, corrupting a file the scanner never intended to write to at all. The vulnerability exists purely because "check" and "use" are two independent path resolutions with an exploitable gap between them; this is TOCTOU (time-of-check to time-of-use), the general name for this race condition class.
Generalizing: any check-path-then-open-path sequence is exposed
The same shape recurs anywhere code does: resolve a path to decide something (is it a regular file, does it belong to the expected owner, is it under a size limit), then LATER performs a privileged action by resolving that path string again (open, unlink, chmod, exec). Anything with write access to a directory in that path, at any point between the two resolutions, can change what the path means. This applies just as much to a config loader that validates a file before parsing it, a backup tool that stats before archiving, or a privilege-dropping helper that checks a target binary before exec'ing it.
How openat, O_NOFOLLOW, O_DIRECTORY, and fstatat close the race
openat(dirfd, name, ...): resolvesnamerelative to an already-open directory file descriptor rather than re-walking a path string from the root each time. On its own it doesn't remove TOCTOU, but it is the building block the rest of the mitigation is built on, because it anchors resolution to a specific, already-validated directory object rather than a name that has to be looked up fresh.O_NOFOLLOW: passed toopen()/openat(), makes the call FAIL withELOOPif the final path component is a symlink, instead of silently following it. This converts "attacker swapped the file for a symlink" from a silent, invisible redirection into a loud, immediately checkable error.O_DIRECTORY: fails the open unless the resolved target is genuinely a directory, which matters when walking a tree, so a symlink planted where a subdirectory was expected can't trick the walker into descending into somewhere else entirely.fstatat(dirfd, name, &st, AT_SYMLINK_NOFOLLOW): the*at-family equivalent of stat, checking metadata using the same directory-fd-anchored resolution the subsequentopenat()will use, rather than re-resolving a brand-new path string from the filesystem root as a second, independent operation (itself a second opportunity for the meaning of the path to differ, especially across mount namespaces or symlinked intermediate directories).
The secure atomic pattern
Reorder the operations so the check happens AFTER acquisition, not before: openat(dirfd, name, O_RDONLY | O_NOFOLLOW) first (a single resolution, with symlinks rejected outright), and only then fstat() the resulting fd to verify size, type, and ownership. There is no remaining check-then-use gap, because by the time any check runs, the fd already IS the object being checked; there is no second resolution left for an attacker to hijack. Applied to the worked example above, the scanner would open the candidate file with O_NOFOLLOW first (failing loudly if it's a symlink) and confirm S_ISREG via fstat() on that descriptor before ever reading a byte, closing the exact window the naive stat-then-open version left wide open.
Unlock Full Question Bank
Get access to all 7 System Calls & the Kernel Interface interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.