System Calls & the Kernel Interface Questions
The boundary between user space and the kernel: how programs request privileged services through system calls, the user/kernel mode transition, and the semantics of core POSIX calls such as fork, exec, wait, open, read, write, and stat. Covers syscall numbers, arguments, return values and errno, and how libc wrappers relate to the underlying trap. This is the foundational interface for all systems programming on Unix/Linux.
You need to rotate logs atomically across multiple processes writing to the same filename. Explain how rename(2) enables atomic rotation, what happens to existing file descriptors after rename, and when copytruncate is required. Propose a robust log rotation strategy that avoids losing logs and works with long-lived processes that keep files open.
Sample Answer
Why rename() enables atomic rotation
rename(oldpath, newpath), when both paths are on the same filesystem, is a single atomic kernel operation: the directory entry newpath is made to point at oldpath's inode, and if newpath already existed, it is atomically replaced (unlinked) as part of that same operation. There is no window where newpath is momentarily missing, and no window where a reader opening newpath could see a partially-written file at that name. This is exactly the property log rotation relies on: rename(current.log, current.log.1) moves the old log out of the way in one step, with no interval where the log path doesn't resolve to something.
What happens to already-open file descriptors after the rename
An fd that was already open on the file before the rename keeps referring to the SAME underlying inode regardless of what any directory entry now says, because rename() only changes which name in the directory tree points at that inode; it does not touch any process's open file descriptors. So a long-lived process that opened current.log before rotation and is still writing through that fd keeps writing to the RENAMED file, now called current.log.1, because the fd was bound to the inode at open time, not to the path string. This is precisely why rename-based rotation requires the writing application to reopen the log file (conventionally on receiving SIGHUP) after rotation; without that reopen, the application's writes silently continue landing in the rotated-away file forever, and the freshly created current.log stays empty from that writer's perspective even though the file on disk with that name is brand new.
When copytruncate is required instead
If the writing process cannot be signaled or otherwise made to reopen its log fd (a third-party binary with no reopen hook, or a process that can't safely be restarted), copytruncate is the fallback: copy the current file's content out to the rotated name, then ftruncate() the ORIGINAL file back to zero length in place, so the still-open fd keeps writing to the same inode and the same name. The caveat: after truncation the fd's write offset is unchanged unless the fd was opened with O_APPEND (in which case O_APPEND makes every write() reposition to the file's current end before writing, regardless of the fd's stored offset, so each write recomputes its position from the file's actual current end), so without O_APPEND the next writes can land at the old offset past the new end-of-file, leaving a gap of null bytes until the application's position catches up on its own. Copytruncate is also not atomic the way rename is: there is a brief window between the copy and the truncate where a concurrent write can land in a region already captured by the copy, or be lost from both copies, so it is a fallback for when rename's atomicity genuinely isn't available, not a strict improvement on it.
A robust rotation strategy
Prefer rename() plus a signal-driven reopen (a SIGHUP handler that calls close() then open(path, O_APPEND | O_CREAT, mode) on the log path) as the primary strategy, since it is atomic and loses no log data; fall back to copytruncate, always paired with O_APPEND on the writer's fd, only when the writer genuinely cannot be made to reopen. Keep enough rotated generations for the retention policy, and compress older generations asynchronously so the rotation step itself stays fast and doesn't block the writer.
A second worked example built on the same atomicity fact: safe config replacement
The same rename() guarantee is used for a related but distinct problem: publishing a new version of a config file (potentially secret-bearing) without ever exposing a partially-written or wrongly-permissioned file at the real path. The pattern: write the new content to a TEMPORARY file in the SAME directory as the target (same filesystem is required here too, since rename() across filesystems is not atomic and in fact fails with EXDEV), fsync() that temp file so its content is durable on disk before it's ever visible under the real name, then fchmod()/fchown() the temp file to the FINAL intended mode and ownership BEFORE the rename, not after. That ordering matters specifically for secrets: if the mode is fixed up after the rename instead of before, the file is briefly visible at its real, well-known name with an overly permissive mode, a real exposure window even if it lasts only milliseconds, and even a very short window is enough for a concurrently running process that happens to be watching that path to read it. Then rename(tmp, target) atomically publishes the new content, and finally fsync() the containing DIRECTORY's own file descriptor as well, a commonly missed step, since the directory entry update itself needs its own fsync to be durable against a crash; without it, a crash immediately after rename could leave the directory entry pointing at the old inode again on some filesystems after recovery, even though the rename() call itself returned success.
A helper process exits with an unexpected status code, and your daemon must decide whether to restart it, log it, or quarantine the host. How would you interpret exit codes, signals, and core-dump indicators in waitpid() results?
Sample Answer
The raw status isn't the exit code
waitpid()'s status output parameter is a packed, platform-defined bitfield, not directly "the exit code." The exact bit layout isn't part of the portable API (which is precisely why the macro family exists), so correct code always decodes it through WIFEXITED/WEXITSTATUS/WIFSIGNALED/WTERMSIG/WCOREDUMP/WIFSTOPPED/WSTOPSIG, never by inspecting the raw integer's bits directly. I'll show the raw bytes below purely to illustrate the packing, not as something to rely on in real code.
Decoding, with real measured examples
I forked three children (a clean exit, a nonzero exit, and a fatal signal with a core dump) and printed the raw status alongside the decoded macros:
static void run_case(const char *label, void (*fn)(void)) {
pid_t pid = fork();
if (pid == 0) { fn(); _exit(0); }
int status;
waitpid(pid, &status, 0);
printf("[%s] pid=%d raw_status=0x%04x", label, pid, status);
if (WIFEXITED(status)) {
printf(" WIFEXITED=1 WEXITSTATUS=%d\n", WEXITSTATUS(status));
} else if (WIFSIGNALED(status)) {
printf(" WIFSIGNALED=1 WTERMSIG=%d(%s) WCOREDUMP=%d\n",
WTERMSIG(status), strsignal(WTERMSIG(status)), !!WCOREDUMP(status));
}
}
/* case_clean: _exit(0); case_error: _exit(17); case_crash: raise(SIGSEGV); */
MEASURED output (Debian 12, ulimit -c unlimited):
[clean-exit] pid=5126 raw_status=0x0000 WIFEXITED=1 WEXITSTATUS=0
[nonzero-exit] pid=5127 raw_status=0x1100 WIFEXITED=1 WEXITSTATUS=17
[fatal-signal] pid=5128 raw_status=0x008b WIFSIGNALED=1 WTERMSIG=11(Segmentation fault) WCOREDUMP=1
WIFEXITEDtrue means the child calledexit()/_exit()/returned frommain()normally (regardless of value);WEXITSTATUSextracts the 8-bit exit code, note only 8 bits,0to255, is the on-the-wire Unix exit-status width, so a language runtime that lets youexit(300)will wrap silently and you'd see44, not300. The17case above lands as0x11in the packed status's upper byte, and decodes cleanly back to17.WIFSIGNALEDtrue means the child was terminated BY a signal, it never reached its own exit path at all.WTERMSIGgives which signal;WCOREDUMPtells you whether the kernel actually wrote a core file, which requires the signal to be one of the core-dumping set (SIGSEGV/SIGABRT/SIGBUS/SIGQUIT and a few others),ulimit -c/RLIMIT_COREto permit it, and enough disk/permissions for the dump to actually land. Above,0x8bdecodes as signal11(SIGSEGV) in the low 7 bits plus the0x80coredump bit set.WIFSTOPPED/WSTOPSIGonly matter if you passedWUNTRACED(or you're aptrace-based tracer): the child was merely stopped, not terminated, still alive, resumable withSIGCONT. A plain supervisor that never passesWUNTRACEDsimply never sees this case, which is usually what you want, you don't want a job-control-suspended child mistaken for one that exited.
Mapping to the restart / log / quarantine decision
WIFEXITEDandWEXITSTATUS == 0: success, normal bookkeeping, no restart needed.WIFEXITEDwith a known "won't-succeed-on-retry" convention code (say, your own agreed convention that exit code 2 means "bad config, deterministic failure"): log at high severity, do NOT restart (retrying a deterministic failure just burns cycles and floods the logs), alert a human instead.WIFEXITEDwith any OTHER nonzero code: treat as transient (a network hiccup, momentary resource contention), log it, restart with backoff, but track a per-worker restart counter so a worker that fails EVERY time still eventually trips a "restart budget exhausted, stop and page someone" threshold instead of crash-looping forever.WIFSIGNALEDwithSIGTERM/SIGINT: usually an intentional, externally requested stop (an operator, or your own supervisor issuing it as part of a controlled restart). Don't treat this as a crash on its own; correlate against whether YOUR supervisor is the one that sent it.WIFSIGNALEDwithSIGSEGV/SIGABRT/SIGBUS/SIGFPE/SIGILL(the classic "bug or tampering" set): this is the case that most warrants the "quarantine the host" branch, especially whenWCOREDUMP == 1(there's an actual core file to investigate) AND the SAME crashing signal recurs across successive restarts of the same worker. One crash might be transient memory corruption from unusual input; the SAME signal repeatedly after restart is a much stronger signal that something is deterministically wrong on that host specifically, a corrupted binary, a bad dependency version, or an exploit attempt targeting that worker, worth pulling the host for forensics rather than restarting it back into the same failure indefinitely.
Implementation note
Gate this on a per-worker restart counter with exponential backoff, and always persist the RAW exit detail (the exit code or the signal name and whether it dumped core, not just "it failed") into whatever your incident timeline uses, "worker died" without the why is close to useless hours into an on-call investigation.
Explain semantics and practical implications of O_DIRECT on Linux: how it bypasses the page cache, alignment constraints for buffers and file offsets, effects on write/read performance (especially for databases), and how to design accurate benchmarks to verify O_DIRECT behavior on a specific filesystem and block device.
Sample Answer
O_DIRECT is an open() flag that tells the kernel to bypass the page cache for I/O on that file descriptor: reads pull data straight from the storage device into the caller's buffer, and writes push straight from the caller's buffer to the device, without going through the kernel's normal buffered-I/O path that would otherwise cache pages in RAM.
Why databases specifically want this
Database engines (Postgres, and MySQL's InnoDB with innodb_flush_method=O_DIRECT, among others) maintain their OWN buffer pool with an eviction policy tuned to their access patterns, named here only as illustrative examples: clock-sweep (a cheap approximation of LRU that scans buffers in a circular order and evicts the first one whose reference bit is unset) and LRU-K variants (which evict based on a page's Kth-most-recent access instead of just its single most recent one). The specific algorithm names aren't the point; what matters is that these policies understand which pages are "hot" for the workload's actual query mix in a way the kernel's generic page-cache LRU does not, and the database layers its own durability/checkpointing logic on top. Without O_DIRECT, the same data ends up double-buffered: once in the database's buffer pool, again in the kernel's page cache, wasting RAM on a redundant copy and adding an extra memcpy on every I/O, while the kernel's generic LRU can evict pages the database's own, workload-aware policy would have kept resident. O_DIRECT removes the redundant layer so the database's cache is the only cache.
Alignment constraints
O_DIRECT requires the userspace buffer's address, the file offset, and the transfer length to all be aligned to the logical block size of the underlying device (commonly 512 bytes on older devices, 4096 bytes on modern "4Kn" native-sector SSDs/NVMe, and sometimes the filesystem imposes its own, potentially larger, alignment requirement on top). A misaligned call fails with EINVAL. This is why database engines allocate their I/O buffers with posix_memalign() (or an equivalent aligned allocator) rather than plain malloc(), which gives no alignment guarantee beyond what's needed for ordinary C types.
Performance effects, and a common misconception
Eliminating double buffering and the extra copy is the upside, but O_DIRECT also eliminates the kernel's readahead and write-back coalescing, so the application takes on the responsibility of its own prefetching and write coalescing, which is exactly what database I/O subsystems implement. The common misconception worth flagging directly: O_DIRECT is about bypassing the page CACHE, it is NOT by itself a durability guarantee. Data written with O_DIRECT can still sit in a volatile on-drive or disk-controller write cache unless the write is also paired with fsync()/O_DSYNC or the drive's own cache is disabled or battery-backed; "bypasses the kernel cache" and "is safely on persistent media" are two separate claims that O_DIRECT alone does not conflate correctly if you assume it does.
Designing an accurate benchmark
- Use a dataset several times larger than available RAM. If the working set fits entirely in page cache, buffered I/O will look artificially fast (everything after the first pass is a RAM hit) and O_DIRECT will look artificially slow by comparison, which is a same-basis violation: you would be comparing "cached reads" against "always-goes-to-device reads" and calling it a comparison of I/O paths.
- Confirm alignment is genuinely honored, not silently downgraded: the test harness should fail loudly (EINVAL) on a deliberately misaligned buffer, proving the code path is real and not falling back to buffered I/O without telling you.
- Match production concurrency (queue depth). O_DIRECT's real advantage typically shows up paired with asynchronous I/O (io_uring, libaio) issuing many requests concurrently; a single-threaded, one-request-at-a-time benchmark under-represents the win because it never lets the storage device's internal parallelism come into play.
- Keep the durability claim separate from the cache-bypass claim: if you care about write durability, benchmark fsync-inclusive latency as its own number, not folded into the raw O_DIRECT write throughput figure, since conflating them mixes two different guarantees into one basis.
- Confirm the specific filesystem and kernel version actually honor O_DIRECT for that mount; some network filesystems or overlay configurations either reject it outright or, more dangerously, accept the flag without truly bypassing the cache. Check the open() return value rather than assuming success implies the semantic you expect, and corroborate with
strace/blktracethat requests are reaching the block layer as unbuffered I/O.
Describe how to set a file descriptor to non-blocking mode in C using fcntl(2) or by passing O_NONBLOCK to open(2). Explain race conditions when setting flags after open, how to avoid them, and common bugs in servers that forget to handle EAGAIN/EWOULDBLOCK properly (especially with edge-triggered epoll).
Sample Answer
There are two ways to put a file descriptor (fd, the small integer handle the kernel hands back for an open file, socket, pipe, etc.) into non-blocking mode, then a real race condition that bites people who use the wrong one, then the class of bugs that shows up once the fd actually is non-blocking and you have to handle the 'nothing to do right now' signal correctly.
1. Setting non-blocking mode with fcntl(2) on an fd you already have
fcntl (file control) is the general-purpose syscall for inspecting and changing properties of an already-open fd. The file status flags (which include O_NONBLOCK, O_APPEND, and a few others) are read and written as a bitmask, so the correct idiom is read-modify-write, not a blind set, because a blind fcntl(fd, F_SETFL, O_NONBLOCK) silently clears every other flag that was already set on that fd (for example O_APPEND on a log file you're also writing to):
int flags = fcntl(fd, F_GETFL, 0);
if (flags == -1) { perror("fcntl(F_GETFL)"); return -1; }
if (fcntl(fd, F_SETFL, flags | O_NONBLOCK) == -1) {
perror("fcntl(F_SETFL)");
return -1;
}
2. Setting non-blocking mode with O_NONBLOCK at open(2) time
If you control the call that creates the fd, you can request non-blocking mode up front instead of a separate fcntl call afterward:
int fd = open("/tmp/myfifo", O_RDONLY | O_NONBLOCK);
Note that O_NONBLOCK is meaningful mainly for FIFOs (named pipes, a filesystem object two processes use to talk to each other without a shared parent), sockets, and terminal/character devices. On a regular disk file it is effectively a no-op: POSIX regular-file reads and writes are defined to not return EAGAIN, so O_NONBLOCK does not make disk I/O asynchronous.
3. The race condition: open() then fcntl(), versus doing it atomically
A race condition is a bug where correctness depends on the relative timing of two things that can happen in either order. Here the bug shows up when you create an fd in blocking mode and only afterward call fcntl to flip on O_NONBLOCK, leaving a window between creation and that fcntl call during which the fd is still fully blocking.
The clearest version of this is accept(). A single-threaded reactor calling accept() then fcntl(connfd, F_SETFL, ... | O_NONBLOCK) looks safe, but the moment you hand connfd off to a worker thread or another event loop for load balancing before that fcntl call runs, that worker can issue a blocking read() on what it assumes is a non-blocking socket and stall a thread that was supposed to never block. The two operations (accept, then set-nonblocking) are not atomic, so anything with access to the fd in between can observe the pre-fcntl (blocking) state.
An even harder version of the same race exists on FIFOs, and there fcntl-after-open cannot fix it at all: opening a FIFO for reading in blocking mode does not even return until a writer opens the other end. If your intent was 'open it non-blocking so I don't stall waiting for a peer,' calling fcntl after open() is too late, because open() itself is the thing that blocks.
How to avoid it: request the flag atomically at creation, using the Linux syscalls built for exactly this
open(2)already supports this directly: pass O_NONBLOCK in the flags argument, no follow-up fcntl needed (shown above).socket(2)has a Linux extension:socket(AF_INET, SOCK_STREAM | SOCK_NONBLOCK, 0)creates the socket already non-blocking.accept4(2)is the fix for the accept() race specifically:accept4(listenfd, addr, addrlen, SOCK_NONBLOCK)sets the flag on the same syscall that creates the connected fd, so there is no window where another thread can see it blocking.pipe2(2)does the same for pipes:pipe2(fds, O_NONBLOCK)instead ofpipe()followed by two fcntl calls.
(These *2/*4 variants need _GNU_SOURCE or _DEFAULT_SOURCE defined before the includes on glibc, since they're Linux-specific, not POSIX. This is the same atomicity motivation as O_CLOEXEC/SOCK_CLOEXEC, which close an analogous race where a forked child could inherit and leak an fd into an exec'd program before the parent got a chance to mark it close-on-exec.)
4. Bugs from mishandling EAGAIN/EWOULDBLOCK
Once an fd is genuinely non-blocking, a read/write/accept call that would otherwise have to wait instead returns -1 immediately and sets errno to EAGAIN (or EWOULDBLOCK: on Linux the two are numerically identical, errno 11, but POSIX does not guarantee that on every platform, so portable code checks errno == EAGAIN || errno == EWOULDBLOCK). This is not a failure, it is the kernel's way of saying 'no data/space right now, try again later.' The bugs cluster around three mistakes:
- Treating EAGAIN as a real error. Logging it as a failure, tearing down the connection, or retrying in a tight spin loop instead of going back to the event loop and waiting for the next readiness notification.
- Not looping until EAGAIN with edge-triggered epoll (EPOLLET).
epollis Linux's readiness-notification API for watching many fds at once; it can run in level-triggered mode (EPOLLLT, the default: it keeps telling you 'still readable' every time you ask, as long as data remains) or edge-triggered mode (EPOLLET: it tells you exactly once, at the moment the fd transitions from not-ready to ready). With EPOLLET, if you read one chunk, see there's more, and just move on to the next fd in your event loop 'to be fair,' the socket buffer still has bytes in it but no new edge is coming until more data arrives from the peer, so you never get woken up again for that leftover data. The connection appears to silently stall. The only correct pattern under EPOLLET is: read (or write) in a loop until the call returns -1/EAGAIN, and only then go back to epoll_wait. - Forgetting to unsubscribe from EPOLLOUT once a deferred write drains. When a write() returns a short count or EAGAIN because the send buffer is full, the standard pattern is to buffer the remainder and register interest in EPOLLOUT so you're notified when there's room again. The corresponding bug is forgetting to remove EPOLLOUT interest once the buffered data has fully drained: the socket is writable almost all the time, so epoll_wait keeps returning immediately for that fd on every iteration, burning CPU in what looks like a busy loop with no work actually happening.
- Confusing EINTR with EAGAIN. EINTR (call interrupted by a signal before any I/O happened) is a different case that also needs a retry, but retrying it correctly means simply re-issuing the same call, not re-arming epoll interest or treating it as 'no data.' Code that lumps EINTR into the same branch as EAGAIN either busy-loops or misses real data.
Checklist for a strace/code review of this pattern: confirm the non-blocking flag was set atomically at creation (or read-modify-write via fcntl if not); confirm every read/write path checks for EAGAIN/EWOULDBLOCK and treats it as 'go back to the event loop,' not an error; confirm any EPOLLET consumer loops to EAGAIN on every readiness event; confirm EPOLLOUT interest is added when a write is short and removed once the buffer is empty.
When a signal interrupts a blocking system call, how do you decide whether to retry, fail, or propagate the error? Explain why this matters for long-running services that handle log files, sockets, or child processes.
Sample Answer
When a blocking syscall gets interrupted by a signal, the kernel hands you back -1 with errno set to EINTR ("interrupted system call"), meaning nothing went wrong with the operation itself, a signal simply arrived and was handled before the call could complete. The decision isn't "always retry" or "always fail". It's a three-way policy question, and the right answer depends on what kind of syscall it was, what the signal meant, and what state the call may have already partially mutated.
The decision framework
- Was this an intentional wakeup, or truly unexpected? Some signals exist specifically to interrupt a blocking call on purpose, for example a supervisor sending
SIGTERMto tell a worker to shut down, or a timeout implemented viaalarm()+SIGALRM. If the signal is a deliberate "stop what you're doing" instruction, EINTR should propagate as "abort this operation and begin shutdown", not be silently retried. If the signal was something incidental (say,SIGWINCHfor a terminal resize reaching a daemon that has no terminal-related state to update), the interruption carries no information and the operation should simply be retried. - Can the call be safely retried at all, or did it already do partial work? This is the sharpest edge.
read()/write()on a byte stream (socket, pipe, regular file) can return a short count rather than EINTR when interrupted after transferring some bytes; that's not an error at all, it's success-with-fewer-bytes-than-requested, and the correct response is to advance your buffer offset and re-issue the call for the remainder, not restart from zero and not treat it as EINTR. A true EINTR (zero bytes transferred, call didn't even start) is always safe to blindly retry. - Does the platform/libc already restart it for you?
sigaction(2)'sSA_RESTARTflag tells the kernel: for certain "slow" syscalls (broadly, most blocking I/O calls, not all,select/poll/epoll_waitare notable syscalls that are NOT restarted even with SA_RESTART and always return EINTR on signal), automatically re-issue the call after the signal handler returns, so your code never even sees EINTR for that class of interruption. That's the least code and the safest default when you do want "ignore incidental signals, keep working", but it means you must not rely on getting woken up by every signal if you're using SA_RESTART broadly, since some signals will now be invisible to a blocking read/write.
Why it matters for long-running services
- Log files and regular file I/O: a log-writing daemon that reopens or rotates logs on
SIGHUPneeds that signal to interrupt an in-flightwrite()cleanly rather than have SA_RESTART swallow it silently and keep writing to the now-unlinked old file descriptor; here, propagate-and-act is correct, not retry. - Sockets: a server accepting connections in a loop with
accept()wants incidental signals (a child reaper'sSIGCHLD, for instance) to not tear down the accept loop; wrapping the call in an EINTR-retry loop (or usingSA_RESTART) keeps the server available. But aread()mid-request that gets EINTR from an operator-issuedSIGTERMshould propagate up as "connection being drained, stop accepting new work", not retry into another blocking read that delays shutdown. - Child processes:
waitpid()is itself interruptible; a supervisor blocked inwaitpid()needsSIGCHLDto interrupt it (that's the whole point of reaping via a signal-driven wakeup), so the correct handling there is specifically "EINTR here usually means go re-check what changed", not a blind retry loop that ignores why it woke up.
The rule of thumb that generalizes across all three: retry when the signal is incidental and the call is idempotent or was a true zero-progress interruption; propagate/abort when the signal itself IS the instruction (shutdown, timeout, reap-now); and always check for a short count before deciding you got EINTR at all, because misclassifying a short write as "clean success" silently drops the unwritten tail of the buffer, one of the most common EINTR-adjacent bugs in production log/network code.
Unlock Full Question Bank
Get access to all 45 System Calls & the Kernel Interface interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.