Kernel Architecture & OS Internals Questions
How an operating system kernel is structured and what it is responsible for: kernel subsystems (scheduler, memory manager, drivers, interrupt handling), kernel vs. user space and the cost of crossing between them, and the boot and initialization path. Includes the core mechanisms the kernel provides: process and thread lifecycle (creation, zombies and orphans, uninterruptible sleep) and context switching, processes vs. threads as an architectural choice, CPU scheduling (CFS, priorities, preemption models, real-time policies on a general-purpose kernel), virtual memory (address translation, multi-level page tables and entry formats, TLB misses and shootdowns, page faults, copy-on-write, memory-mapped files, userfaultfd, swapping, page replacement, kernel allocators, process address-space layout, page-level protection such as NX and ASLR as mechanisms, hugepages, page coloring, overcommit and the OOM killer as kernel mechanisms), interrupts and softirqs, and how loadable modules and device drivers extend the kernel. Includes worked exercises such as page-replacement traces, address-translation arithmetic and small page-table, TLB or scheduler simulations. Covers the concepts and internals, not microcontroller interrupt and ISR design, RTOS and real-time scheduling theory, day-to-day host administration, system-call and POSIX API semantics, OS-level performance tuning, or forensic and security analysis of a host.
What does the kernel track for a process versus a thread, and what do threads of one process share? When would you choose several processes over several threads for a server, and why?
Sample Answer
Direct answer
A process is a running program with its own address space, file-descriptor table, credentials and PID. A thread is a schedulable flow of execution inside a process. Threads of one process share the address space (code, globals, heap), the open files and the signal handlers, but each has its own registers, stack, thread ID and signal mask (the set of signals that thread currently blocks from being delivered). Choose threads when tasks need to share memory cheaply and a crash is acceptable to take down all of them. Choose several processes when you want fault or security isolation, or when sharing is minimal.
What the kernel tracks
On Linux the kernel schedules one task (task_struct) per thread. A "process" is a group of tasks that share resources. The clone system call creates a task, and flags decide what is shared. Each flag switches on one kind of sharing; a thread is a task created with all of them, and a separate process with none: CLONE_VM (same address space), CLONE_FILES (same descriptor table), CLONE_FS (same working directory), CLONE_SIGHAND (same signal handlers) and CLONE_THREAD (same process, so same PID as seen by getpid()). fork() is a clone that shares nothing: the child gets a copy-on-write copy of the address space (the kernel does not copy memory pages up front; parent and child share them read-only, and a page is copied only when one side writes to it), which is why it is cheap in proportion to the size of the page tables rather than the size of memory used.
| Resource | Per process | Shared by all its threads | Per thread |
|---|---|---|---|
| Address space (code, heap, globals) | Yes | Yes | No |
| File-descriptor table | Yes | Yes | No |
| Credentials (user and group IDs) | Yes | Normally the same | Kernel stores them per task |
| Signal dispositions (what to do when each signal arrives: ignore, run a handler, or terminate) | Yes | Yes | No |
| Registers, program counter, stack | No | No | Yes |
Signal mask, thread-local storage (variables declared so that each thread gets its own copy, abbreviated TLS in this table only), errno | No | No | Yes |
| Scheduling unit | No | No | Yes |
On Windows the same split exists differently: a process object owns the address space and handle table, a thread is the unit of scheduling, and there is no fork; creating a process (CreateProcess) is heavier and starts a fresh image.
A tiny experiment shows the difference. A global counter starts at 0. A thread sets it to 100 and finishes; a forked child sets it to 200 and exits. The parent then prints it each time.
#include <pthread.h>
#include <stdio.h>
#include <sys/wait.h>
#include <unistd.h>
int counter = 0; /* a global: lives in the address space */
void *bump(void *arg) { counter = 100; return NULL; }
int main(void) {
pthread_t t;
pthread_create(&t, NULL, bump, NULL); /* thread: same address space */
pthread_join(t, NULL);
printf("after thread: counter = %d\n", counter);
counter = 0;
pid_t pid = fork(); /* process: child gets its own copy */
if (pid == 0) { counter = 200; _exit(0); }
waitpid(pid, NULL, 0);
printf("after fork child: counter = %d\n", counter);
return 0;
}
Compiled with gcc and run in a Linux container, it printed:
after thread: counter = 100
after fork child: counter = 0
The thread wrote into the one shared address space, so the parent sees 100. The child wrote into its own copy, so the parent still sees 0.
Process versus threads for a server
| Concern | Threads in one process | Several processes |
|---|---|---|
| Crash isolation | A segmentation fault or abort in one thread kills them all | A crashing worker leaves siblings running |
| Creation cost | Lower (no new address space) | Higher (page tables, fork or exec) |
| Sharing data | Direct memory access, which requires locks | Needs inter-process communication (IPC) |
| IPC overhead | None for shared memory | Pipes and sockets copy through the kernel; shared-memory mappings avoid the copy but need their own synchronization |
| Security | A memory bug lets one thread read secrets of any other; the whole process has one set of permissions | Each process can have its own user ID, sandbox and filesystem view, so a compromised parser reaches little |
| Race hazards | Data races are possible on every shared variable | Races only on explicitly shared regions |
Recommendation. For a latency-sensitive service whose workers share a large in-memory cache, use threads (or an event loop on a few threads). For a component that handles untrusted input, such as a file parser, an image decoder, or a browser tab, run it in a separate low-privilege process (one given only the few permissions its job needs): if an attacker exploits a memory bug the damage stays in that process, and a crash restarts one worker. Real systems mix both: a browser uses a process per site and threads inside each, and Postgres uses a process per connection partly for isolation.
Worked example
A server decodes uploaded images. With thread-per-request, a malformed file that triggers a heap overflow (writing past the end of a heap buffer into neighbouring memory) runs the attacker's code with access to the process's encryption keys for HTTPS connections (the TLS protocol, unrelated to thread-local storage above), session cache and database credentials, all in the same address space, and a plain crash drops every in-flight request. With a pool of decoder processes running under a restricted user ID with a system-call filter (a kernel-enforced list of system calls the process may make, such as seccomp on Linux; anything else is refused or kills it), the same overflow reaches only that decoder's memory; the parent sees a child exit and starts a replacement. The price is copying each upload to the decoder through a pipe or a shared mapping.
Pitfalls and notes
- "Threads are always faster" ignores the locking they require; "processes are always safer" ignores IPC bugs and cost.
- Memory-forensics note: a memory image lists each thread as its own task entry, and because threads share one address space a dump of the process contains every thread's stack.
- Fork in a multithreaded program only copies the calling thread, which leaves locks held by other threads stuck in the child; call
execsoon after.
After a process forks, parent and child appear to have separate memory but no copying has happened yet. How does copy-on-write make that work at the page-table level, and what goes wrong for a workload that writes heavily to memory right after the fork?
Sample Answer
Direct answer
At fork, the kernel does not copy the parent's memory. It gives the child its own page tables (the per-process maps from virtual pages to physical frames) that point at the same physical frames, and it marks every writable private page read-only in both processes. Reads cost nothing extra. The first write by either side hits a read-only entry, which raises a page fault (a CPU trap into the kernel); the fault handler sees the page is shared copy-on-write, allocates a new frame, copies the 4 KiB, points the writer's entry at the copy with write permission, and resumes the instruction. A workload that writes heavily right after the fork therefore takes one fault plus one page copy per page it touches, so latency spikes and memory use climbs toward double the parent's size. Resident size (RSS) means the physical memory a process currently has mapped.
At the page-table level
forkduplicates the memory-region list (VMAs, the kernel's records of each mapping) and copies page-table entries (PTEs) for private anonymous memory. For each writable private page, the PTE is made read-only in the parent and the child, and the physical page's reference count (how many holders keep the frame alive) and map count (how many page tables map it) are increased because two address spaces now map it.- Neither process sees a difference on reads. A write to a read-only PTE inside a region the VMA says is writable is recognised as a copy-on-write fault.
- The fault handler checks who else maps the page. If another process still maps it, it allocates a new frame, copies the 4 KiB, installs the copy in the writer's PTE as writable, and drops one reference on the old frame. If the writer is the only remaining mapper (the other process already copied or exited), it just makes the PTE writable again without copying.
- The stale read-only translation must be removed from the TLB (the CPU's cache of translations), at least for that address, or the CPU would keep faulting.
Cost per first-touched page: a trap, a frame allocation, a 4 KiB copy, a PTE update and a TLB invalidation. The parent also pays this on its first write to each page after the fork, even if the child never writes, because the parent's entries were made read-only too.
What goes wrong for a write-heavy workload
- Memory blow-up. Worst case is about 2 times the parent's resident memory. A database that forks a 10 GiB process to write a snapshot while the parent keeps taking updates: if the parent dirties 30 percent of pages while the child runs, the extra is 10 GiB times 0.3 = 3 GiB, so the machine needs 13 GiB for the pair. The page tables themselves are small by comparison: 10 GiB is 2,621,440 pages, and at 8 bytes per entry that is 20 MiB.
- Latency spikes. Each first write to a shared page stalls the writer for a fault and copy, and a burst of writes right after
forkturns into a burst of faults, often visible as p99 latency jumps (the 99th-percentile response time: the delay that only 1 request in 100 exceeds) just after a snapshot starts. - Allocation failures. Under strict overcommit (
vm.overcommit_memory=2, "never overcommit") the child's private writable mappings must be accounted for atforktime, so a fork can fail withENOMEMeven though copy-on-write would rarely copy everything. Default mode 0 uses a heuristic that rejects only obvious overcommits. - Huge pages change the grain. (Huge pages are page sizes larger than 4 KiB, commonly 2 MiB on x86-64; transparent huge pages is the kernel feature that uses them automatically.) Whether a write copies 4 KiB or something larger depends on the kernel version and on transparent huge page settings (
/sys/kernel/mm/transparent_hugepage/enabled), so check your host when fork pauses look too large.
Worked example (executed)
The program maps 16,384 anonymous pages (64 MiB) and writes to all of them, forks, then in the child reads its proportional share (Pss_Anon from /proc/self/smaps_rollup; Pss is the process's share of each page divided among the processes mapping it), writes one byte into every page, and counts its own minor page faults (faults that need no disk read) via getrusage.
#define _GNU_SOURCE
#include <stdio.h>
#include <string.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/resource.h>
#include <sys/wait.h>
static long minflt(void) { struct rusage r; getrusage(RUSAGE_SELF, &r); return r.ru_minflt; }
/* Read one field (in KiB) from /proc/self/smaps_rollup */
static long rollup(const char *key) {
FILE *f = fopen("/proc/self/smaps_rollup", "r");
char line[256]; long v = -1;
while (fgets(line, sizeof line, f))
if (!strncmp(line, key, strlen(key))) { sscanf(line + strlen(key), " %ld", &v); break; }
fclose(f);
return v;
}
int main(void) {
long page = sysconf(_SC_PAGESIZE);
size_t pages = 16384, len = pages * page;
char *buf = mmap(NULL, len, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
memset(buf, 1, len); /* parent touches every page */
printf("page size %ld, buffer %zu pages = %zu KiB\n", page, pages, len / 1024);
fflush(stdout);
pid_t pid = fork();
if (pid == 0) {
long f0 = minflt(), p0 = rollup("Pss_Anon:");
printf("child, right after fork: Pss_Anon %ld KiB\n", p0);
for (size_t i = 0; i < len; i += page) buf[i] = 2; /* one byte per page */
long f1 = minflt(), p1 = rollup("Pss_Anon:");
printf("child, after writing every page: new minor faults %ld, Pss_Anon %ld KiB\n", f1 - f0, p1);
fflush(stdout);
_exit(0);
}
waitpid(pid, NULL, 0);
return 0;
}
Compiled with gcc -O2 -Wall -Wextra cow.c -o cow && ./cow (GCC 14, Linux, 4 KiB pages). Output from one run (the exact counts drift by a few between runs, for example 16,402 or 16,403 faults and Pss_Anon values a few KiB apart, because the program's own startup and printf allocate a little; the proportions do not change):
page size 4096, buffer 16384 pages = 65536 KiB
child, right after fork: Pss_Anon 32838 KiB
child, after writing every page: new minor faults 16403, Pss_Anon 65610 KiB
Right after the fork each page is mapped by two processes, so the child's proportional share is about half of 65,536 KiB (32,838 KiB; the extra 70 KiB or so is the program's other anonymous memory, and a few KiB of that varies from run to run). After writing one byte per page, the child took 16,403 faults (16,384 for the buffer plus about 19 other faults from the child's own stack, libc and output buffers) and its share doubled to the full 64 MiB: every page was copied.
Telling copy-on-write growth from a leak
| Signal | Copy-on-write growth | Leak |
|---|---|---|
| Time shape | rises after a fork, then saturates (cannot exceed the parent's resident size at fork time) | rises without bound over hours or days in one process |
| After the child exits | system memory falls back | stays high |
smaps_rollup / smaps | Shared_Dirty moves into Private_Dirty as pages are copied; Rss (resident set size) for both processes counts the same shared frames, so compare Pss | Private_Dirty and Anonymous grow steadily in one process |
| Faults | minflt (field 10 of /proc/PID/stat) jumps in step with writes, or perf stat -e minor-faults | no particular fault burst |
Sample Pss over time with grep -E 'Rss|Pss|Private_Dirty|Shared_Dirty' /proc/PID/smaps_rollup; the sum of Pss across all processes is the real memory, while the sum of Rss counts shared frames more than once.
Mitigations and interactions for fork-heavy services
- Skip the copy when you only want a new program:
vforkorposix_spawninstead offorkplusexec, so no page tables are duplicated just to be discarded. madvise(MADV_DONTNEED)(a hint telling the kernel a range's contents are no longer needed) on a private anonymous range drops the pages; later accesses see zero-filled pages. In a child that is a quick way to release memory it will never need. In the process that owns the data it destroys the data, so apply it only where discarding is correct.mlockall(a call that locks all of a process's pages in RAM so they are never swapped out) pins pages in RAM, but the man page says locks are not inherited by a forked child and warns that forking aftermlockormlockallis dangerous for real-time work: the copy-on-write faults that follow can cause high latencies. Lock memory in a process that does not fork, or pre-fault and avoid forking late.- OOM killer. The kernel scores each process by the memory it uses (adjusted by
oom_score_adj, a per-process setting that raises or lowers its score; the OOM killer is the kernel routine that kills a process when memory runs out). Pages shared copy-on-write count against each process that maps them, so a large parent can be picked even when much of its memory is shared with a child, and a growing child can push the machine into an OOM kill. WatchPss, not justRss, and size the host for the 2x worst case.
Trade-offs and pitfalls
- Copy-on-write makes
forkfast to return, not free: the cost moves to the first write of each page. - Forking a very large, write-heavy process is a poor idea; prefer a design that snapshots by copying selected data, or fork when the write rate is low.
That is every published Kernel Architecture & OS Internals question for Backend Developer so far. Browse the other topics in this category, or practice this one interactively.