Kernel Architecture & OS Internals Questions
How an operating system kernel is structured and what it is responsible for: kernel subsystems (scheduler, memory manager, drivers, interrupt handling), kernel vs. user space and the cost of crossing between them, and the boot and initialization path. Includes the core mechanisms the kernel provides: process and thread lifecycle (creation, zombies and orphans, uninterruptible sleep) and context switching, processes vs. threads as an architectural choice, CPU scheduling (CFS, priorities, preemption models, real-time policies on a general-purpose kernel), virtual memory (address translation, multi-level page tables and entry formats, TLB misses and shootdowns, page faults, copy-on-write, memory-mapped files, userfaultfd, swapping, page replacement, kernel allocators, process address-space layout, page-level protection such as NX and ASLR as mechanisms, hugepages, page coloring, overcommit and the OOM killer as kernel mechanisms), interrupts and softirqs, and how loadable modules and device drivers extend the kernel. Includes worked exercises such as page-replacement traces, address-translation arithmetic and small page-table, TLB or scheduler simulations. Covers the concepts and internals, not microcontroller interrupt and ISR design, RTOS and real-time scheduling theory, day-to-day host administration, system-call and POSIX API semantics, OS-level performance tuning, or forensic and security analysis of a host.
A machine powers on, the bootloader menu appears, but the system never reaches a login prompt and the screen is nearly silent. Given what you know of the boot stages, how do you work out which stage failed and what do you check at each?
Sample Answer
Direct answer
Seeing the bootloader menu already tells you the early stages worked: the firmware (UEFI or BIOS, the code that starts the machine) found a disk, ran the bootloader, and the bootloader drew its menu. The failure is in a later stage, and the job is to find which one by making the machine talk, then reading where its output stops. The stages after the menu are: (1) the bootloader loads the kernel image and an initramfs (a small in-memory root filesystem with early tools and drivers) into RAM and starts the kernel; (2) the kernel decompresses and initialises itself; (3) the kernel runs /init from the initramfs, which loads storage drivers, assembles RAID (several disks combined into one device) or LVM (Linux's logical volume manager, which carves flexible volumes out of disks), unlocks encryption and finds the real root filesystem; (4) the kernel switches to the real root and starts PID 1 (the first process, usually systemd); (5) PID 1 starts services until a login prompt appears. A nearly silent screen usually means the quiet option hides the messages or the console points at the wrong device, so the first move is to turn the output on.
Step 1: make it loud (one-time edit, nothing persisted)
At the bootloader menu, edit the kernel line for this boot only (in GRUB, press e) and:
- Remove
quiet(documented as disabling most log messages) and any splash option; for systemd's own messages also addplymouth.enable=0(Plymouth is the graphical boot splash; this turns it off so text messages show). - Add
ignore_loglevel, which the kernel docs describe as printing all kernel messages to the console. - Point the console at where you are looking:
console=tty0for the screen, orconsole=ttyS0,115200for a serial or virtual-machine console. The kernel docs listearlycon=andearlyprintk=for output before the normal console exists; theearlyprintkentry says it is "useful when the kernel crashes before the normal console is initialized".earlyprintk=is only available on some architectures (x86, 32-bit Arm, s390 and others); on arm64 useearlycon(with no value it takes the console from the device tree or ACPI SPCR table). - If the kernel seems to die in early init, add
initcall_debug(an initcall is one of the kernel's start-up functions, run in a fixed order during boot), which traces each initialisation function as it runs and is "useful for working out where the kernel is dying during startup".
Reboot with this edit and note the last thing printed. Where output stops identifies the stage.
Step 2: map the last message to a stage
| Last output | Stage that failed | What to check |
|---|---|---|
Nothing after the bootloader says it is loading the kernel and initramfs, or after Booting the kernel | Kernel did not start or has no working console (stage 1 or early 2) | try earlycon/earlyprintk and a different console=; boot the previous kernel entry from the menu; look for an unsupported kernel/hardware combination, a corrupted or truncated image in /boot (a full /boot during an update is a common cause), Secure Boot rejecting an unsigned kernel or module |
| Kernel messages scroll, then stop at an initcall | Stage 2, a driver or hardware init hangs | the last line from initcall_debug names the function; skip the offending driver with initcall_blacklist=<function name> (for a driver built into the kernel) or module_blacklist=<module> (for a loadable module the kernel loads); both are kernel parameters, whereas modprobe.blacklist=<module> is read by user-space modprobe and so only helps for modules loaded later, for example from the initramfs, not for a hang inside a built-in initcall; or boot a different kernel |
Messages then a wait: waiting for device ... to appear, a timeout, or a dracut/initramfs emergency prompt (dracut is the tool Red Hat-family systems use to build the initramfs) | Stage 3: initramfs cannot find or open the root device | the root= UUID or label is wrong or stale; the disk driver (NVMe, virtio, RAID controller) is missing from the initramfs; an LVM volume or encrypted device is not activated. In the emergency shell, ls /dev/disk/by-uuid, lsmod, dmesg searched for nvme, sd or virtio lines. Slow devices may need rootwait, documented as waiting indefinitely for the root device |
Kernel panic - not syncing: VFS: Unable to mount root fs on ... (the kernel source also prints Cannot open root device with a list of available partitions) | Stage 3 to 4: the kernel could not mount root | root device missing, no driver for the disk or the filesystem type built in or in the initramfs, wrong root= or rootfstype= (the filesystem type to expect on the root device, such as ext4) |
Failed to execute /init or No working init found. Try passing init= option to kernel | Initramfs or root has no runnable init | the initramfs is empty or corrupt (rebuild it); rdinit= selects a different program to run from the initramfs (the ramdisk) in place of /init, init= overrides the real-root init; the kernel tries /sbin/init, /etc/init, /bin/init, /bin/sh in turn |
Root mounted, then A start job is running for ... or a drop to emergency mode | Stage 5: a systemd unit or an /etc/fstab mount failed | boot with systemd.unit=rescue.target (a systemd option that boots to a minimal single-user mode instead of the normal target) or emergency, systemd.log_level=debug systemd.log_target=console; check journalctl -xb and fstab entries for missing devices (nofail for non-critical mounts) |
| Login prompt never appears but the system is up | Display manager or getty failure | other virtual terminal (Ctrl+Alt+F2), ssh, systemctl status of the display manager |
The kernel's built-in panic strings and init search order come from its init/do_mounts.c and init/main.c sources; panic= (with a positive number of seconds) makes a crashed machine reboot after a delay, and panic=0 waits forever, so use it to capture the message rather than losing it to an instant reboot.
Step 3: separate "new" from "old"
- If it worked yesterday, boot the previous kernel from the menu. If that boots, the new kernel or its regenerated initramfs is the culprit (missing driver, wrong module set, interrupted update).
- If an old kernel fails too, suspect the disk, the filesystem,
/etc/fstabor the bootloader configuration, not the kernel. - Boot rescue media and check
dmesgfor I/O errors, runfsckon an unmounted root, and inspect the initramfs contents (lsinitrdon dracut-based systems,lsinitramfson Debian and Ubuntu). - If the machine got far enough to write logs,
journalctl -b -1shows the previous boot, provided the journal is persistent.
Worked example
A server updated its kernel yesterday and now stops after the GRUB menu with a blank screen. Edit the entry: remove quiet, add ignore_loglevel and console=ttyS0,115200 (this server's console is a serial line; on a physical screen use console=tty0 instead). Output now reaches the initramfs and then stalls with a message of the form Warning: /dev/disk/by-uuid/... does not exist followed by Entering emergency mode (on Debian and Ubuntu the equivalent is ALERT! /dev/disk/by-uuid/... does not exist. Dropping to a shell!). The previous kernel boots fine. In the emergency shell ls /dev/disk/by-uuid is empty and lsmod shows no nvme or virtio_blk. Diagnosis: stage 3, the new initramfs lacks the storage driver (the disk controller module), so /init ran but never saw the root disk. Fix: boot the old kernel, regenerate the initramfs, confirm the driver is listed inside it (lsinitrd | grep nvme), and test the new entry before deleting the old one.
Variant to tell apart: if the screen instead ends at Kernel panic - not syncing: VFS: Unable to mount root fs on unknown-block(0,0) (block device number 0,0, meaning no real disk was identified as root), the kernel never got a working initramfs to run, typically because the new image is missing, zero-length or truncated after an interrupted update or a full /boot, or because the initrd line is absent from the bootloader entry. The fix path is the same (boot the old entry, rebuild the initramfs) but the check is different: compare the image's size in /boot and the entry's initrd line, not the driver list.
Trade-offs and pitfalls
- Debug options such as
ignore_loglevelandinitcall_debugare noisy and slow a boot; use them to find the stage, then remove them. - Do not start by reinstalling the bootloader. If the menu appears, the bootloader works.
- Recovery commands (rescue shells, repairing GRUB) matter only after the failing stage is known; naming the stage first prevents fixing the wrong layer.
Describe how the kernel hands out memory to itself. What are the buddy allocator and the slab family for, how do kmalloc and vmalloc differ, and how does fragmentation show up and hurt?
Sample Answer
The problem. The kernel cannot call a C library malloc; it is the thing that hands out memory. It needs memory for its own data (process descriptors, file and network structures, buffers) in sizes from a few dozen bytes to megabytes, often from contexts that must not sleep. Two layers solve this: a page-level allocator underneath and an object-level allocator on top.
Layer 1: the buddy allocator (whole pages). Physical memory is divided into pages (4 KiB on x86-64). The buddy allocator tracks free memory as blocks of 2^order contiguous pages: order 0 is one page, order 3 is eight pages (32 KiB), and so on. To allocate an order-3 block, it takes a free order-3 block, or splits an order-4 block into two order-3 'buddies', using one and keeping the other free. When a block is freed, if its buddy (the adjacent same-size block) is also free, the two merge into the next order. Splitting and merging makes both operations fast and keeps contiguous space available as long as frees are not scattered.
A traced example: suppose the only free memory is one order-4 block (16 pages). Allocating one page (order 0) splits order 4 into two order-3 blocks (keep one free, split the other), that into two order-2 blocks, then two order-1, then two order-0. The free lists now hold one block each of order 3, 2, 1 and 0, which is 8 + 4 + 2 + 1 = 15 pages, plus the 1 page handed out, 16 in all. Freeing that page makes it merge with its free order-0 buddy, then order 1 with its buddy, order 2 with its, and order 3 with its, ending as the single order-4 block again. If instead a stray long-lived allocation sat in one of those pages, the merge would stop there and no order-4 block could be rebuilt: that is fragmentation.
Layer 2: the slab family (small objects). Asking the buddy allocator for a 4 KiB page to hold a 200-byte object wastes most of it. A slab allocator (SLUB is the Linux implementation) takes pages from the buddy allocator and carves each into many equal-size objects. A cache (a kmem_cache) serves one object type or size: for example one cache for a process descriptor, one for file metadata. Benefits: no per-object page waste, objects of one type packed together, freed objects are reused quickly (warm in cache) typically from per-CPU free lists (a private list of free objects for each CPU, so taking one needs no shared lock) with little locking, and a type's memory can be tracked and reclaimed separately. 'The slab family' means the allocator implementations that share this model (the kernel documentation for monitoring caches describes the SLUB implementation).
kmalloc vs vmalloc
kmalloc | vmalloc | |
|---|---|---|
| Virtual contiguity | yes | yes |
| Physical contiguity | yes | no (kernel docs: not physically contiguous) |
| Backed by | slab caches of fixed size classes (requests are rounded up to a class) | individual pages, mapped into a separate virtual range |
| Cost | fast, no page-table building per call | slower: it must assemble many separate pages into one virtual range, and a range built from 4 KiB pages needs one TLB entry per page |
| Size | meant for objects smaller than a page; practical limit is bounded | for large allocations |
| Hardware use | usable for DMA buffers | not usable by a device that needs one contiguous physical range |
Rule of thumb: use kmalloc for small objects and anything a device will access by physical address; use vmalloc for big software-only buffers (for example a large table) where physical contiguity is not needed; kvmalloc (a helper that tries kmalloc first and falls back to vmalloc, which is the right default when the size is unknown but the memory is only touched by the CPU). The GFP flags (get-free-page flags, passed with every allocation) say how the allocator may behave: GFP_KERNEL may sleep and reclaim memory (so not allowed in interrupt context, the code that runs in answer to a hardware interrupt or a softirq, a deferred software interrupt, where there is no process that can wait), while GFP_ATOMIC never sleeps and may dip into reserves, so it fails more readily. | ||
DMA implication. DMA (direct memory access) lets a device read or write RAM without the CPU. A device that does not support scatter-gather (the ability to read or write a list of separate memory pages as one transfer) needs one physically contiguous buffer, so you need kmalloc-style memory, and for old devices that can only address low memory, flags such as GFP_DMA32 restrict where it comes from (to the first 4 GiB of physical addresses, which a device with 32-bit addressing can reach). Physical memory is also split into zones, such as DMA (the first 16 MiB on x86-64), DMA32 and Normal, which are address ranges with different usage rules; the buddy allocator keeps separate free lists per zone. A vmalloc buffer can look contiguous to the CPU while its pages are scattered in RAM. |
How fragmentation shows up
- External fragmentation: plenty of free pages, but scattered, so no free block of the needed order exists.
/proc/buddyinfoprints the count of free blocks at each order per zone; a healthy host has nonzero counts in the high columns, a fragmented one has large counts in the first columns and zeros at the end. Illustrative output in the format of/proc/buddyinfo(the counts are invented for this example and the small DMA zone's row is omitted; a real machine will differ):
Node 0, zone DMA32 1 0 1 1 1 2 2 1 1 1 500
Node 0, zone Normal 86 619 7301 11504 5665 3437 2113 1324 739 516 1873
Each row is a zone; the 11 numbers are counts of free blocks of order 0 through 10 (order 10 = 1,024 pages = 4 MiB). The Normal zone has 1,873 free 4 MiB blocks, so a large contiguous request succeeds: this is a healthy host. The free memory in a column is count x 2^order x 4 KiB, for example order 5 is 3,437 x 32 pages x 4 KiB = 429.6 MiB, and the Normal row's columns together come to about 11,716 MiB (the DMA32 row's 500 order-10 blocks add about 2,000 MiB more). A fragmented host with the same total would show most of that memory in the left columns and 0 in the right ones, so a request for order 8 or above would fail or force compaction even though free reports gigabytes. Symptoms: allocation failures or stalls for high-order requests, transparent huge page fallbacks, compact_stall counts rising in /proc/vmstat. The fix is to avoid high-order allocations (use kvmalloc or scatter-gather) and let compaction (moving pages to join free space) run.
- Internal fragmentation: waste inside the allocation. A
kmallocrequest is rounded up to its size class, and slab pages can be mostly empty (a few live objects pin a whole page)./proc/slabinfo(which needs root) shows active versus total objects per cache. Its columns arename active_objs num_objs objsize objperslab pagesperslab. An illustrativedentryrow (dentries are the kernel's cached directory entries, mapping path names to files):dentry 50000 100000 192 21 1 ...means 100,000 object slots of 192 bytes held in memory (100,000 x 192 B = 18.3 MiB) of which only 50,000 are in use (9.2 MiB), so half the cache's memory is unused slack that only frees when whole slab pages empty out;/proc/meminfosplitsSlabintoSReclaimable(caches that can be dropped under pressure) andSUnreclaim(cannot).
Why it hurts: a slab page cannot return to the buddy allocator until every object in it is freed, so long-lived stragglers keep memory unavailable for large requests even though 'free memory' looks fine.
Closing example: a high-performance network stack. Receive and transmit paths allocate and free a packet buffer descriptor millions of times per second, often from interrupt or softirq context where sleeping is forbidden. This is exactly what dedicated slab caches with per-CPU free lists and non-sleeping flags are for; designs on top of this typically recycle buffers instead of freeing them, so the hot path avoids the allocator altogether. The cost of getting it wrong is visible as GFP_ATOMIC failures under bursts, which the stack answers by dropping packets. Concretely: at 1 million packets per second, a 1 ms burst needs about 1,000 buffer descriptors from non-sleeping allocations; if the per-CPU lists and reserves hold fewer, the rest are dropped and the drop counters rise.
What is a loadable kernel module and how does it differ from code built into the kernel? What does it mean for a kernel to be tainted, and what are the stability and security consequences of loading third-party modules?
Sample Answer
Direct answer
A loadable kernel module is a piece of kernel code compiled separately as a .ko file (an ELF object, ELF being the standard Linux executable and object file format) that can be inserted into and removed from a running kernel with insmod, modprobe and rmmod. Once loaded it is part of the kernel: same privilege, same address space, same crash domain. Code built into the kernel image (=y in the configuration) is always present from boot; a module (=m) is loaded on demand. A kernel is "tainted" when something has happened that makes its behaviour harder to trust or debug, such as loading a proprietary, out-of-tree or unsigned module; the state is a bitmask (one number whose individual binary digits are separate yes/no flags) readable from /proc/sys/kernel/tainted, and a nonzero value tells maintainers to be cautious with bug reports. Third-party modules carry stability risk (a bug panics the whole machine) and security risk (a malicious module is a rootkit (software that hides an intruder's presence) with full power), which signing and policy are meant to control.
Module versus built-in
| Built into the image | Loadable module | |
|---|---|---|
| Present | From boot, always | After insmod or modprobe, or automatic load |
| Needed for boot | Yes if it is required before the root filesystem can be mounted (unless it is in the initramfs) | Only if packaged in the initramfs |
| Memory | Always resident | Only when loaded; can be unloaded if nothing uses it |
| Trust | Covered by the signature of the kernel image | Verified separately if module signing is enabled |
| Update | New kernel and reboot | Replace the file and reload (no reboot, if unloadable) |
| ABI (the binary-level contract between compiled code and the kernel: structure layouts, function signatures) | Compiled together with the kernel | Must match the kernel it was built for |
The kernel has no stable internal interface, so a module must be built against the exact kernel (or compatible headers) it will run on: the module records a version string and, with CONFIG_MODVERSIONS, checksums of the kernel symbols it uses. That is why out-of-tree drivers are rebuilt by tools such as DKMS (Dynamic Kernel Module Support, which recompiles registered modules automatically when a new kernel is installed) on every kernel update. Modules can only use symbols the kernel exports; some are exported to GPL-compatible modules only, so a module without a GPL-compatible MODULE_LICENSE is treated as proprietary.
How loading happens: the loading process calls finit_module (or init_module), a system call that requires the CAP_SYS_MODULE capability (the Linux privilege bit that permits loading kernel modules, normally held only by root). The kernel loads the ELF image into kernel memory, resolves its symbols, checks the signature if enforcement is on, and runs its init function. At boot, modules come in three ways: from the initramfs (the early root image carrying the storage and filesystem drivers needed to mount the real root), from udev (the device manager daemon) loading by hardware alias (an identifier string a device advertises, matched to a module) when a device is discovered, and from lists in /etc/modules-load.d. Automatic loading can be suppressed with a modprobe blacklist or the modprobe.blacklist= boot parameter, but both are read by user-space modprobe, and a blacklist entry only makes modprobe ignore the module's internal hardware aliases (so alias-based loading by udev skips it): an explicit modprobe name or insmod can still load it. To refuse a module outright, use the kernel's own module_blacklist= boot parameter (read by the kernel itself: a comma-separated list of module names that the load path then rejects, whichever tool asked), or an install name /bin/false line in a modprobe.d file for loads that go through modprobe.
What taint means
Each cause sets a bit; the kernel documentation lists the flags. The ones that matter for modules are bit 0 (letter P, a proprietary module was loaded), bit 1 (F, a module was force-loaded), bit 12 (O, an externally built "out-of-tree" module was loaded) and bit 13 (E, an unsigned module was loaded); bit 7 (D) means the kernel has oopsed (hit a recoverable internal error and printed an "oops" report, as with a bad pointer dereference), and bit 9 (W) means it issued a warning. Decoding a value by hand:
12289=8192+4096+1=213+212+20
so bits 13, 12 and 0 are set: unsigned (E), out-of-tree (O) and proprietary (P). As a program:
FLAGS = {0: "P proprietary module loaded", 12: "O out-of-tree module loaded", 13: "E unsigned module loaded",
1: "F module force-loaded", 7: "D kernel oopsed or BUG hit", 9: "W kernel warning issued"}
def decode(value):
return [FLAGS.get(b, f"bit {b}") for b in range(19) if value >> b & 1]
for v in (0, 4097, 12289):
print(v, decode(v))
print(2**13 + 2**12 + 2**0)
Output:
0 []
4097 ['P proprietary module loaded', 'O out-of-tree module loaded']
12289 ['P proprietary module loaded', 'O out-of-tree module loaded', 'E unsigned module loaded']
12289
Taint is sticky: it stays until reboot even after the module is unloaded, because the damage a bad module could have done to kernel memory does not go away. Upstream developers generally treat reports from kernels tainted by proprietary modules with suspicion because they cannot inspect that code. Taint is information, not a block: the kernel keeps running.
Stability and security consequences
The two risks call for different controls: stability comes from where the module came from and how it is tested; security comes from signing, the one-way sysctl and who holds CAP_SYS_MODULE. The bullets below follow that split.
- Stability. There is no isolation. A null pointer or a missed lock in a third-party module can corrupt any kernel structure or panic the machine, and the fault can show up far from the cause. A module also ties you to a kernel version: every kernel upgrade means a rebuild and a test, and a GPU or storage driver that lags a release blocks the upgrade.
- Security. A loaded module runs at the highest privilege and can hide processes, hook system calls or read any memory, which is how kernel rootkits work. Controls: module signing (
CONFIG_MODULE_SIG; withoutCONFIG_MODULE_SIG_FORCEor themodule.sig_enforce=1boot parameter an unsigned module still loads but taints the kernel with E, with enforcement only validly signed modules load); the sysctlkernel.modules_disabled=1, which is one-way: after it is set, modules can be neither loaded nor unloaded until reboot; and restricting who holds CAP_SYS_MODULE. - Secure Boot. Secure Boot verifies the bootloader and kernel signature, and distribution kernels commonly extend that by refusing unsigned modules while it is on. To run a third-party driver such as an out-of-tree GPU or virtualization module you then sign it with your own key and enrol that key with the firmware's machine-owner-key (MOK) mechanism, using the
mokutiltool to queue the key for enrolment at the next boot, rather than disabling Secure Boot.
Recommendation
Prefer drivers that are in the mainline kernel, load only what the machine needs, enable signing with enforcement on production fleets, pin third-party modules to tested kernel versions, and record /proc/sys/kernel/tainted in your monitoring so a tainted host is visible before a crash report is filed. What would change this: if a vendor module is the only way to use the hardware, keep it on a fixed kernel and put the rebuild-and-test step into your upgrade pipeline.
During an incident many processes are stuck in the D state, a key service is unresponsive, and CPU usage is low. What does uninterruptible sleep mean in the kernel, why can such tasks not be killed, and what would you check next to find what they are waiting on?
Sample Answer
What the D state means. Every task (a process or thread) the kernel schedules has a state. R is running or runnable, S is an interruptible sleep (waiting for something, and a signal can wake it), and D is uninterruptible sleep (the kernel flag is TASK_UNINTERRUPTIBLE; ps documents it as "uninterruptible sleep (usually I/O)" and proc(5) as "uninterruptible disk sleep"). A task in D has told the scheduler "do not run me until the event I am waiting for happens". It is off the CPU, so it uses no CPU time. That is why your service is dead while CPU usage is low.
Why it cannot be killed. A signal, including SIGKILL, is only acted on when the task is woken and about to return to user space. An uninterruptible sleep is not woken by signals, so kill -9 just leaves SIGKILL pending. It is delivered the moment the awaited event arrives, and never if the event never arrives. The design exists because kernel code in the middle of an operation (holding a lock, with a disk or network request in flight that will write into a buffer when it completes, halfway through a filesystem update) often cannot safely abort. Applications also assume that file I/O is not interrupted by signals. Linux added a middle state, TASK_KILLABLE (merged in 2.6.25), where only a fatal signal ends the sleep. The NFS client was the first code converted, but any wait that still uses the plain uninterruptible sleep stays unkillable. Two consequences: the task is not a zombie (Z, already dead, waiting for its parent to collect it), and for a plain uninterruptible wait (a request to a hung disk or a stuck driver) you cannot fix this by killing it; you must make the wait end. The wait name can tell you which kind you have: a wchan ending in _killable (such as the NFS ones below) means kill -9 will work even though ps shows D (a killable wait need not carry that suffix; the kernel_clone wait in the demo below is killable too). That matters for NFS: since 2.6.25 a hard-mounted NFS task waiting on a dead server sleeps in the killable form (nfs_wait_bit_killable and rpc_wait_bit_killable return -ERESTARTSYS when a fatal signal is pending in kernel source), so kill -9 normally clears those, although the mount itself stays hung and new tasks that touch it will pile up again.
Why load is high with idle CPUs. Linux's load average counts tasks that are runnable (R) plus tasks in D (the kernel adds the running and uninterruptible task counts in kernel/sched/loadavg.c). Fifty tasks stuck in D push load to 50 while the CPU sits idle. A high load with low CPU utilisation is a strong hint that you have a D-state pile-up, not a compute problem.
What I would check, in order.
- Count and identify the stuck tasks, and whether they share a wait location:
ps -eo pid,stat,wchan:32,comm | awk 'NR==1 || $2 ~ /^D/'
wchan (wait channel) is the name of the kernel function where the task sleeps. If forty tasks all show the same NFS or block-layer function (the block layer is the kernel code between file systems and storage devices), you have one root cause. Illustrative output (not from a run) of a pile-up behind a dead NFS server; these waits are killable, so kill -9 would remove those processes, while the same pile-up behind a hung disk would show block-layer or driver wait functions that ignore it:
PID STAT WCHAN COMMAND
2210 D nfs_wait_bit_killable java
2214 D nfs_wait_bit_killable java
2301 D rpc_wait_bit_killable rsync
... 37 more lines, 36 of them nfs_wait_bit_killable ...
One function name repeated across nearly every line says the tasks are waiting on one thing, so the next step is that thing (the NFS mount), not each process.
- Get the full kernel stack of one or two of them (needs root, and a kernel built with CONFIG_STACKTRACE, the option that lets the kernel record call stacks, per proc_pid_stack(5)):
cat /proc/<pid>/wchan; echo
cat /proc/<pid>/stack
Read it bottom to top: the bottom is the system call the application made (read, write, fsync, open), the top is where it is parked. Frames like nfs_* or rpc_* mean an NFS server or network stall, io_schedule (the kernel function that puts a task to sleep until its I/O finishes) under a filesystem function means waiting for a block device, a driver function means a hung device. A stuck NFS read might look like this illustrative stack (top frame first):
[<0>] rpc_wait_bit_killable+0x1c/0x60
[<0>] nfs_wait_on_request+0x40/0x90
[<0>] nfs_file_read+0x88/0xd0
[<0>] vfs_read+0x9c/0x1a0
[<0>] __x64_sys_read+0x1a/0x30
The bottom line is the read system call, the middle shows the NFS client, and the top shows it parked waiting on the RPC (remote procedure call) reply from the server.
- Dump all blocked tasks at once, including stacks, to the kernel log:
echo w > /proc/sysrq-trigger # sysrq ('magic SysRq', a kernel debugging interface, needs CONFIG_MAGIC_SYSRQ) 'w' = dump tasks in uninterruptible state
dmesg | tail -200
Also search dmesg for "blocked for more than ... seconds". The hung-task detector (a kernel watchdog thread that periodically looks for tasks stuck in D; it exists only in kernels built with CONFIG_DETECT_HUNG_TASK, which most distribution kernels enable) prints that warning with a stack for a task stuck in D longer than kernel.hung_task_timeout_secs (check your value with sysctl; 0 disables it). It skips tasks in the killable form of D (a comment in kernel/hung_task.c reads "skip the TASK_KILLABLE tasks -- these can be killed"), so a pile-up behind a hard NFS mount, whose waits are killable, can raise the load average without any such warning, while a plain uninterruptible wait (a hung disk or driver) does produce one.
-
Follow the stack to the resource. For storage:
iostat -x 1(is one device at 100% utilisation with a growing queue,aqu-sz, or is it at zero throughput because requests are not completing? Illustrative bad sign:%util100 withr/sandw/snear 0 andaqu-szlarge),dmesgfor SCSI/NVMe resets, timeouts and multipath path failures (multipath means one disk reachable over several network or cable paths, and a failure of all paths stalls I/O), andlsblk/cat /proc/mountsfor the device. For NFS:grep nfs /proc/mountsto seehardorsoft, then check whether the server answers (ping;rpcinfo -p <server>, which lists the server's RPC services and hangs or errors if it is unreachable;nfsiostat, which shows per-mount operation counts and round-trip times, so a mount with requests outstanding and no replies is stalled). With the defaulthardmount (retry forever; the alternativesoftgives up and returns an error after a timeout), NFS requests are retried indefinitely (nfs(5)), so a dead server freezes every task that touches the mount, silently, until it returns. For a driver hang:dmesgaround the first error, and the stack's top function names the driver. -
Find the common trigger: what changed just before the first stuck task (a storage failover, a network switch change, a server reboot, a full RAID rebuild (RAID: several disks combined into one volume), a new kernel module)? Look at the timestamp of the oldest
blocked for more thanmessage, when the waits are of the kind that produces one.
Fix. Restore the thing being waited on (bring the NFS server or the network path back, fix or fail over the storage path, reset the device). As soon as the awaited event completes, the tasks wake and any pending SIGKILLs are acted on, so stuck tasks usually disappear on their own. For tasks stuck on a dead NFS server specifically, kill -9 on the killable waits frees the processes without waiting for the server, but the mount keeps hanging every new access until it is fixed or detached. If the server is gone for good, a forced or lazy unmount (umount -f forces it; umount -l detaches the mount point immediately and cleans up later) can detach the mount, but be careful: it can lose unwritten data, and a task already blocked in the kernel may stay stuck. A reboot is the last resort. A soft NFS mount avoids the freeze but, as nfs(5) warns, can cause silent data corruption, so I would choose it only where responsiveness matters more than data integrity. For the service itself, put timeouts and a health check that does not touch the suspect mount, so the load balancer drains the node while you repair it.
A reproducible D state. This program parks itself in a killable D wait with vfork() (a variant of fork where the parent is suspended until the child execs or exits, because the child borrows the parent's memory) and has a monitor process look at it from outside. It is the easiest D state to produce safely on a laptop; it is a stand-in for the mechanics, and real incidents look like the NFS examples above.
// dstate.c: park a process in a killable D-state wait and look at it from outside.
// Build: gcc -O2 -Wall -Wextra dstate.c -o dstate Run as root (needed for /proc/<pid>/stack)
#define _GNU_SOURCE
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/types.h>
#include <sys/wait.h>
int main(void) {
char pid[32];
snprintf(pid, sizeof pid, "%d", (int)getpid());
if (fork() == 0) { // monitor: inspects the parent while it is stuck
sleep(1);
execl("/bin/sh", "sh", "-c",
"ps -o pid,stat,wchan:32,comm -p $0; grep State /proc/$0/status; "
"echo \"wchan: $(cat /proc/$0/wchan)\"; head -3 /proc/$0/stack; "
"kill -KILL $0; sleep 0.2; ps -p $0 >/dev/null && echo still-there || echo gone-after-SIGKILL",
pid, (char *)NULL);
_exit(127);
}
if (vfork() == 0) { // the parent is suspended until this child execs or exits
sleep(5);
_exit(0);
}
return 0;
}
Built and run as root in a --privileged gcc:14 container on arm64 (/proc/<pid>/stack needs root and CONFIG_STACKTRACE; without root the head line fails with Permission denied while the first three views still work), it prints the following, apart from the PID and a Killed notice from the shell:
PID STAT WCHAN COMMAND
15 D kernel_clone dstate
State: D (disk sleep)
wchan: kernel_clone
[<0>] kernel_clone+0x1a0/0x430
[<0>] __arm64_sys_clone+0x64/0x98
[<0>] do_el0_svc+0x70/0xd0
gone-after-SIGKILL
This shows what the three views look like (ps STAT D, /proc/<pid>/status "disk sleep", the stack with the system call at the bottom). Note that "disk sleep" is just the label for D and this wait involved no disk. This particular wait is killable, so kill -KILL removes the process at once (gone-after-SIGKILL). A true TASK_UNINTERRUPTIBLE wait, such as a request outstanding to a hung disk or a driver stuck in the kernel, would ignore it, and a dead NFS server is not reproduced here.
Explain Linux memory overcommit and the vm.overcommit_memory modes. Why can a process be killed by the OOM killer or its cgroup limit even though its resident size looks modest, and what does this mean for containers?
Sample Answer
Direct answer
With overcommit, Linux lets a process reserve (commit) more virtual address space than there is RAM plus swap, because most reserved memory is never touched. A call like malloc or mmap therefore succeeds at once, and the real memory is only taken page by page when the program first touches each page (a page fault: the CPU traps to the kernel, which finds a physical frame and maps it). If the machine or the container's cgroup (control group: the kernel's resource-limit group) has no frame left at that moment, there is no error code to return to the faulting instruction, so the kernel runs the OOM killer (out-of-memory killer) and kills a process. That is why a process can die at a line of code that looks harmless, long after the allocation call that "succeeded". Containers matter because the overcommit mode is machine-wide, while the limit that kills a container is its cgroup's memory.max, and that limit counts more than the number ps calls resident size.
The three modes of vm.overcommit_memory
(From the kernel's overcommit documentation.)
| Value | Name | Behaviour |
|---|---|---|
| 0 | Heuristic (default) | Obvious overcommits of address space are refused: in the kernel source (__vm_enough_memory), a single request larger than total RAM plus swap fails. Everything else is allowed |
| 1 | Always | No check at allocation time; the sparse-array style programs that rely on mostly-zero memory are the stated use |
| 2 | Never (strict) | Total address-space commit may not exceed swap plus a percentage of physical RAM (vm.overcommit_ratio, default 50), or swap plus a fixed amount if vm.overcommit_kbytes is set |
For a concrete case in mode 0, on a machine with 16 GiB of RAM and 2 GiB of swap, one mmap of 20 GiB is refused (20 > 18) but a 17 GiB request is allowed, and so are several 10 GiB requests in a row, because each is checked alone and only the size of that one request is compared. That gap is the overcommit. Mode 1 would allow the 20 GiB request too.
Commit accounting. The kernel tracks Committed_AS: the sum of address space that the system has promised to processes. In mode 2 it compares that to CommitLimit and makes mmap, brk or fork return ENOMEM up front. CommitLimit is swap plus RAM times the ratio. Example from the Docker VM used for the runs below (values in kB, host-specific; read yours with grep -E 'MemTotal|SwapTotal|CommitLimit' /proc/meminfo and cat /proc/sys/vm/overcommit_ratio; the kernel also subtracts huge pages reserved for hugetlb from the RAM term):
CommitLimit = SwapTotal + MemTotal x 50 / 100
= 17,462,192 + 16,413,624 x 0.50
= 17,462,192 + 8,206,812 = 25,669,004
which matches the CommitLimit: 25669004 kB that /proc/meminfo printed there. The limit is only enforced in mode 2; in modes 0 and 1 it is informational.
Why allocation succeeds and the kill comes later
This script reserves 2 GiB of address space, touches only 64 MiB, and then (given a second argument) touches everything:
import mmap, os, sys
def status(label):
f = {l.split(":")[0]: l.split(":")[1].strip() for l in open("/proc/self/status") if l.startswith(("VmSize", "VmRSS"))}
print(f"{label:<34} VmSize={f['VmSize']:>12} VmRSS={f['VmRSS']:>10}", flush=True)
print("overcommit_memory =", open("/proc/sys/vm/overcommit_memory").read().strip())
status("start")
big = mmap.mmap(-1, 2 * 1024**3, flags=mmap.MAP_PRIVATE | mmap.MAP_ANONYMOUS) # reserve 2 GiB of address space
status("after mmap 2 GiB (untouched)")
mb = int(sys.argv[1])
for off in range(0, mb * 1024**2, 4096): # touch only mb MiB
big[off] = 1
status(f"after touching {mb} MiB")
if len(sys.argv) > 2: # keep going until the limit bites
for off in range(mb * 1024**2, 2 * 1024**3, 4096):
big[off] = 1
status("touched all 2 GiB")
Run in a python:3.12-slim container with no limit (python overcommit.py 64), then with a 256 MiB cgroup limit (docker run --rm -m 256m --memory-swap 256m ... sh -c 'python overcommit.py 64 all; echo "exit status=$?"; cat /sys/fs/cgroup/memory.max; grep -E "^(oom|oom_kill) " /sys/fs/cgroup/memory.events'):
--- no limit, touch 64 MiB
overcommit_memory = 1
start VmSize= 13704 kB VmRSS= 9612 kB
after mmap 2 GiB (untouched) VmSize= 2110856 kB VmRSS= 9676 kB
after touching 64 MiB VmSize= 2110856 kB VmRSS= 75332 kB
--- memory limit 256m, touch everything
overcommit_memory = 1
start VmSize= 13704 kB VmRSS= 9640 kB
after mmap 2 GiB (untouched) VmSize= 2110856 kB VmRSS= 9704 kB
after touching 64 MiB VmSize= 2110856 kB VmRSS= 75360 kB
Killed
exit status=137
268435456
oom 1
oom_kill 1
The command pieces: -m 256m caps the container's memory at 256 MiB (this becomes memory.max); --memory-swap 256m sets memory plus swap to the same figure, so no swap is allowed and the limit is hard; sh -c '...' runs several commands in order inside the container; echo "exit status=$?" prints the exit code of the Python run; cat /sys/fs/cgroup/memory.max prints the limit in bytes; and the grep pulls the oom and oom_kill counters out of memory.events, a file of counters the kernel keeps per cgroup.
Read it like this: VmSize (virtual size, the address space reserved) jumped by about 2 GiB at mmap while VmRSS (resident set size, the pages actually in RAM) barely moved. The mmap of 2 GiB succeeded inside a 256 MiB container because the limit counts used pages, not reserved ones. The process was killed (exit status 137 = 128 + signal 9) only when touching pushed usage to memory.max = 268,435,456 bytes, and memory.events recorded oom 1 and oom_kill 1. The Docker VM used for these runs was in mode 1, so this run shows mode 1 plus a cgroup limit, not the default heuristic mode.
Why the resident size can look modest and the process still dies
memory.max limits memory.current, which per the cgroup v2 documentation includes more than anonymous memory: page cache (file contents the kernel keeps in RAM), kernel data structures such as dentries (cached directory-entry lookups) and inodes (the kernel's per-file records), and network socket buffers (queues of data waiting to be sent or read). Things ps RSS for one process does not show:
- Other processes in the same cgroup. The limit is shared by every process in the container, including sidecars, shells and forked workers.
- Page cache and tmpfs. tmpfs is a file system that lives entirely in RAM, and
/dev/shmis a tmpfs directory used for shared memory between processes. Files written, files read, and anything under/dev/shmor a tmpfs mount are charged to the cgroup. Clean cache can be reclaimed, but dirty or shared-memory pages cannot be dropped quickly. - Kernel memory: page tables for a large mapping, slab objects (small fixed-size kernel allocations such as the dentries and inodes above), socket buffers.
- A small numeric case. A container with a 512 MiB limit runs one process with 200 MiB RSS, but it also wrote a 250 MiB log file (250 MiB of page cache) and holds 30 MiB of tmpfs and socket buffers.
memory.currentis 200 + 250 + 30 = 480 MiB. The clean 250 MiB of cache can be reclaimed, so this is survivable, and even dirty log pages can be reclaimed once they are written back (reclaim stalls the allocating thread meanwhile). What cannot be reclaimed is tmpfs and shared memory when there is no swap, plus kernel memory: if the tmpfs share grows past what the cache reclaim can make room for, the kernel kills a process even thoughpsstill shows 200 MiB. - Memory that was not resident when you looked: a spike between samples, or pages touched after your last measurement.
Also, the OOM killer picks a victim by its OOM score (a badness number the kernel computes per process, mainly from how much memory it uses, adjustable through oom_score_adj), so with several processes in a cgroup the one killed is not necessarily the one that grew. By default a cgroup OOM kills one process, not the whole group, unless memory.oom.group is set to 1, which makes the kernel kill every process in that cgroup together.
What this means for containers
- Set the memory limit from measured
memory.currentat peak load plus headroom, not from the process RSS. Watchmemory.eventsforoom_kill, and know thatmemory.highthrottles with heavy reclaim but never invokes the OOM killer, so it makes a good soft ceiling belowmemory.max. Of the cgroup files,memory.max(the hard limit) andmemory.events(the evidence) are the two to know first;memory.highandmemory.oom.groupare tuning. - Managed runtimes (the JVM, Go, Node) need their heap caps set from the container limit, not the machine's RAM, or they size themselves for the host and are killed at the cgroup line.
- Do not rely on
vm.overcommit_memory=2to protect one container: it is machine-wide and makes unrelatedforkandmmapcalls fail (a large process forking needs a second copy's worth of commit even though copy-on-write would share it). Use cgroup limits for per-container isolation and overcommit settings for the host's global policy. - Exit status 137 and an event in
memory.eventsare the evidence to look for when a container "just disappears".
Unlock Full Question Bank
Get access to all 49 Kernel Architecture & OS Internals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.