Assembly and Low-Level Language Fundamentals Questions
Programming and debugging at the instruction level. Covers reading and writing assembly for x86-64 and ARM (A32, Thumb, AArch64), including small hand-written routines, vector (SIMD) code and exclusive load/store atomics, registers, the stack and frame layout, prologues and epilogues, calling conventions and ABIs (System V, Microsoft x64, AAPCS), variadic calls, inline assembly and its constraints and clobbers, Cortex-M exception entry and context-switch code written in assembly, and how compilers translate and optimize source into machine code (optimization flags, inlining, tail calls, LTO, aliasing, strength reduction, virtual dispatch, memcpy lowering, stack spills). Also covers object files and linking as they affect generated code (ELF, PE/COFF and Mach-O, relocations, GOT and PLT, position-independent code, static linking, stack unwinding), the compiler backend ideas behind it (SSA, register allocation, instruction selection, peephole passes) and emitting machine code from a minimal JIT. On the debugging side: reading disassembly, using gdb and lldb for registers, frames, breakpoints and watchpoints, analyzing core dumps and stripped binaries with addr2line and build IDs, and diagnosing crashes from instruction-level state such as corrupted returns, stack smashing, ABI mismatches and optimizer-induced bugs. Debugging method in general, hardware probe tooling, malware analysis and exploit-mitigation design are covered elsewhere.
In a hardened binary, how would you determine whether a crash is caused by stack smashing, heap corruption, use-after-free, or a corrupted return address using only the assembly view, register contents, memory layout, and debugger evidence? Describe the distinctions you would look for in each case.
Sample Answer
Direct answer
In a hardened build (stack canary, _FORTIFY_SOURCE (compiler and libc checks that abort when a copy would overflow a buffer of known size), PIE (position-independent executable, loaded at a random address), full RELRO (the dynamic-linking table, the GOT, is made read-only after startup), a checking allocator) many corruptions no longer crash where they happen: they are detected by a check and turned into a deliberate abort, so the first discriminator is how the process died, not where the pc is. Read in this order: (1) the signal (6 SIGABRT, the signal abort() raises, means a runtime check fired; 11 SIGSEGV or 7 SIGBUS means a wild access or jump), (2) the top frames (a named check function such as __stack_chk_fail, __chk_fail or free, or the program's own code, or no valid frame at all), (3) the abort message the runtime saved, (4) whether the crashing thread's stack still unwinds, (5) what the faulting pointer or pc value is made of. A canary (a random value placed between locals and the saved return address, checked before return) is the thing that makes the stack-smash and corrupted-return cases look different.
If you keep only three signatures, keep these: a named check function (__stack_chk_fail, __chk_fail, an allocator abort inside free) with SIGABRT means the runtime caught it; a pc and link register made of the same fill bytes, no check frame and an intact canary means the return slot alone was overwritten; a pc or register holding a small number or a shifted heap address with a valid caller means a freed object was used. The full table follows the five cases.
The test program
One program with one defect per mode, built hardened and crashed on purpose:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
struct handler { void (*fn)(void); long id; };
static void hello(void) { puts("hello"); }
__attribute__((noinline)) static void stack_loop(volatile int n) {
char buf[16];
for (int i = 0; i < n; i++) buf[i] = 'A';
printf("%c\n", buf[0]);
}
__attribute__((noinline)) static void stack_memcpy(const char *src, size_t n) {
char buf[16];
memcpy(buf, src, n);
printf("%c\n", buf[0]);
}
__attribute__((noinline)) static void heap_overflow(volatile int n) {
volatile char *a = malloc(24);
volatile char *b = malloc(24);
b[0] = 'x';
for (int i = 0; i < n; i++) a[i] = 'B';
free((void *)b);
free((void *)a);
}
__attribute__((noinline)) static void use_after_free(void) {
struct handler *h = malloc(sizeof *h);
h->fn = hello;
h->id = 1;
free(h);
h->fn();
}
__attribute__((noinline)) static void indexed_write(volatile long idx) {
volatile long slot[2] = {0, 0};
slot[idx] = 0x4242424242424240L; /* idx is attacker-controlled and unchecked */
printf("%ld\n", slot[0]);
}
int main(int argc, char **argv) {
int mode = argc > 1 ? atoi(argv[1]) : -1;
char big[256];
memset(big, 'A', sizeof big);
switch (mode) {
case 0: stack_loop(64); break;
case 1: stack_memcpy(big, 64); break;
case 2: heap_overflow(40); break;
case 3: use_after_free(); break;
case 4: indexed_write(atoi(argv[2])); break;
}
return 0;
}
Saved as crash_modes.c and built with gcc -O1 -g -fstack-protector-strong -D_FORTIFY_SOURCE=2 -fPIE -pie -Wl,-z,relro,-z,now (GCC 14.4, glibc 2.41, AArch64 container, ulimit -c unlimited) and read with gdb 16.3. gdb cannot ptrace under x86-64 emulation, so these are AArch64 sessions: pc is the program counter (x86-64 rip), sp the stack pointer (rsp), x29 the frame pointer (rbp), x30 the link register, where a call leaves the return address (x86-64 pushes it on the stack instead), and x0 to x7 the first eight arguments (x86-64 uses rdi, rsi, rdx, rcx, r8, r9). On x86-64 the same reading applies with those names. The x/ command used below is gdb's memory examine: x/10gx ADDR shows 10 units of g (8-byte, giant) size in x (hex) format starting at ADDR, and x/s prints a C string. Addresses start with 0xaaaa and 0xffff because of PIE and ASLR; subtract the module's load base (info proc mappings) to compare with objdump offsets.
Case A: stack smashing caught by the canary (mode 0)
*** stack smashing detected ***: terminated
Aborted (core dumped)
#4 0x0000ffffa19e9e80 in __fortify_fail () from /lib/aarch64-linux-gnu/libc.so.6
#5 0x0000ffffa19eaf38 in __stack_chk_fail () from /lib/aarch64-linux-gnu/libc.so.6
#6 0x0000aaaab7fe09f4 in stack_loop (n=<optimized out>) at crash_modes.c:13
#7 0x4141414141414141 in ?? ()
Backtrace stopped: previous frame identical to this frame (corrupt stack?)
$1 = 6
(This block is the program's own two terminal lines followed by gdb's output for bt and p $_siginfo.si_signo on the core. Omitted: gdb's start-up lines (Core was generated by ..., Program terminated with signal SIGABRT, the thread-library lines), the command lines, the frame 0 line gdb prints when it opens the core, and frames 0 to 3 (an unnamed libc frame, raise, abort and a second unnamed libc frame). $1 = 6 is the signal number.) The signal is 6, the check is __stack_chk_fail called from the function that owns the overflowed buffer (frame 6, stack_loop), and that function's saved return address is already overwritten with the fill byte (0x4141... in frame 7). The abort message is recoverable from the core even if stderr was lost. glibc saves it in a buffer pointed to by the variable __abort_msg; the first 4 bytes of that buffer hold the buffer's size and the text follows, so *(char**)&__abort_msg reads the pointer and +4 skips the size field. x/s *(char**)&__abort_msg+4 printed "*** stack smashing detected ***: terminated\n". In the disassembly the guard is checked in the epilogue by a load, compare and conditional branch just before the return. The same pattern appears in every protected function; here it is from indexed_write (abridged to the check and the return):
a58: ldr x2, [sp, #40]
a5c: ldr x1, [x0]
a60: subs x2, x2, x1
a64: mov x1, #0x0 // #0
a68: b.ne a78 <indexed_write+0x84> // b.any
a6c: ldp x29, x30, [sp, #48]
a70: add sp, sp, #0x40
a74: ret
(objdump -d --no-show-raw-insn on the PIE binary; the lines before a58 load the address of the guard from the global offset table, the GOT, a table of addresses the dynamic linker fills in, and the failure call a78: bl 7c0 <__stack_chk_fail@plt> follows the ret; the function header line and everything before a58 are omitted). Reading it: ldr x2, [sp, #40] loads the canary copy from the stack, ldr x1, [x0] loads the reference value from the guard, subs subtracts them setting the flags, mov x1, #0x0 clears the register that held the guard value (it does not change the flags), b.ne jumps to the failure call if they differ, and otherwise ldp restores x29 and x30, add sp pops the frame and ret jumps to x30. The canary copy is at [sp, #40], the saved x29 and x30 at [sp, #48] and [sp, #56]: damage that reaches x30 by a linear overflow from a lower address must pass through the canary first. In the mode 0 core, x/gx &__stack_chk_guard printed 0xd142f11ffc02a500, the reference value the stack copy is compared with.
Case B: overflow caught before any damage (mode 1)
*** buffer overflow detected ***: terminated
#5 0x0000ffff8bb995f8 in __chk_fail () from /lib/aarch64-linux-gnu/libc.so.6
#6 0x0000ffff8bb9a7a0 in __memcpy_chk () from /lib/aarch64-linux-gnu/libc.so.6
#7 0x0000aaaabcaa0ab0 in memcpy (__dest=0xffffe252ad78, __src=0xffffe252ac78, __len=64) at /usr/include/aarch64-linux-gnu/bits/string_fortified.h:29
#8 stack_memcpy (src=src@entry=0xffffe252ada8 'A' <repeats 200 times>..., n=n@entry=64) at crash_modes.c:17
#9 0x0000aaaabcaa0c48 in main (argc=<optimized out>, argv=0xffffe252b048) at crash_modes.c:50
(This block is the program's abort message followed by gdb's bt output on the core. Omitted: the shell's Aborted (core dumped) line, gdb's start-up lines, the command line and the frame 0 line gdb prints when it opens the core, and frames 0 to 4 (an unnamed libc frame, raise, abort, another unnamed libc frame and __fortify_fail). The ... in frame 8 is gdb's own truncation of the repeated A.)
This is _FORTIFY_SOURCE rather than the canary: memcpy was rewritten to __memcpy_chk, which knew the destination was 16 bytes and the length 64 and aborted before copying. The stack is intact (frames 8 and 9 are the real callers) and the arguments are visible in the frame. If you see __chk_fail or *_chk frames, no memory was damaged by that call.
Case C: heap corruption (mode 2)
munmap_chunk(): invalid pointer
Aborted (core dumped)
#6 0x0000ffffbd1a8158 in free () from /lib/aarch64-linux-gnu/libc.so.6
#7 0x0000aaaae3600b48 in heap_overflow (n=<optimized out>, n@entry=40) at crash_modes.c:26
(gdb) x/10gx $x19-16
0xaaaadeed1290: 0x0000000000000000 0x0000000000000021
0xaaaadeed12a0: 0x4242424242424242 0x4242424242424242
0xaaaadeed12b0: 0x4242424242424242 0x4242424242424242
0xaaaadeed12c0: 0x4242424242424242 0x0000000000000000
0xaaaadeed12d0: 0x0000000000000000 0x0000000000020d31
(The block above is the program's two terminal lines, then frames 6 and 7 of gdb's bt on the core, then the x/10gx $x19-16 command typed at an interactive gdb prompt after frame 7; omitted are gdb's start-up lines, frames 0 to 5 and frame 8, and the output of the frame 7 command itself, which repeats frame 7 and prints its source line.) The abort comes from inside the allocator (free; frames 0 to 5 are raise, abort and unnamed libc-internal frames, whose symbols are missing because libc is stripped, with frame 8 being main), the user stack is intact, and the message names an allocator check (here the allocator read the forged size field and rejected the pointer). The heap words show why. A chunk header is the allocator's bookkeeping just before the user data: 8 bytes of previous-chunk size, then 8 bytes of this chunk's size whose low bits are flags. Reading the dump (run with frame 7 selected so that x19 is heap_overflow's register; each row is two 8-byte words, the left one at the address shown): the first row 0x0 0x21 is the header of this block, with 0x21 meaning size 0x20 (32 bytes) plus the flag bit 1, "previous chunk in use"; row 2 and the first word of row 3 are the 24 bytes of user data, all 0x42, the byte for B; the second word of row 3 is the next chunk's size field, which now holds 0x4242424242424242 instead of a small size such as 0x21; and rows 4 and 5 are the next block's data area and what follows it. gdb printed x19 as the pointer to the overflowed block because the optimizer kept it in that register. The overflowed block here sits directly before the one being freed, so look for pattern bytes at the start of the neighbouring chunk. The message was again recovered with the __abort_msg command above ("munmap_chunk(): invalid pointer\n"), and exact allocator wording varies by glibc version.
Case D: use-after-free (mode 3)
Bus error (core dumped)
#0 0x0000000aaaaf45fb in ?? ()
#1 0x0000aaaac62f0b80 in use_after_free () at crash_modes.c:35
#2 0x0000aaaac62f0c94 in main (argc=<optimized out>, argv=0xffffe2d92b88) at crash_modes.c:52
$1 = 7
x0 0xaaaaf45fb 45813286395
x2 0xaaaaf45fb 45813286395
x30 0xaaaac62f0b80 187650446134144
(The block above is the program's terminal line, the complete bt (three frames), p $_siginfo.si_signo as $1 = 7, and info registers lines for x0, x2 and x30 only; gdb's start-up lines, the command lines and the frame 0 line gdb prints when it opens the core are omitted.) The stack is intact: frame 1 is the real caller and x30 points just after the call instruction in it. The pc is not a code address but 0xaaaaf45fb, and the same value is in x0 and x2. The source loads the function pointer from the object h and calls it. The freed object's first 8 bytes now hold the allocator's free-list link (glibc 2.41 stores the chunk address shifted right 12 bits, XORed with the next link, here null), so the "function pointer" is the heap address shifted right by 12 bits (its page number), which is what the register shows. A worked example with illustrative numbers: for a freed chunk at 0xaaaaf45fb2a0, shifting right by 12 bits drops the last three hex digits and gives 0xaaaaf45fb; XORed with a null next link (0) the stored link is 0xaaaaf45fb, the value in x0. This scheme is glibc's safe-linking, a hardening that stops an attacker forging free-list links without knowing the heap address. The signature is: an indirect call or load through a pointer whose value looks like a small number or a random 64-bit value that corresponds to the freed chunk, a valid caller, and a valid link register. The signal was SIGBUS because the value is not a multiple of 4 (it was odd in the transcript above, but any value with a non-zero low two bits qualifies, and an even value that is not a multiple of 4 is still SIGBUS; ASLR changes the heap address on every run, so these low bits vary, most runs give SIGBUS, and an occasional run whose shifted address happens to be a multiple of 4 gives SIGSEGV instead), and AArch64 instruction addresses must be 4-byte aligned, so jumping to such an address faults on alignment (Linux reports it as SIGBUS, as the transcript shows); a value that is aligned but unmapped gives SIGSEGV.
Case E: corrupted return address with the canary intact (mode 4, index 4)
An indexed write that skips over the canary (index 2 is the canary, 3 the saved x29, 4 the saved x30 in this frame; index 2 aborts with __stack_chk_fail, as shown by running the same binary with that argument):
Segmentation fault (core dumped)
#0 0x0042424242424240 in ?? ()
#1 0x4242424242424240 in ?? ()
Backtrace stopped: previous frame identical to this frame (corrupt stack?)
$1 = 11
pc 0x42424242424240 0x42424242424240
x29 0xffffd633e870 281474275469424
x30 0x4242424242424240 4774451407313060416
(The block above is the program's terminal line, the complete bt (two frames and gdb's "Backtrace stopped" line), p $_siginfo.si_signo as $1 = 11, and info registers lines for pc, x29 and x30 only; gdb's start-up lines, the command lines and the frame 0 line gdb prints when it opens the core are omitted.) No abort, no check function in the frames, SIGSEGV with pc and x30 both holding the 0x42 pattern the write stored (gdb shows pc as 0x0042424242424240 and x30 as 0x4242424242424240; the value ends in 0, so it is 4-byte aligned and the fault is an ordinary unmapped-address fault instead of the bus error seen in case D), and no valid caller above. The frame pointer x29 is still a sane stack address, which shows only the return slot was hit. That exactly matches what the epilogue does: restore x29/x30 from the stack, then ret. The canary compare passed because the canary word was never touched. The same core shows it: after the epilogue sp points just above the dead frame, so the canary copy is at $sp-0x18, saved x29 at $sp-0x10 and saved x30 at $sp-8:
(gdb) x/gx &__stack_chk_guard
0xffff98373b60 <__stack_chk_guard>: 0x32b29a1b57379e00
(gdb) x/8gx $sp-0x30
0xffffd3997410: 0x0000ffffd3997708 0x0000000000000000
0xffffd3997420: 0x0000000000000000 0x32b29a1b57379e00
0xffffd3997430: 0x0000ffffd3997550 0x4242424242424240
0xffffd3997440: 0x0000000000000000 0x4141414141414141
(Typed at the (gdb) prompt on the core of that second crash; no lines are omitted from these two commands' output.) The word at 0xffffd3997428 equals the guard, the saved x29 slot at 0xffffd3997430 is a sane stack address, and only the return slot at 0xffffd3997438 holds the written value. (This transcript is from a second crash of the same program with the same arguments; addresses and the guard value change on every run.)
How to tell them apart (decision table)
| Observation | Stack smashing (A) | Heap corruption (C) | Use-after-free (D) | Corrupted return (E) |
|---|---|---|---|---|
| Signal | 6 | 6 usually | 11 or 7 | 11 or 7 |
| Frames at the top | __stack_chk_fail over the buggy function | free / malloc internals over the caller | the program's own function, pc in no function | no check frame; pc not in any function |
| Caller frame valid? | The buggy function's own saved return is usually trashed | Yes | Yes, x30 valid | No, x30 and pc both hold the written pattern (printed as 0x4242424242424240 and 0x0042424242424240), backtrace breaks |
| What the odd value is made of | Fill pattern in the saved return slot | Fill pattern in the neighbouring chunk's header | Allocator link (address >> 12) or freed-chunk contents | Fill pattern or attacker value |
| Canary | Different from the guard value | Not involved | Not involved | Equal to the guard value |
For the canary row, compare the canary slot in the frame with the global guard using x/gx on both (x/gx &__stack_chk_guard worked here, where print could not use the symbol because libc has no type information for it).
Next steps
Rebuild the failing path with AddressSanitizer (-fsanitize=address), which names the defect (stack-buffer-overflow for the two stack cases, heap-buffer-overflow, heap-use-after-free) and prints the offending write. On the test program above it reports modes 0 to 3, and mode 4 with index 2 or 3, but it is silent for mode 4 with index 4 (AArch64): that indexed write jumps past the redzone ASan places around slot, so nothing is reported. In the ASan build itself that run prints 0 and exits normally, because the ASan-instrumented frame keeps the saved x29 and x30 below slot, where an upward write cannot reach them; only the hardened build's frame layout puts the return slot in range of index 4. So a write that skips the redzone is the case ASan can miss, and for the corrupted return with an intact canary the core from the hardened build is what identifies it. -fsanitize=bounds catches that particular write (index 4 out of bounds for type 'long int [2]'), because it checks the index against the array's declared size rather than watching memory. If you cannot rebuild, the core plus the matching binary and libraries (same build ID) is what you need; keep stderr or journal logs, since the abort message names the check.
What is the practical difference between a software breakpoint and a hardware watchpoint in GDB or LLDB, and when would you prefer one over the other while investigating a security issue?
Sample Answer
Direct answer
A software breakpoint stops the program when execution reaches an instruction: the debugger overwrites that instruction with a trap instruction (a special instruction that raises an exception the debugger catches), lets the program run, and restores the original on the stop. It asks "did execution arrive here?". A hardware watchpoint stops the program when a memory location is read or written: the CPU's own debug registers (a few special registers inside the processor that each hold an address to watch, with comparison hardware wired to them) check every memory access against that address, with no change to code and almost no slowdown. It asks "who touched this data?". Use a breakpoint to answer where a function is entered or a branch taken. Use a hardware watchpoint to catch the unknown writer of a corrupted variable, canary, saved return address or function pointer, which is the typical security investigation.
Mechanism and what you can observe
On x86-64 the trap instruction is the one-byte int3, whose opcode (the number that encodes the instruction in machine code) is 0xCC; on AArch64 it is the 4-byte BRK. The session below is AArch64, so on an x86-64 machine the same experiment reads differently: GDB overwrites only the first byte of target with 0xCC, x/i $pc still shows the original instruction (for the -O0 build of this program, push rbp, byte 0x55), and the first 32-bit word the program reads back changes in its lowest byte only, from 0xe5894855 to 0xe58948cc (x86 is little-endian, so the first byte is the low byte of the word). The unpatched word 0xe5894855 is what this program prints on x86-64 without a debugger; the patched word follows from the one-byte patch and was not captured on x86-64. Run natively on AArch64 Linux (GDB 16.3), this program reads the first instruction word of its own function target before and after a breakpoint hit, with and without the debugger:
#include <stdio.h>
int counter;
__attribute__((noinline)) int target(int x)
{
counter += x;
return counter;
}
int main(void)
{
unsigned int *code = (unsigned int *)(void *)target;
printf("first instruction word of target: 0x%08x\n", code[0]);
target(1);
printf("first instruction word of target: 0x%08x\n", code[0]);
return 0;
}
Built with gcc -O0 -g -fno-stack-protector (GDB's startup warnings about address-space randomization and the thread library, and the Starting program: and Continuing. lines, are removed from the transcripts here, and the process number in the last line differs from run to run). In the transcript, break *target puts the breakpoint on the function's very first instruction, x/i $pc means "examine memory at the program counter as one instruction", and the inferior in the last line is GDB's word for the program being debugged:
$ ./bp
first instruction word of target: 0xd10043ff
first instruction word of target: 0xd10043ff
$ gdb -q ./bp
Reading symbols from ./bp...
(gdb) break *target
Breakpoint 1 at 0x400644: file bp.c, line 6.
(gdb) run
first instruction word of target: 0xd4200000
Breakpoint 1, target (x=65535) at bp.c:6
6 {
(gdb) x/i $pc
=> 0x400644 <target>: sub sp, sp, #0x10
(gdb) continue
first instruction word of target: 0xd4200000
[Inferior 1 (process 184) exited normally]
The 0xd10043ff word is sub sp, sp, #16. The first printed line appears before the stop because main reads the word before it calls target, and the breakpoint is already inserted by then (on a terminal; when the program's output goes to a pipe it is buffered, and both lines appear when the program exits). While the breakpoint is inserted the running program sees 0xd4200000 (brk #0) instead, even though GDB's x/i shows the original instruction because GDB hides its own patch. That is the practical cost of a software breakpoint: it writes to code memory. It therefore cannot be placed in read-only flash (use hbreak, a hardware breakpoint, which uses a comparator register: a debug register that matches the address of the instruction about to run, so nothing in memory has to change), it can break self-checking or self-modifying code, and a program (or malware) that checksums its own code or scans for trap opcodes can detect it. A hardware watchpoint does none of that.
A hardware watchpoint finding a buffer overflow
A buffer overflow is a write past the end of an array that lands on whatever sits next to it. This program has an unchecked strcpy that overruns an 8-byte buffer into a neighbouring flag:
#include <stdio.h>
#include <string.h>
__attribute__((noinline)) void copy_name(const char *src)
{
int is_admin = 0;
char name[8];
strcpy(name, src); /* no bounds check */
printf("is_admin = %d\n", is_admin);
}
int main(int argc, char **argv)
{
copy_name(argc > 1 ? argv[1] : "bob");
return 0;
}
Built with gcc -O0 -g -fno-stack-protector over.c -o over, ./over prints is_admin = 0 and ./over AAAAAAAAAAAAAAAAAAAA (twenty A characters) prints is_admin = 1094795585, which is 0x41414141. The overwrite happens inside a libc function, so a breakpoint in your own code can only show the damage afterwards. Watching the variable shows the culprit (abridged GDB 16.3 session on AArch64: the Breakpoint 1 at line, the output of run, the Continuing. lines, the Hardware watchpoint 2: is_admin line that precedes each Old value together with the blank lines around it, and the source line under the first stop are left out, the arguments of copy_name are replaced by (...), and the library address differs from run to run; the first stop is the legitimate initialization). watch is_admin asks GDB to stop whenever that variable's value changes. The first Old value = 65535 is not a value the program ever wrote: is_admin is a stack slot, and before the line int is_admin = 0; runs it still holds leftover bytes from earlier calls, so the first change reported is the program's own initialization to 0. (The same leftover value appears as x=65535 in the breakpoint transcript above, for the same reason: the argument slot had not been written yet at the function's first instruction.)
(gdb) break copy_name
(gdb) run AAAAAAAAAAAAAAAAAAAA
(gdb) watch is_admin
Hardware watchpoint 2: is_admin
(gdb) continue
Old value = 65535
New value = 0
copy_name (...) at over.c:8
(gdb) continue
Old value = 0
New value = 1094795585
0x0000ffff9ed0ea9c in strcpy () from /lib/aarch64-linux-gnu/libc.so.6
(gdb) x/i $pc
=> 0xffff9ed0ea9c <strcpy+92>: str q1, [x0, x5]
The stop is inside strcpy at the vector store that wrote the flag (q1 is a 128-bit SIMD register, and libc's strcpy copies 16 bytes per store, so one instruction overwrites the neighbouring flag together with the buffer), and bt from there leads back to copy_name, which is the fix location: bound the copy. The watch command printed Hardware watchpoint, which means the CPU is doing the comparison. For a region too large for the debug registers the same command prints plain Watchpoint, i.e. a software watchpoint: for example watch -l *(char(*)[64])name produced Watchpoint 2: -location *(char(*)[64])name. Reading that syntax from the inside out: (char(*)[64]) is a C cast to "pointer to an array of 64 chars", applied to name, and the leading * dereferences it, giving GDB a 64-byte object to watch. -l (short for -location) tells GDB to watch the memory address that expression refers to right now, rather than re-evaluating the expression by name, which would also stop the watch when the frame goes away. A 64-byte region is far wider than the debug registers can match, so GDB falls back to the software watchpoint.
Choosing between them
| Question | Use | Why |
|---|---|---|
| Is this function reached, and with what arguments? | software breakpoint (break) | cheap, unlimited in number |
| Code is in ROM or flash, or is checksummed | hardware breakpoint (hbreak) | no code write |
| Who modified this variable, pointer, canary or return slot? | hardware watchpoint (watch, or watch -l on an address) | stops at the writing instruction, no slowdown |
| Who reads this secret or flag? | read or access watchpoint (rwatch, awatch) | catches reads as well as writes |
| Region larger than the CPU can match | software watchpoint | works but is single-stepped (the debugger executes one instruction at a time and re-reads the value after each) |
GDB's manual states the limits of the alternative: a software watchpoint works by single-stepping the program and testing the value each time, which is "hundreds of times slower than normal execution" and reports the change at the next statement rather than the instruction. Hardware watchpoints are limited by the number of debug registers and by the width they can watch (for example some systems only watch up to 4 or 8 bytes), and insertion can fail when you ask for too many (the manual's wording is "Could not insert watchpoint"; GDB 16.3 prints Could not insert hardware watchpoint N.), in which case you delete or disable some watchpoints. Watchpoints on locals are deleted when the function returns.
rwatch stops on reads only and awatch on reads or writes; watch alone is the one used in the examples above. LLDB equivalents: breakpoint set --name target (add --hardware for a hardware breakpoint) and watchpoint set variable is_admin or watchpoint set expression -s 8 -- <address>; LLDB's help also warns that hardware resources for watchpoints are limited.
Pitfalls
- Setting
watch is_adminbefore the variable is initialized stops on the legitimate store first, as above. Skip it withcontinue, or add a condition (watch is_admin if is_admin != 0). - A hardware watchpoint on a stack slot needs to be set after the frame exists, and its address is only valid for this activation.
- On threaded programs a watchpoint may need to apply to all threads, and the stop is on the thread that wrote.
- A stack canary is a known value the compiler places between a buffer and the saved return address and checks before the function returns; watching the canary's address catches the overrun that would trigger the check. (The programs above were built with
-fno-stack-protector, which turns canaries off so the overflow is visible.) - For corruption you cannot reproduce under a debugger, sanitizers (AddressSanitizer, a compiler-inserted checker for out-of-bounds and use-after-free memory accesses) find the same overrun at the faulty write without a debugger, which is a complement, not a substitute, when you need to understand the surrounding state.
A stack buffer overflow is suspected in a service, but the crash is intermittent. How would you use the debugger to catch the exact instruction that corrupts the stack, and how would you tell whether the overwrite happened before or after the function's return address was last used?
Sample Answer
Direct answer
Find the memory slot that holds the return address (the saved copy of the address the CPU jumps back to when the function finishes), then ask the CPU to stop the instant anything writes to it. That is a hardware watchpoint: gdb programs the processor's debug registers so the write traps right after the offending instruction completes. The stop shows you the exact instruction (and its caller chain). To decide "before or after the return address was last used", note that the address is used once, by the ret at the end of the function that owns the slot: if the watchpoint fires while that function (or its callees) is still on the backtrace, the overwrite came first and ret then jumps to the corrupted value; if the write only shows up after that ret ran, the slot belongs to a dead frame and the stack was reused, which is a different bug (a stale pointer into a returned frame), not a smashed return address.
Where the return address lives (it differs by architecture)
A stack frame is the stack space one function call owns: its saved registers, its locals and its saved return address. On x86-64, call pushes the return address, so it sits above the callee's locals. For the program below, GCC 14.4 at -O0 puts name at [rbp-16], the saved rbp at [rbp] and the return address at [rbp+8], so the bytes from name upward are: offsets 0 to 15 are name itself, 16 to 23 are the saved rbp, and 24 to 31 are the return address. A copy that writes 25 or more bytes (counting the terminating NUL that strcpy adds) reaches the return address; exactly 24 bytes stops after the saved rbp. (That -O0 layout is what the x86-64 compiler's assembly output (-S) shows for this program.) On AArch64 the return address arrives in register x30 (the link register, which bl fills in instead of pushing onto the stack) and a non-leaf function saves it with stp x29, x30, [sp, #-48]! at the bottom of its frame, with locals above it. There, the same overflow runs upward out of parse into the caller's frame and corrupts main's saved x30, not parse's. Read the layout from the debugger, not from a diagram.
The program
#include <stdio.h>
#include <string.h>
static void parse(const char *input) {
char name[16];
strcpy(name, input); /* no length check */
printf("hello %s\n", name);
}
int main(void) {
parse("short");
parse("this input is much longer than sixteen bytes");
return 0;
}
Built with gcc -O0 -g -fno-stack-protector -no-pie -o overflow overflow.c (GCC 14.4, aarch64 Linux). The stack protector is off so that nothing stops the overflow before gdb sees it.
The ordered procedure
- Reproduce under the debugger and stop in a state where the slot exists. Break at the entry of the suspect function. A breakpoint condition such as
$_streq(input, "...")(gdb's built-in string-equals function) stops only for the bad input; when no input distinguishes the bad call,ignore N countskips the first N hits instead. Attach withgdb -p <pid>for a running service. - Locate the return-address slot.
info framefor the frame that owns it prints "saved registers" with the stack address of the saved return register (x30 at 0x...on AArch64,rip at 0x...on x86-64). - Set a hardware watchpoint on that address, using
watch -l(location: evaluate the expression once, take its address, and watch that fixed memory) so it stays attached to the memory even when the frame changes, and run. The expression*(unsigned long *)0xfffffffff978means "treat this number as the address of an 8-byte unsigned integer". - At the stop, read the evidence. Hardware watchpoints report after the writing instruction has completed, so the culprit is the instruction before
$pc, the program counter ($pcis the address of the next instruction to run).x/2i $pc-4means "examine 2 instructions starting at$pcminus 4": on fixed-width AArch64 (every instruction is 4 bytes) the first line is the instruction that just executed and wrote the slot, and the second line, marked=>, is the one about to run. On x86-64 instructions vary in length, so a start address a few bytes back can land inside an instruction and decode as nonsense (x/3i $pc-8did exactly that in the x86-64 session described below); usedisassembleon the function that gdb names, or on a range that starts at a known instruction boundary, and read the instruction just above the=>marker, or usebtand the function it names.btshows who called it, the old and new values show what was written (ASCII-looking bytes mean text from a string copy), andx/son the buffer shows the source. - Place a breakpoint on the
retof the function whose slot you watch. If the watchpoint fires first, the overwrite was before the return address was used.info breakpointsalso records hit counts. - Confirm and fix. Continue to see the jump to the bad address, then replace the unbounded copy with a bounded one, and rerun the same session to see the watchpoint stay silent.
A real transcript (AArch64, gdb 16.3)
The transcript omits the two libthread_db lines (gdb loading its thread-debugging helper library), the Breakpoint 1 at 0x400690: file overflow.c, line 6. and Breakpoint 2 at 0x4006e4: file overflow.c, line 14. lines that the two break commands print, the blank line that run prints before the first stop line, any Starting program and Continuing. progress lines, the program's own hello short line (printed by the first call to parse, between run and the first stop, because stdout is line-buffered on a terminal), and the source-line echo that gdb prints after each stop and after up (for example 6 strcpy(name, input);). In the info frame output, the ... line stands for the lines above Saved registers: (the frame address, pc and saved pc, the caller frame, the source language, and the argument and locals addresses with the previous sp). The stack addresses are from one run and differ on other runs. main's ret is at 0x4006e4 (from disassemble main).
(gdb) break parse if $_streq(input, "this input is much longer than sixteen bytes")
(gdb) break *0x4006e4
(gdb) run
Breakpoint 1, parse (input=0x400720 "this input is much longer than sixteen bytes") at overflow.c:6
(gdb) up
#1 0x00000000004006dc in main () at overflow.c:12
(gdb) info frame
...
Saved registers:
x29 at 0xfffffffff970, x30 at 0xfffffffff978
(gdb) watch -l *(unsigned long *)0xfffffffff978
Hardware watchpoint 3: -location *(unsigned long *)0xfffffffff978
(gdb) continue
Hardware watchpoint 3: -location *(unsigned long *)0xfffffffff978
Old value = 281474840470108
New value = 8295751878259777650
0x0000fffff7e8eb14 in strcpy () from /lib/aarch64-linux-gnu/libc.so.6
(gdb) bt
#0 0x0000fffff7e8eb14 in strcpy () from /lib/aarch64-linux-gnu/libc.so.6
#1 0x000000000040069c in parse (input=0x400720 "this input is much longer than sixteen bytes") at overflow.c:6
#2 0x00000000004006dc in main () at overflow.c:12
(gdb) x/2i $pc-4
0xfffff7e8eb10 <strcpy+208>: str q0, [x3], #32
=> 0xfffff7e8eb14 <strcpy+212>: ldr q0, [x2, #16]
(gdb) continue
Breakpoint 2, 0x00000000004006e4 in main () at overflow.c:14
(gdb) p/x $x30
$1 = 0x73206e6168742072
(gdb) stepi
0x00206e6168742072 in ?? ()
Reading it: the slot at 0xfffffffff978 is main's saved x30. The old value, 281474840470108, is 0xfffff7e1225c in hex: a valid libc address (the point in main's caller where main should return), which is what a healthy return address looks like. The write came from a 16-byte vector store (str q0, [x3], #32) inside strcpy, called from parse at line 6. The new value, 8295751878259777650, is 0x73206e6168742072: the bytes 72 20 74 68 61 6e 20 73, which are the ASCII characters "r than s" from the input string. main was still on the backtrace and its own ret had not yet executed: the breakpoint on ret fires after the watchpoint, and by then the ldp x29, x30, [sp], #16 just before that ret has loaded the corrupted word into x30, which is why p/x $x30 prints 0x73206e6168742072. So the overwrite came before the return address was used, and ret jumps to garbage (the next stepi lands at 0x00206e6168742072: in this session the program counter holds the target without the 0x73 top byte that x30 still shows, and the fault address the kernel reports in siginfo has the same form, so this is not just gdb's formatting; running freely instead of stepping ends in SIGBUS, a bus-error signal, and the garbage target here is not a multiple of 4 while AArch64 instruction addresses must be). On x86-64 the overflow reaches the return address of parse itself, so no up is needed: info frame at the breakpoint in parse reads rip at 0x..., watch -l on that slot stops inside strcpy with an old value that is a valid code address in main and a new value made of the same ASCII bytes, and the ret to break on is parse's own. That x86-64 session was run in qemu-x86_64 user mode with gdb-multiarch; QEMU's debug stub could not insert a hardware watchpoint: gdb printed Hardware watchpoint 2 when the watch was set, then continue failed with Could not insert hardware watchpoint 2. After set can-use-hw-watchpoints 0, gdb set a software watchpoint (printed as Watchpoint N, without the word Hardware) and the stop inside strcpy described above followed. The transcript above is the native AArch64 session.
What goes wrong, and limits of the method
- The processor has only a few debug registers, each covering a small aligned region, so watch only the slot or a few bytes, not the whole buffer. If hardware watchpoints are unavailable, gdb falls back to single-stepping (
set can-use-hw-watchpoints 0forces this) and the program runs far slower. - A watchpoint written as
watch name(a local) is dropped when its frame exits; use-lwith an address. - Without a debugger available in production: compile with
-fstack-protector-strong, whose canary (a guard value placed between locals and the return address and checked at function exit) aborts with "stack smashing detected", or build with AddressSanitizer, which reports the overflow at the writing instruction. Both tell you the function and line, but only the watchpoint gives the exact instruction and the order relative toret. - Intermittent crashes: enable core dumps (
ulimit -c unlimited) so a crash leaves a file, then use the watchpoint on a deterministic replay if you can reconstruct the triggering input.
How does a call to a shared-library function such as printf actually get resolved at runtime on x86-64 Linux? Walk through what the dynamic linker does, what the call site looks like, and how lazy binding differs from immediate binding.
Sample Answer
Direct answer
The executable cannot contain the address of printf or puts, because the C library is loaded at an address chosen at run time. So the call site calls a small stub in the executable, the PLT (procedure linkage table). The stub jumps through a pointer stored in the GOT (global offset table, a writable array of addresses the dynamic linker fills in). With lazy binding the pointer initially points back into the stub, which hands control to the dynamic linker (ld.so, the program that loads shared libraries and resolves symbols); the linker looks the symbol up, writes the real address into the GOT slot and jumps there. Every later call goes through the same stub but lands directly in libc. With immediate binding (LD_BIND_NOW or the -z now link flag) the linker fills in every slot before main starts, so no call ever takes the slow path.
The artifacts on x86-64 Linux
The program (greet.c):
#include <stdio.h>
int main(void)
{
puts("first call");
puts("second call");
return 0;
}
Built with gcc -O1 -o greet greet.c (GCC 14.4 in a linux/amd64 container, emulated; this toolchain produced a non-PIE executable with lazy binding, and readelf -d shows no BIND_NOW flag). The main function from objdump -d --no-show-raw-insn greet (the flag drops the column of machine-code bytes):
0000000000401126 <main>:
401126: sub $0x8,%rsp
40112a: mov $0x402004,%edi
40112f: call 401030 <puts@plt>
401134: mov $0x40200f,%edi
401139: call 401030 <puts@plt>
40113e: mov $0x0,%eax
401143: add $0x8,%rsp
401147: ret
A relocation is a record that tells the loader "patch this address with the real address of that symbol". The linker emitted one of type R_X86_64_JUMP_SLOT, which tells ld.so which GOT slot belongs to puts. It is in the .rela.plt section of the readelf -rW greet output (the .rela.dyn section that readelf prints before it is omitted here). In the Info column, the high 32 bits (00000002) are the symbol-table index of puts and the low 32 bits (00000007) are the relocation type, 7 = R_X86_64_JUMP_SLOT:
Relocation section '.rela.plt' at offset 0x468 contains 1 entry:
Offset Info Type Symbol's Value Symbol's Name + Addend
0000000000404000 0000000200000007 R_X86_64_JUMP_SLOT 0000000000000000 puts@GLIBC_2.2.5 + 0
The PLT (objdump -d --no-show-raw-insn -j .plt greet, with its file-header lines (greet: file format elf64-x86-64, Disassembly of section .plt:) and the blank lines around them omitted). The first entry, puts@plt-0x10, is PLT0, the shared resolver stub, and puts@plt is the stub for puts. This listing is in AT&T syntax, so the source operand comes first and the destination last, $ marks an immediate value, * marks an indirect jump through memory, and 0x2fca(%rip) is a memory operand at the address of the next instruction plus 0x2fca; the # 403ff0 <...> text is objdump's annotation with the address that operand works out to. nopl is a multi-byte no-operation used as padding, and psABI below means the processor-specific ABI supplement, the document that fixes these rules for x86-64:
0000000000401020 <puts@plt-0x10>:
401020: push 0x2fca(%rip) # 403ff0 <_GLOBAL_OFFSET_TABLE_+0x8>
401026: jmp *0x2fcc(%rip) # 403ff8 <_GLOBAL_OFFSET_TABLE_+0x10>
40102c: nopl 0x0(%rax)
0000000000401030 <puts@plt>:
401030: jmp *0x2fca(%rip) # 404000 <puts@GLIBC_2.2.5>
401036: push $0x0
40103b: jmp 401020 <_init+0x20>
The push $0x0 at 0x401036 pushes relocation index 0. The GOT as stored in the file (objdump -s -j .got.plt greet, with its file-header lines and blank lines omitted) holds 8-byte little-endian entries starting at 0x403fe8, and the right-hand column is objdump's ASCII rendering of the bytes:
Contents of section .got.plt:
403fe8 083e4000 00000000 00000000 00000000 .>@.............
403ff8 00000000 00000000 36104000 00000000 ........6.@.....
Each 8-byte entry is stored with its least significant byte first, so the bytes 08 3e 40 00 00 00 00 00 are reversed to 0x0000000000403e08, and 36 10 40 00 00 00 00 00 to 0x0000000000401036. That reads: slot 0 = 0x403e08 (the address of the _DYNAMIC section), slots 1 and 2 = 0, slot 3 (0x404000, the slot for puts) = 0x401036, the address of the push $0 instruction right after the jmp in the puts stub.
The first call, step by step
The System V x86-64 psABI (the ABI document) describes this sequence:
maincallsputs@plt(0x401030).- The first instruction jumps to the address in the GOT entry for
puts. Initially that entry "holds the address of the followingpushqinstruction, not the real address" of the function, so the jump falls straight through to the next instruction. push $0pushes the relocation index (the position ofputs's relocation in.rela.plt), then the stub jumps to PLT0.- PLT0 pushes the second GOT entry (GOT+8, which the psABI describes as one word of identifying information for the dynamic linker, set when the program image is created) and jumps through the third GOT entry (GOT+16), which the dynamic linker set to its own resolver. The psABI says that on x86-64 GOT entries one and two are reserved for this purpose.
- The resolver uses the relocation index to find the
R_X86_64_JUMP_SLOTentry (the relocation record shown above), looks the symbol up in the loaded libraries, stores the real address in the GOT slot (0x404000), and jumps toputs. - On the second call, step 2's jump goes directly to
putsin libc; the dynamic linker is never entered again for this symbol.
The last call returns to main normally: the resolver's work is invisible except for the first call being slower.
Observing the GOT slot change
This program reads slot 3 of the GOT before and after the first call (non-PIE, so _GLOBAL_OFFSET_TABLE_ is the start of .got.plt, and puts is the first PLT function referenced):
#define _GNU_SOURCE
#include <dlfcn.h>
#include <stdio.h>
/* Linker-defined symbol: the start of .got.plt. Slots 0-2 are reserved for the
dynamic loader; the first PLT function (here puts) owns slot 3. */
extern void *_GLOBAL_OFFSET_TABLE_[];
int main(void)
{
void *before = _GLOBAL_OFFSET_TABLE_[3];
puts("hello"); /* first call goes through the PLT */
void *after = _GLOBAL_OFFSET_TABLE_[3];
void *real = dlsym(RTLD_DEFAULT, "puts");
printf("slot changed by first call: %s\n", before != after ? "yes" : "no");
printf("slot now equals libc puts: %s\n", after == real ? "yes" : "no");
return 0;
}
Built as gcc -O1 -no-pie -Wl,-z,lazy -o got_peek got_peek.c and as -Wl,-z,now, on linux/amd64 (emulated). Output (the two -- heading lines name the build that was run and were printed by a shell echo; the rest is the program's own output):
-- lazy:
hello
slot changed by first call: yes
slot now equals libc puts: yes
-- -z now:
hello
slot changed by first call: no
slot now equals libc puts: yes
With lazy binding the slot changes on the first call; with -z now it already held the final address before the call.
LD_DEBUG=bindings shows the same moment from the loader's side. With lazy binding, the symbol is bound after transferring control (when main is already running); with LD_BIND_NOW=1 it is bound before. Run as LD_DEBUG=bindings ./greet 2>&1 | grep -E 'puts|first|second|calling init|transferring' (and the same with LD_BIND_NOW=1 in front), with the process-id prefix and leading whitespace of each line removed and the lazy: and LD_BIND_NOW=1: headings added by hand; every other line the filter drops is omitted. Read the lines as follows: calling init lines show the loader running each library's start-up code; transferring control: ./greet is the point where the loader hands over to the program's own entry point, so everything after it is the program running; binding file ./greet ... to .../libc.so.6 ...: normal symbol puts'is the loader resolving the reference fromgreettoputsin libc, and the version tag[GLIBC_2.2.5]is the symbol version it matched (the bracketed[0]numbers can be ignored here).first callandsecond callare the program's own output. Because the pipe makesputsbuffer its output until the program exits, these two lines print last in both modes, so they do not mark when each call happened; thetransferring controlline is the marker for when the program's own code starts running, and the binding line shows whetherputs` was resolved before or after it:
lazy:
calling init: /lib64/ld-linux-x86-64.so.2
calling init: /lib/x86_64-linux-gnu/libc.so.6
transferring control: ./greet
binding file ./greet [0] to /lib/x86_64-linux-gnu/libc.so.6 [0]: normal symbol `puts' [GLIBC_2.2.5]
first call
second call
LD_BIND_NOW=1:
binding file ./greet [0] to /lib/x86_64-linux-gnu/libc.so.6 [0]: normal symbol `puts' [GLIBC_2.2.5]
calling init: /lib64/ld-linux-x86-64.so.2
calling init: /lib/x86_64-linux-gnu/libc.so.6
transferring control: ./greet
first call
second call
Lazy versus immediate binding
| Lazy | Immediate (-z now, LD_BIND_NOW) | |
|---|---|---|
| When symbols resolve | on first call of each function | all, before main |
| Startup cost | low | higher: every imported function is looked up |
| First-call cost | one trip through the resolver | none |
| Failure mode | an unresolvable function fails when first called, possibly late | fails at startup |
| GOT writability | .got.plt slots must stay writable at run time | GOT can be made read-only after relocation |
RELRO (relocation read-only: the loader makes parts of the GOT read-only once it is done) shows the last row. Linking greet.c three ways (gcc -O1 -o greet greet.c with -Wl,-z,norelro, -Wl,-z,relro,-z,lazy and -Wl,-z,relro,-z,now) and reading readelf -lW and -SW gives the table below, which is a hand-built summary of those two outputs (the GNU_RELRO segment is the address range the loader write-protects after relocation; the address and size columns of the program header were combined into a start-to-end range):
| Link flags | GNU_RELRO range | Where the puts slot is | Is it write-protected? |
|---|---|---|---|
-z norelro | no such segment | .got.plt | no |
-z relro -z lazy | 0x403df8 to 0x404000 | 0x404000, the first byte after the range | no, so lazy binding can still write it |
-z relro -z now | 0x403dd0 to 0x404000 | inside .got (0x403fd0 to 0x404000), which is covered | yes |
With lazy binding the protected range still includes .got (0x403fd8) and the reserved first three slots of .got.plt (0x403fe8 to 0x404000), but the lazily filled slot at 0x404000 sits just past the end and stays writable. With -z now there is no lazy slot, so the function slots move into .got and the whole table is protected. For a penetration tester this is the practical difference: overwriting a function's GOT slot to redirect a later call requires the slot to be writable at run time, which lazy binding requires and "full RELRO" (relro plus now) removes.
Pitfalls
LD_DEBUGandLD_BIND_NOWare environment variables read byld.so; they are useful for diagnosis, not for deployment.
How does a variadic C function such as printf find its extra arguments on x86-64 System V when some arrive in registers and some on the stack, and why does the caller set up an extra register before the call?
Sample Answer
Direct answer
A variadic function is one declared with ..., like printf: it has some named parameters and then accepts however many extra arguments the caller supplies (printf("%d %s\n", 7, "x") passes two extras after the format string). An ABI (application binary interface) is the set of rules for how code passes arguments and results in registers and on the stack.
On x86-64 System V (the ABI used on Linux, BSD and macOS, whose rules come from the AMD64 psABI, the processor-specific supplement to the System V ABI), the compiler gives a variadic function no special way to receive its extra arguments. They arrive exactly like any other call: the first six integer-class arguments (integers, characters and pointers, as opposed to floating-point values) in rdi, rsi, rdx, rcx, r8, r9, the first eight floating-point ones in xmm0 to xmm7, the rest on the stack. The variadic function's prologue (its entry code) spills all of those argument registers into a block in its own frame called the register save area, and builds a small va_list structure that records where the next register argument and the next stack argument are. va_arg then walks the register save area first and switches to the stack area when the registers are used up. The caller sets al (the low byte of rax) because saving eight 16-byte vector registers is the expensive part: al is a hidden argument giving an upper bound on how many vector registers carry arguments, so the callee can skip the vector saves entirely when it is zero.
What the psABI specifies
From the System V AMD64 psABI text:
- For calls that may reach a variadic function,
al"is used as hidden argument to specify the number of vector registers used. The contents ofaldo not need to match exactly the number of registers, but must be an upper bound on the number of vector registers used and is in the range 0–8 inclusive." - The
va_listtype is an array of one structure with four fields:unsigned int gp_offset,unsigned int fp_offset,void *overflow_arg_areaandvoid *reg_save_area. - The register save area holds
rdiat offset 0,rsiat 8,rdxat 16,rcxat 24,r8at 32,r9at 40, thenxmm0at 48,xmm1at 64, and so on up toxmm7at 160, 176 bytes in all. gp_offsetis the offset to the next unused integer register slot (at most 48),fp_offsetthe offset to the next unused vector slot (at most 176), andoverflow_arg_areapoints at the next stack-passed argument.
The calling side, from the compiler
GCC 14.4 -O2 -S -masm=intel (an x86-64 Linux container run under emulation; Intel syntax, so the destination operand comes first and [rsp+24] is memory at rsp plus 24) for these printf calls with no, one and two floating-point arguments. In the listings that follow, assembler directives, the local .LFB/.LFE labels and the string and constant data (.LC0 to .LC4) are removed, and the text after each ; is a hand-added annotation, not compiler output; the instructions are otherwise as emitted, in emitted order:
#include <stdio.h>
void say_int(void) { printf("%d\n", 7); }
void say_dbl(void) { printf("%d %f\n", 7, 2.5); }
void say_two(void) { printf("%f %f\n", 1.5, 2.5); }
say_int: ; printf("%d\n", 7)
mov esi, 7
mov edi, OFFSET FLAT:.LC0
xor eax, eax
jmp printf
say_dbl: ; printf("%d %f\n", 7, 2.5)
mov esi, 7
mov edi, OFFSET FLAT:.LC2
mov eax, 1
movsd xmm0, QWORD PTR .LC1[rip]
jmp printf
say_two: ; printf("%f %f\n", 1.5, 2.5)
movsd xmm1, QWORD PTR .LC1[rip]
mov edi, OFFSET FLAT:.LC4
mov eax, 2
movsd xmm0, QWORD PTR .LC3[rip]
jmp printf
al is 0, 1 and 2: the number of vector registers used, which is a valid upper bound. A caller that always loads 8 would also be correct; it would only make the callee do unneeded work. The calls are jmp printf because nothing follows them: a tail call is a call made as the very last action, compiled as a plain jmp so that the callee's ret returns straight to the caller.
The receiving side, from the compiler
A variadic function that reads doubles, compiled with -O2 -S -masm=intel (same abridgement as above: directives and the .LFB0/.LFE0 labels removed, ; comments hand-added):
#include <stdarg.h>
double mix(int n, ...) {
va_list ap; va_start(ap, n);
double t = 0;
for (int i = 0; i < n; i++) t += va_arg(ap, double);
va_end(ap);
return t;
}
mix:
sub rsp, 48
test al, al
je .L8
movaps XMMWORD PTR [rsp-88], xmm0
movaps XMMWORD PTR [rsp-72], xmm1
movaps XMMWORD PTR [rsp-56], xmm2
movaps XMMWORD PTR [rsp-40], xmm3
movaps XMMWORD PTR [rsp-24], xmm4
movaps XMMWORD PTR [rsp-8], xmm5
movaps XMMWORD PTR [rsp+8], xmm6
movaps XMMWORD PTR [rsp+24], xmm7
.L8:
lea rax, [rsp+56] ; overflow_arg_area = first stack argument
lea r8, [rsp-136] ; reg_save_area
mov DWORD PTR [rsp-108], 48 ; fp_offset = 48 (no named vector argument)
mov QWORD PTR [rsp-104], rax
mov QWORD PTR [rsp-96], r8
test edi, edi
jle .L7
mov rcx, rax
mov edx, 48
pxor xmm0, xmm0
xor eax, eax
jmp .L4
.L14: ; next double from the register save area
mov esi, edx
add eax, 1
add edx, 16
addsd xmm0, QWORD PTR [r8+rsi]
cmp edi, eax
je .L1
.L4:
cmp edx, 175
jbe .L14
.L5: ; next double from the stack
mov rdx, rcx
add eax, 1
add rcx, 8
addsd xmm0, QWORD PTR [rdx]
cmp edi, eax
jne .L5
.L1:
add rsp, 48
ret
.L7:
pxor xmm0, xmm0
add rsp, 48
ret
Walking the non-obvious lines. movaps stores a 16-byte vector register and requires a 16-byte-aligned address. The save addresses are negative offsets from rsp: the function reserves only 48 bytes with sub rsp, 48, and the save area sits mostly below rsp, in the 128-byte red zone that System V allows a function to use without adjusting rsp (the area is safe here because mix makes no calls; the highest-addressed saves, xmm6 and xmm7, land inside the 48-byte frame). lea r8, [rsp-136] is the start of the 176-byte register save area, where rdi would be at offset 0 and xmm0 at offset 48 (that is [rsp-88], matching the first movaps); the integer part of the area is never written because the function does not read integers. lea rax, [rsp+56] is the first stack-passed argument: 48 bytes of frame plus the 8-byte return address. The three mov stores build the va_list fields: fp_offset = 48, overflow_arg_area and reg_save_area. There is no store for gp_offset (it would be 8, because one named integer argument used rdi): GCC dropped it because this function never reads an integer argument. mov edx, 48 keeps fp_offset in a register for the loop, and .L4 compares it with 175: once it passes the last vector slot (the slot at 160 ends at 176), the loop moves to .L5 and reads from the stack instead.
Here test al, al and je skip all eight vector saves when al is 0. Because this function reads only doubles, GCC dropped the integer register spills it does not use; a function using va_arg(ap, int) would keep them. The loop reads doubles at [r8 + fp_offset], adds 16 to fp_offset each time, and when fp_offset exceeds 175 (the cmp edx, 175 / jbe) it falls through to reading [overflow_arg_area] and advancing that by 8.
Watching the structure move
GCC's va_list can be cast to the psABI structure (GCC on x86-64 only; this is not portable C):
#include <stdarg.h>
#include <stdio.h>
struct va_info {
unsigned int gp_offset;
unsigned int fp_offset;
char *overflow_arg_area;
void *reg_save_area;
};
static void probe(int n, ...) {
va_list ap;
va_start(ap, n);
struct va_info *v = (struct va_info *)ap;
char *start = v->overflow_arg_area;
printf("after va_start: gp_offset=%u fp_offset=%u\n", v->gp_offset, v->fp_offset);
for (int i = 0; i < n; i++) {
int x = va_arg(ap, int);
printf("int #%d = %d: gp_offset=%u, stack bytes consumed=%ld\n",
i + 1, x, v->gp_offset, (long)(v->overflow_arg_area - start));
}
double d = va_arg(ap, double);
printf("double = %.1f: fp_offset=%u\n", d, v->fp_offset);
va_end(ap);
}
int main(void) {
probe(7, 1, 2, 3, 4, 5, 6, 7, 2.5);
return 0;
}
Compiled gcc -O0 -Wall -o va va.c and run in an emulated linux/amd64 container (GCC 14.4):
after va_start: gp_offset=8 fp_offset=48
int #1 = 1: gp_offset=16, stack bytes consumed=0
int #2 = 2: gp_offset=24, stack bytes consumed=0
int #3 = 3: gp_offset=32, stack bytes consumed=0
int #4 = 4: gp_offset=40, stack bytes consumed=0
int #5 = 5: gp_offset=48, stack bytes consumed=0
int #6 = 6: gp_offset=48, stack bytes consumed=8
int #7 = 7: gp_offset=48, stack bytes consumed=16
double = 2.5: fp_offset=64
Reading it: n consumed rdi, so gp_offset starts at 8 (offset of rsi). Five integers (rsi, rdx, rcx, r8, r9) are taken from registers and gp_offset reaches 48. Integers 6 and 7 come from the stack, each advancing overflow_arg_area by 8. The double was in xmm0, found at offset 48, and fp_offset moves to 64 (a vector slot is 16 bytes, even though the double uses 8).
Why the caller sets al, and what goes wrong
- Cost. Without
al, every variadic function would pay to save eight vector registers on every call, evenprintf("%d"). With it, integer-only calls skip them. - Wrong or missing
al. Ifalis smaller than the real number of vector arguments, the callee may not save a register that holds an argument, andva_arg(ap, double)reads stale data. This happens when a variadic function is called through a pointer or prototype that declares fixed parameters (for exampleint (*)(const char *, double)), so the compiler treats the call as non-variadic and never setsal. An old-style declaration with no parameter list (int printf();) is different: GCC 14.4 still setsalfor such a call (it loads 1 for a call with onedouble), because the callee might be variadic. A largeralis harmless. - Alignment. The vector saves are
movaps, which needs 16-byte aligned addresses, so the callee relies on the caller having aligned the stack. With a misaligned stack, a call such asprintf("%f\n", 2.5)(al= 1) faults in glibc's register saves, whileprintf("%d\n", 7)(al= 0) can get through, though any other aligned access in the callee can still fault. - Type mismatches. Arguments through
...undergo the default argument promotions (C's rule that afloatis widened todoubleandcharorshortvalues tointbefore being passed). Passing anintwhereva_arg(ap, double)is read picks the wrong register class entirely.
Microsoft x64 and the other ABIs
Microsoft's x64 convention, as its documentation states, passes the first four arguments in rcx, rdx, r8, r9 (or xmm0 to xmm3 for floating-point values), reserves four 8-byte slots of "shadow space" on the stack for the callee, and has no al convention. For variadic calls, "both the integer register and the floating-point register must contain the value", so a double passed through ... is copied into both, and the fifth and later arguments go on the stack after the shadow space. The al register plays no part in it. Other ABIs (AAPCS64, AAPCS for 32-bit Arm) define their own variadic rules in their own documents.
Unlock Full Question Bank
Get access to all 18 Assembly and Low-Level Language Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.