Language-Level Memory Management (C/C++/Rust) Questions
How C and C++ expose and manage memory at the language level: pointers and pointer arithmetic, pointers to pointers and function pointers, arrays versus pointers and array decay, stack versus heap storage, manual allocation and freeing in C (malloc, realloc, free), new and delete versus malloc and free, placement new, RAII and owning smart pointers (unique_ptr, including custom deleters for C resources), shallow versus deep copies and move semantics in C++. Covers reasoning about who owns a buffer and designing ownership and lifetime contracts across function, library and plugin boundaries; dangling and uninitialized pointers, leaks, double frees, use-after-free, off-by-one errors and buffer overruns and how to prevent them; undefined behavior, strict aliasing and safe byte reinterpretation, endianness; struct layout, alignment and padding as the language defines them; and finding and diagnosing memory bugs with sanitizers, Valgrind-style tools, tracing allocators and heap-corruption triage, including allocator design and heap fragmentation at the language level. Boundary: garbage-collector behavior and tuning, embedded memory budgets, register access and packing structs to hardware layouts, OS virtual memory and paging, and lock-based concurrency are covered elsewhere.
What are the main techniques to prevent and detect buffer overflows in C beyond swapping in safer-looking library functions? Cover both coding practice and build-time or runtime defenses.
Sample Answer
Direct answer
Defense for buffer overflows has four layers: write code where the length always travels with the pointer and is checked before every copy; get the compiler to warn and to insert checks (-Wall -Wextra, -D_FORTIFY_SOURCE, -fstack-protector-strong); catch the bugs you did not see in review with AddressSanitizer plus fuzzing before release; and keep OS-level mitigations on as a last line that turns many exploits into crashes. (A fuzzer is a program that feeds a target huge numbers of generated or mutated inputs looking for crashes; fuzzing is running one.) No single layer is enough, and the safe-looking library function swap is the weakest of them.
Why "safer functions" alone fall short
strncpy does not terminate when the source fills the count (https://en.cppreference.com/w/c/string/byte/strncpy). Any bounded function still needs the correct bound, and the common bug is passing the wrong number (the pointer's size, the source length, or the buffer size without the terminator). The fix is structural: make the size impossible to forget.
Layer 1: coding practice
- Pass pointer and length together (a small
struct buf { uint8_t *p; size_t len; }), and checklenbefore copying, not after. - Use
size_tfor sizes and check arithmetic before it happens:if (n > cap - used) return ERR;rather thanused + n > cap, which can wrap. - Prefer functions that report truncation (
snprintfreturns the length it wanted), and check the result. - Allocate with the size computed from the same constant used to bound the copy, so the two cannot drift apart.
- Parse by validating the length field against what you actually received, then copy.
Layer 2: build-time defenses
-Wall -Wextrafor static diagnostics. They are not complete: in the example below, a plainstrcpyinto an 8-byte array fromargv[1]compiled with no warning.-fstack-protector-strongputs a canary (a guard value) between local arrays and the saved return address and checks it on function exit. Picture the stack frame asbuf[8], then the canary, then the saved return address (the place the function jumps back to when it ends): an overflow that runs pastbuftoward the return address must overwrite the canary first, and the check at exit sees the changed value and aborts instead of returning to an attacker-chosen address. GCC defines it as-fstack-protectorplus functions that have local arrays or reference local frame addresses (GCC manual: https://gcc.gnu.org/onlinedocs/gcc/Instrumentation-Options.html).-D_FORTIFY_SOURCE: glibc and the compiler add lightweight checks to string and memory functions when the destination size is known. The recommendation below uses level 2. Level 1 needs-O1or higher, level 2 adds more checks (some conforming programs can fail), and level 3 adds checks for buffers whose size is only known at run time (for example frommalloc(n)) and needs GCC 12 or later with glibc 2.33 or later (https://man7.org/linux/man-pages/man7/feature_test_macros.7.html).
Layer 3: test-time detection
-fsanitize=address instruments memory accesses to detect out-of-bounds and use-after-free bugs. Combine with -fsanitize=undefined, then run unit tests and a fuzzer so the sanitizer sees hostile inputs. This is the layer that finds the bug rather than merely containing it.
Worked example: one overflow, four builds
#include <stdio.h>
#include <string.h>
int main(int argc, char **argv)
{
char buf[8];
if (argc < 2) return 1;
strcpy(buf, argv[1]); /* no length check: overflows for inputs of 8 or more chars */
printf("copied: %s\n", buf);
return 0;
}
Run with a 32-character argument of A in a Linux arm64 container (GCC 14.4.0):
gcc -O0 -fno-stack-protector -> copied: AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (exit 0, silent corruption)
gcc -O0 -fstack-protector-strong -> *** stack smashing detected ***: terminated (Aborted)
gcc -O2 -D_FORTIFY_SOURCE=2 -fno-stack-protector -> *** buffer overflow detected ***: terminated (Aborted)
gcc -O1 -g -fsanitize=address -> ERROR: AddressSanitizer: stack-buffer-overflow ... WRITE of size 33
The unprotected build overflowed an 8-byte array with 33 bytes (the 32 A characters plus the terminating zero byte that strcpy also writes, which is where WRITE of size 33 in the sanitizer line comes from), printed success and exited 0: this is why overflows are dangerous, since nothing signalled the corruption. The other three builds stopped it, at different costs and with different information (the sanitizer names the access, the production flags just abort).
Layer 4: OS and platform mitigations
Two mitigations are normally on by default. A non-executable stack marks stack memory as data, so an attacker who injects machine code into an overflowed buffer cannot simply run it. Address-space randomization (ASLR) loads the program, libraries, stack and heap at different addresses on each run, so an attacker cannot hard-code the address of the code they want to jump to. They make exploiting a surviving overflow harder; they are not a fix and can be bypassed (for example, an information leak reveals the randomized addresses), so treat them as damage limiters.
Recommendation
For a new C codebase handling untrusted data: length-carrying buffer types and checked arithmetic in the code, -Wall -Wextra -fstack-protector-strong -D_FORTIFY_SOURCE=2 in every release build, and a CI job that runs tests and a fuzzer under -fsanitize=address,undefined. If it is a hot, high-exposure parser, consider a memory-safe language for that component. What would change this: on a tiny embedded target without room for canaries, lean harder on coding rules, static analysis and host-side sanitizer testing of the same source.
Pitfalls
- Hardening flags are not tests: a program that aborts on overflow has still got the bug.
- Sanitizers slow programs and use more memory, so they belong in test builds, not production.
- Heap overflows, off-by-one writes inside a struct, and overruns that stay within one allocation may not trip a stack canary at all.
What is undefined behavior in C and C++, and why is it especially dangerous in code that handles untrusted input? Give examples that look harmless in review but can become exploitable or non-deterministic at runtime.
Sample Answer
Direct answer
Undefined behavior (UB) means the language standard places no requirements on what the program does once an execution hits certain constructs. The program may crash, may appear to work, or may do something that changes from one compiler version or flag to the next. It is especially dangerous with untrusted input because the attacker chooses the values, and the compiler is allowed to optimize on the assumption that UB never happens, so the checks you wrote can disappear and the memory errors that remain can be steered.
What UB is, precisely
The standard does not say "this crashes". It says there are no restrictions on behavior, and compilers are not required to diagnose it (cppreference, "Undefined behavior": https://en.cppreference.com/w/c/language/behavior). Two consequences matter in practice:
- The compiler reasons backwards from the assumption that UB cannot occur. If a branch can only be reached by overflowing a signed integer, the compiler may treat that branch as dead and delete it.
- Behavior is not stable. A bug that "works" at one optimization level, on one compiler version, or with one stack layout can break after an unrelated change.
Common UB, why it looks harmless in review, and a mitigation for each
| UB | Why a reviewer passes it | What goes wrong with attacker input | Mitigation |
|---|---|---|---|
| Out-of-bounds read or write (including off-by-one) | The loop "obviously" stops at the end | Write past the end overwrites neighbouring data, a saved return address (the stack slot a function jumps back to when it returns, so overwriting it redirects execution) or a function pointer (a variable holding the address of code to call); a read past the end leaks memory | Carry the length with the pointer, check len against the size before every copy, build tests with AddressSanitizer |
| Use-after-free | The pointer is still in scope and compiles fine | The freed block is reallocated to attacker-influenced data, so the stale pointer now reads or writes someone else's object | Set the pointer to NULL after free, give every block one clear owner, test with AddressSanitizer |
| Double free | Two cleanup paths each look correct | Corrupts the allocator's bookkeeping, which can become a write primitive (a flaw that lets an attacker write chosen bytes to chosen addresses) | One owner frees; NULL after free; AddressSanitizer reports it |
| Signed integer overflow | a + b on int looks like plain arithmetic | A size computed as count * width wraps, a small buffer is allocated, and a later copy overruns it. The compiler may also delete a check written as x + 1 > x | Use size_t for sizes; test before the operation (a > INT_MAX - b) or use __builtin_add_overflow (GCC and Clang) |
| Strict-aliasing violation | A cast from char * to uint32_t * "just reinterprets bytes" | Reading an object through an incompatible type is UB, so the optimizer may reorder or drop the load and use stale data | Copy with memcpy into a correctly typed object; access through char types, which the rule exempts (https://en.cppreference.com/w/c/language/object) |
| Null dereference before the null check | The check is right there, just one line late | The compiler assumes the pointer is non-null after the dereference and removes the check | Check first, dereference second |
| Reading an uninitialized variable | It "is always zero in testing" | Contents come from whatever the stack held earlier, which can be attacker influenced | -Wall -Wextra warnings, and initialize at declaration |
Strict aliasing in one example: the table's char * to uint32_t * cast reads bytes through the wrong type. The defined way is to copy the bytes into an object of the right type. Built as a complete program with gcc -Wall -Wextra (no warnings) in an arm64 Linux container (GCC 14.4.0; little-endian, so the first byte is the least significant), it prints v = 1:
#include <stdint.h>
#include <stdio.h>
#include <string.h>
int main(void) {
unsigned char bytes[4] = { 1, 0, 0, 0 };
uint32_t v;
memcpy(&v, bytes, sizeof v); /* defined: copy the bytes into a real uint32_t */
printf("v = %u\n", v); /* prints: v = 1 */
return 0;
}
Unsigned arithmetic is the contrast case: it is defined to wrap modulo 2^n, so UINT_MAX + 1 is 0, not UB. Signed overflow is UB and may wrap, trap, or be optimized out (https://en.cppreference.com/w/c/language/operator_arithmetic). That asymmetry is why size and offset math belongs in size_t.
Worked example 1: a guard the compiler removes
#include <limits.h>
#include <stdio.h>
/* Looks like an overflow guard. Signed overflow is UB, so the compiler may fold it (replace it by the constant) to 1. */
int will_not_wrap(int x) { return x + 1 > x; }
int main(void) {
printf("INT_MAX + 1 > INT_MAX -> %d\n", will_not_wrap(INT_MAX));
return 0;
}
Compiled and run in a Linux arm64 container with GCC 14.4.0:
gcc -O0 ub.c -> INT_MAX + 1 > INT_MAX -> 1
gcc -O2 ub.c -> INT_MAX + 1 > INT_MAX -> 1
gcc -O2 -fwrapv ub.c -> INT_MAX + 1 > INT_MAX -> 0
At -O2 the whole function compiled to mov w0, 1; ret. On arm64, w0 is the 32-bit register that carries a function's return value (the Arm AAPCS64 calling standard uses r0 to r7 for arguments and results), so this reads "put 1 in the return register, then return": it never even looks at x, and it returns true for every input, so the "overflow guard" guards nothing. With -fwrapv (tell GCC signed overflow wraps) the result is the 2's-complement answer, 0. (2's-complement is the usual way machines store signed integers: adding 1 to the largest value gives the most negative one, so INT_MAX + 1 > INT_MAX is false.) The point is not that wrapping is "right", it is that the program's meaning depended on a compiler assumption nobody wrote down. Notice also that on this GCC the check was folded away even at -O0: do not assume "debug builds are safe".
Worked example 2: a null check deleted
#include <stdio.h>
struct pkt { int len; };
/* Dereference first, check later: the check can be deleted. */
int get_len(struct pkt *p) {
int n = p->len;
if (p == NULL) return -1;
return n;
}
int main(void) {
struct pkt a = { 42 };
printf("%d\n", get_len(&a));
return 0;
}
Built with gcc -O2 -S for arm64, get_len came out as just ldr w0, [x0] then ret. x0 holds the first argument, which is p, so ldr w0, [x0] means "load the 4 bytes at the address in p into the return register", that is p->len, and ret returns it. The p == NULL test is gone, because after p->len the compiler may assume p is not null. In a real parser the same shape appears when a length field is read from a header before the header pointer is validated.
Worked example 3: catching it at runtime
#include <limits.h>
#include <stdio.h>
int main(int argc, char **argv) {
(void)argv;
int x = INT_MAX - 1 + argc; /* argc is 1 when run with no arguments: INT_MAX */
int y = x + argc; /* INT_MAX + 1: signed overflow, undefined */
printf("y = %d\n", y);
return 0;
}
gcc -O1 -fsanitize=undefined ub3.c && ./ub3 prints ub3.c:7:9: runtime error: signed integer overflow: 1 + 2147483647 cannot be represented in type 'int' and then y = -2147483648. With -fno-sanitize-recover=undefined it stops with exit status 1 instead of continuing. -fsanitize=undefined instruments various computations to detect UB at runtime, and -fsanitize=signed-integer-overflow is the specific overflow check (GCC manual, Instrumentation Options: https://gcc.gnu.org/onlinedocs/gcc/Instrumentation-Options.html).
Detection and prevention toolbox
- Compile-time:
-Wall -Wextra(and the compiler's static analyzer, or clang-tidy / Coverity-class tools). They catch only some patterns, so a clean compile is not evidence that the code is free of UB. - Test-time: build unit tests and fuzz targets (fuzzing feeds the code huge numbers of generated inputs; a fuzz target is the small function that receives each one) with
-fsanitize=address,undefined. AddressSanitizer finds out-of-bounds and use-after-free; UBSan finds overflow, bad shifts, misaligned access. Fuzzing with attacker-shaped input is what makes the sanitizers find input-dependent bugs. - Production hardening (reduces impact, does not remove the bug). Treat the layers above as first priority and these as the last line:
-fstack-protector-strong: GCC places a guard value (a canary) next to the return address in functions with local arrays or addresses of locals, and checks it before returning; a stack overflow that overwrites it aborts the program.-D_FORTIFY_SOURCE=2: with optimization on, glibc adds lightweight checks to string and memory functions such asmemcpyandstrcpyto detect some buffer overflows (per the glibc feature-test-macros man page).- A non-executable stack: the memory holding locals cannot be run as code, so injected bytes cannot simply be executed.
- Address randomization (ASLR): the kernel places the stack, heap and libraries at unpredictable addresses (
randomize_va_spacein the Linux kernel documentation), so an attacker cannot hard-code them. These turn many overflows into a crash, which is a denial of service rather than code execution, but the underlying UB is still there.
- Language choice: for new parsers of untrusted data, a memory-safe language removes whole classes of the table above.
Constrained environments (embedded, no OS)
Sanitizers need runtime support and memory, which a small device usually lacks. The usual pattern is: keep protocol parsing and buffer logic in plain portable C, compile it on a host with the sanitizers and a fuzzer, then cross-compile the same source for the target with warnings as errors and static analysis. Where you cannot avoid hardware-specific behavior, isolate it in one small file so the rest stays testable.
Trade-offs and pitfalls
-fwrapvand-fno-strict-aliasing(the second tells the compiler not to assume that pointers of unrelated types point to different memory) are safety nets that make some UB behave predictably; they do not make the code correct, and they cost some optimization. Use them to buy time, then fix the code.- "It passes the test suite" proves nothing about UB: the optimizer decides, not the test. The same source can behave differently on a new compiler release.
- Do not write "defensive" checks that themselves rely on UB (
if (x + 1 < x)). Check the operands before the operation.
You receive a crash report showing a use-after-free in a high-privilege security agent. Walk through how you would confirm the root cause, reproduce the issue, determine whether it is exploitable, and prioritize a fix without introducing regressions.
Sample Answer
Direct answer. A use-after-free (UAF) is a crash or corruption caused by touching memory through a pointer after it has been freed. The triage is: confirm it is a UAF (not an overflow) with AddressSanitizer, get a minimal reproduction, then assess it as a security bug by what code can run between the free and the stale use, not just whether it crashes today.
Confirming the root cause. A raw core dump (the saved memory image a crashed process leaves behind) from production often just shows a segfault or garbage data, which looks the same for several different heap bugs (use-after-free, double free, heap buffer overflow, the three common ways code goes wrong around malloc/free). Do not trust the crash address alone. Rebuild a debug build with -fsanitize=address,undefined (AddressSanitizer, or ASan, a compiler instrumentation that tracks every allocation; plus UBSan for other undefined behaviour) and run the same input: ASan intercepts every malloc/free, marks freed memory as off-limits in its shadow memory (a side table recording which bytes are valid; this is "poisoning") and parks the block in a quarantine instead of reusing it at once, so a later read is caught even though the bytes themselves are unchanged, and on a bad access prints the faulting access, the exact free() call site, and the original allocation site, all with line numbers (the excerpt in the worked example below is abridged). That turns "crashed somewhere near a session object" into "read of size 8 at the on_close field, freed at line 13, allocated at line 10": an unambiguous UAF, not a guess.
Reproducing it. Heap bugs are often intermittent because the freed block is not reused immediately, so the stale bytes happen to still look valid. Three levers make the bug show up reliably, all of them debugging aids rather than fixes: (1) run under ASan, whose quarantine (the holding area where freed blocks wait instead of being reused) and poisoning make the stale access fail on the first touch however the program's own allocations interleave (the quarantine is bounded in size, so a block that stays freed long enough through heavy allocation traffic can eventually be recycled); (2) on glibc, set the environment variable MALLOC_PERTURB_ to a nonzero value. The mallopt(3) manual page says: "If this parameter is set to a nonzero value, then bytes of allocated memory (other than allocations via calloc(3)) are initialized to the complement of the value in the least significant byte of value, and when allocated memory is released using free(3), the freed bytes are set to the least significant byte of value." The intent is that a stale read sees an obviously wrong pattern instead of plausible old data. Check that it actually bites on the block you care about: on glibc 2.41 (run in a Debian container) a small freed block goes onto glibc's per-thread cache of freed blocks, the tcache, which stores its own bookkeeping in the first 16 bytes of the block, and the fill is skipped for it, so with MALLOC_PERTURB_=165 a freed 32-byte block still held its old contents after those 16 bytes, and the stale call in the example below still ran. A larger block was filled with a5 bytes. Setting GLIBC_TUNABLES=glibc.malloc.tcache_count=0 as well turns the tcache off and the fill then reaches small blocks too (the first 8 bytes still hold the allocator's own list pointer). So treat this lever as a supplement that needs the tunable for small objects, and rely on ASan as the dependable one; (3) if the crash is timing-dependent (one thread frees while another is mid-use), run under ThreadSanitizer (-fsanitize=thread, a compiler instrumentation that detects data races between threads), which reports the free and the use as a race. ASan reports a use that happens after the free; a use that is concurrent with the free is a data race, and that is what TSan is for. Then reduce the input (delta-debugging, or a fuzzer corpus minimizer) until the report still fires with the fewest steps, and keep that input as the regression test.
Assessing security impact. Treat a use-after-free in a high-privilege process as security-relevant, and potentially exploitable, until shown otherwise: freed memory can be reallocated for other data before the stale use, so if an outsider can influence what the program allocates or processes in that window, the stale pointer may read or act on data the outsider shaped. The triage step is to establish, from the code and the ASan report, which code can run between the free() and the stale use, whether any of it handles input an unprivileged party controls (network, files, IPC, plugin data), and whether the stale use is a read, a write or a call through the object. Escalate it to the security team as exploitable until proven otherwise, because proving it is not exploitable needs a bounded argument (the window cannot be reached with outside input) rather than an unsuccessful attempt to trigger it, and the agent's privilege raises the cost of being wrong. Do not build or publish a working exploit to settle the question; the report, the reachability analysis and the fix are what the interview-level answer and a real triage need. Then grep every other code path that holds the same pointer or can run between the two points. A high-privilege agent that processes untrusted input (the scenario given) is exactly the profile where this matters, not just "occasionally crashes".
Worked example. A session object holds a callback pointer. One path frees the session; another still has the old pointer and calls through it.
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
struct session { char user[24]; void (*on_close)(const char *); };
static void say(const char *u) { printf("closing %s\n", u); }
int main(void) {
struct session *s = malloc(sizeof *s); /* 32 bytes requested */
strcpy(s->user, "alice");
s->on_close = say;
free(s); /* path A frees the session ... */
s->on_close(s->user); /* ... path B still calls through */
return 0;
}
Compiled with gcc -g -O0 (no sanitizer; GCC 14 with -Wall already warns pointer 's' used after 'free' for this trivial case, but real bugs span functions and threads, where that warning cannot see them), this does not crash: it prints closing followed by a few garbage bytes that differ from run to run, because nothing has reused the freed 32 bytes yet. glibc's tcache (its per-thread cache of freed blocks) writes its own list bookkeeping into the first 16 bytes of a freed block, which is why the user string comes out as garbage and varies, while the rest of the block, including the callback pointer, is untouched, so the stale call still runs. That is the dangerous case: it looks fine in a quick manual test. Compiled with -fsanitize=address,undefined, the same run reports (abridged: the PID, the stack frames inside the C library and the trailing shadow-memory dump are left out, and the ... stand for elided addresses):
==PID==ERROR: AddressSanitizer: heap-use-after-free on address 0x503000000058
READ of size 8 at 0x503000000058 thread T0
#0 ... in main session.c:14
0x503000000058 is located 24 bytes inside of 32-byte region [...040,...060)
freed by thread T0 here: ... main session.c:13
previously allocated by thread T0 here: ... main session.c:10
pinning the bug to s->on_close(s->user) (the exact source line of the bad read), with the free() site and the original allocation site each reported separately so you can trace the whole lifetime from one report.
Contrast with a different heap bug: if the code instead wrote past the end of a buffer (smashing the allocator's bookkeeping for the neighbouring block), plain glibc typically aborts later, at some unrelated free or malloc, with a message such as "free(): invalid next size" or "double free or corruption", while ASan reports a heap-buffer-overflow immediately at the bad write. The fix is different (bounds-check the write), and the late, distant symptom looks nothing like the UAF's stale read, which is why the first step is to let ASan name the bug class rather than guessing from the crash.
Trade-offs and pitfalls. Prioritizing the fix: a UAF on a high-privilege agent's hot path gets fixed before a cosmetic one even if both "only crash" today, because its security impact can change with unrelated code changes elsewhere in the binary (a new allocation appearing on that path tomorrow). Do not fix it by adding a null check after free without also auditing every other place that holds a copy of the same pointer (an "alias": a second variable holding the same address). The structural fix is to null out every known alias at the single free point, or move to a smart pointer (a wrapper type, such as C++'s std::unique_ptr, that owns the object and destroys it exactly once; it does not make a stale use a compile error, but copying it is rejected at compile time so there is a single owner, and a moved-from or reset pointer is null, so a late use becomes an immediate, easily diagnosed null dereference instead of silent undefined behaviour on freed memory; std::shared_ptr or weak_ptr suit the case where several holders legitimately share the object). Watch for a fix that moves the bug rather than closing it: freeing later (delaying a free to "be safe") can convert a UAF into a leak if the delayed free point is never reached on an error path. Verify the fix by rerunning the same ASan reproduction: a real fix removes the specific ASan report, not just the end-to-end crash.
That is every published Language-Level Memory Management (C/C++/Rust) question for Penetration Tester so far. Browse the other topics in this category, or practice this one interactively.