Language-Level Memory Management (C/C++/Rust) Questions
How C and C++ expose and manage memory at the language level: pointers and pointer arithmetic, pointers to pointers and function pointers, arrays versus pointers and array decay, stack versus heap storage, manual allocation and freeing in C (malloc, realloc, free), new and delete versus malloc and free, placement new, RAII and owning smart pointers (unique_ptr, including custom deleters for C resources), shallow versus deep copies and move semantics in C++. Covers reasoning about who owns a buffer and designing ownership and lifetime contracts across function, library and plugin boundaries; dangling and uninitialized pointers, leaks, double frees, use-after-free, off-by-one errors and buffer overruns and how to prevent them; undefined behavior, strict aliasing and safe byte reinterpretation, endianness; struct layout, alignment and padding as the language defines them; and finding and diagnosing memory bugs with sanitizers, Valgrind-style tools, tracing allocators and heap-corruption triage, including allocator design and heap fragmentation at the language level. Boundary: garbage-collector behavior and tuning, embedded memory budgets, register access and packing structs to hardware layouts, OS virtual memory and paging, and lock-based concurrency are covered elsewhere.
A packet inspection component must parse untrusted bytes without copying them. How do you handle pointer casting, endianness, alignment, and strict aliasing so the code stays correct under compiler optimization and hostile input?
Sample Answer
Direct answer. Never cast the raw byte pointer to a struct pointer and dereference it. Read each field with explicit byte access (shift-and-or for endianness) into a local, size-checked struct, and let the payload stay a plain pointer into the original buffer (no copy). This sidesteps strict-aliasing UB, alignment faults, and endianness bugs all at once, at the cost of one small accessor function per field layout.
Why the obvious cast is wrong, field by field.
- Alignment. A
uint32_tmember inside a struct typically needs 4-byte alignment (the CPU, or some CPUs, can only load a multi-byte value efficiently, or at all, from an address that is a multiple of its size; ARM historically faults on misaligned access for some instructions, x86 is more tolerant but still slower). A byte buffer coming off the wire or out of a capture file has no alignment guarantee: it can start at any offset. Casting(struct Hdr *)bufand reading->stream_idis a misaligned access on a pointer the compiler assumes is properly aligned, which-fsanitize=undefined(UBSan) flags explicitly. - Strict aliasing. The C standard says an object may only be accessed through its "effective type" (the declared type, or a type compatible with it, with a byte-type exception) - so a
uint8_tbuffer reinterpreted as a different struct type violates that rule. Because this is undefined rather than merely risky, the optimizer is allowed to assume the two accesses do not alias and can reorder or eliminate reads, so the cast is undefined behavior even when it happens to work at-O0. - Endianness. Endianness is which end of a multi-byte number is stored first in memory: big-endian stores the most-significant byte at the lowest address (write the number the way you'd write it on paper); little-endian stores the least-significant byte first. The wire format here is big-endian, the convention networking protocols call "network byte order" (per the Internet protocol suite); the host CPU may be little-endian (most x86/ARM machines are). A direct struct read gives you the bytes in host order, silently wrong, with no sanitizer catching it: this bug does not crash, it just produces the wrong number. Concretely, the four wire bytes
0x12 0x34 0x56 0x78are meant to be read as0x12345678; a little-endian host reading them as a rawuint32_twould instead get0x78563412, since it assumes the lowest address holds the least-significant byte. - Hostile input. Untrusted bytes can claim any length field. If you trust a length read via a misaligned/aliased cast before bounds-checking, you can walk past the buffer end on the very next access.
The fix. Read bytes explicitly, not through a pointer cast: (p[0]<<8)|p[1] for a 16-bit big-endian field, the four-shift version for 32-bit. This is alignment-safe (byte access is always aligned), aliasing-safe (no reinterpretation of the underlying object, only arithmetic on uint8_t), and endianness-explicit (the shifts assemble the network-order bytes into a host-order value regardless of host endianness). Validate the length field against the actual buffer length before computing payload, and only ever expose payload as a pointer into the original buffer, never copied ("zero-copy": the payload bytes are read directly from the input buffer instead of being duplicated into a new one), so cost stays O(1) per packet.
Worked example and run. An 8-byte header (version, flags, 16-bit payload length, 32-bit stream id, all big-endian) followed by a payload, parsed from a deliberately misaligned offset the way a real capture buffer would hand it to you. The input bytes: 01 00 00 03 12 34 56 78 61 62 63 (version=1, flags=0, payload_len=3, stream_id bytes 12/34/56/78, payload bytes 61/62/63 = ASCII "abc"), starting one byte into a larger buffer so the header is at a misaligned address:
#include <stddef.h>
#include <stdint.h>
#include <stdio.h>
#include <string.h>
enum { HDR_LEN = 8 };
static uint16_t rd16be(const uint8_t *p) { return (uint16_t)((p[0] << 8) | p[1]); }
static uint32_t rd32be(const uint8_t *p) {
return ((uint32_t)p[0] << 24) | ((uint32_t)p[1] << 16) | ((uint32_t)p[2] << 8) | p[3];
}
typedef struct { uint8_t version, flags; uint16_t payload_len; uint32_t stream_id; } Hdr;
static int parse(const uint8_t *buf, size_t len, Hdr *h, const uint8_t **payload) {
if (len < HDR_LEN) return -1;
h->version = buf[0]; h->flags = buf[1];
h->payload_len = rd16be(buf + 2);
h->stream_id = rd32be(buf + 4);
if (h->version != 1) return -2;
if (h->payload_len > len - HDR_LEN) return -3; /* len >= HDR_LEN already checked */
*payload = buf + HDR_LEN; /* zero copy */
return 0;
}
/* Driver: the header starts one byte into a larger buffer, so it is misaligned. */
typedef struct { uint8_t version, flags; uint16_t payload_len; uint32_t stream_id; } WireHdr;
int main(void) {
uint8_t big[16] = {0};
const uint8_t pkt[11] = {1, 0, 0, 3, 0x12, 0x34, 0x56, 0x78, 0x61, 0x62, 0x63};
memcpy(big + 1, pkt, sizeof pkt);
Hdr h; const uint8_t *pl;
int rc = parse(big + 1, 11, &h, &pl);
printf("rc=%d len=%u stream=0x%x payload=%.*s\n", rc, h.payload_len, h.stream_id, (int)h.payload_len, pl);
printf("claiming only 9 bytes: rc=%d\n", parse(big + 1, 9, &h, &pl));
printf("cast version: 0x%x\n", ((const WireHdr *)(big + 1))->stream_id); /* the bug, for contrast */
return 0;
}
Built with gcc -O2 -fsanitize=address,undefined (GCC 14.4, Linux container, aarch64) and run on the 11-byte packet above starting at a misaligned address, the program prints (UBSan's message for the last line goes to stderr and so appears first, followed by a note: pointer points here line and a byte dump, omitted here) rc=0 len=3 stream=0x12345678 payload=abc, claiming only 9 bytes: rc=-3 and cast version: 0x78563412. So parse returns rc=0 len=3 stream=0x12345678 payload=abc, correctly reconstructing the big-endian stream id (0x12 0x34 0x56 0x78 assembled in that byte order) on a little-endian host, with no sanitizer warning. Feeding it the same buffer but claiming only 9 bytes available (so the declared payload_len=3 would read past it) returns rc=-3 and never touches memory past the buffer. For contrast, a version that does ((const WireHdr *)buf)->stream_id on the same misaligned pointer fails under UBSan with runtime error: member access within misaligned address ... for type 'const struct WireHdr', which requires 4 byte alignment, and separately reads back 0x78563412 instead of 0x12345678 because it never does the byte-order conversion at all, confirming the cast version is broken exactly where the safe version is clean.
Trade-offs and pitfalls. The byte-accessor approach costs a handful of shift/or instructions per field, which GCC at -O2 folds into one load plus one byte-reverse instruction when the access pattern is simple: for rd32be that is ldr + rev on AArch64 and movl + bswap on default x86-64 (movbe, a single load-and-byte-swap instruction, appears only when the target allows it, for example -mmovbe or -march=haswell). Either way it costs about the same as the cast, so the zero-copy win is not undone by slow field reads. The common wrong turn is "just add __attribute__((packed)) to the struct and cast it" (a compiler extension that removes padding so the struct's layout matches the wire bytes): packing fixes the layout mismatch and tells the compiler the fields may be misaligned (so on strict-alignment targets it emits safe, often byte-at-a-time, accesses and UBSan stays quiet), but it does nothing for endianness and still leaves the strict-aliasing violation: the packed variant of the cast on the 11-byte packet above (program below) gives 0x78563412 with no sanitizer warning, i.e. silently wrong. You pay for the byte-wise access and still get the wrong number. If you need to parse the same wire format in many places, generate the accessor functions once (or use a vetted zero-copy parsing library) rather than hand-rolling the shifts at every call site, since a copy-pasted shift with a typo'd bit count is a silent correctness bug that no sanitizer catches.
The packed variant of the cast, as a separate program:
#include <stdint.h>
#include <stdio.h>
#include <string.h>
typedef struct __attribute__((packed)) { uint8_t version, flags; uint16_t payload_len; uint32_t stream_id; } P;
int main(void) {
uint8_t big[16] = {0};
const uint8_t pkt[11] = {1, 0, 0, 3, 0x12, 0x34, 0x56, 0x78, 0x61, 0x62, 0x63};
memcpy(big + 1, pkt, sizeof pkt);
printf("packed cast: 0x%x\n", ((const P *)(big + 1))->stream_id);
return 0;
}
Built with gcc -O2 -fsanitize=address,undefined (GCC 14.4, Linux container, aarch64) it prints packed cast: 0x78563412 with no sanitizer message.
You receive a crash report showing a use-after-free in a high-privilege security agent. Walk through how you would confirm the root cause, reproduce the issue, determine whether it is exploitable, and prioritize a fix without introducing regressions.
Sample Answer
Direct answer. A use-after-free (UAF) is a crash or corruption caused by touching memory through a pointer after it has been freed. The triage is: confirm it is a UAF (not an overflow) with AddressSanitizer, get a minimal reproduction, then assess it as a security bug by what code can run between the free and the stale use, not just whether it crashes today.
Confirming the root cause. A raw core dump (the saved memory image a crashed process leaves behind) from production often just shows a segfault or garbage data, which looks the same for several different heap bugs (use-after-free, double free, heap buffer overflow, the three common ways code goes wrong around malloc/free). Do not trust the crash address alone. Rebuild a debug build with -fsanitize=address,undefined (AddressSanitizer, or ASan, a compiler instrumentation that tracks every allocation; plus UBSan for other undefined behaviour) and run the same input: ASan intercepts every malloc/free, marks freed memory as off-limits in its shadow memory (a side table recording which bytes are valid; this is "poisoning") and parks the block in a quarantine instead of reusing it at once, so a later read is caught even though the bytes themselves are unchanged, and on a bad access prints the faulting access, the exact free() call site, and the original allocation site, all with line numbers (the excerpt in the worked example below is abridged). That turns "crashed somewhere near a session object" into "read of size 8 at the on_close field, freed at line 13, allocated at line 10": an unambiguous UAF, not a guess.
Reproducing it. Heap bugs are often intermittent because the freed block is not reused immediately, so the stale bytes happen to still look valid. Three levers make the bug show up reliably, all of them debugging aids rather than fixes: (1) run under ASan, whose quarantine (the holding area where freed blocks wait instead of being reused) and poisoning make the stale access fail on the first touch however the program's own allocations interleave (the quarantine is bounded in size, so a block that stays freed long enough through heavy allocation traffic can eventually be recycled); (2) on glibc, set the environment variable MALLOC_PERTURB_ to a nonzero value. The mallopt(3) manual page says: "If this parameter is set to a nonzero value, then bytes of allocated memory (other than allocations via calloc(3)) are initialized to the complement of the value in the least significant byte of value, and when allocated memory is released using free(3), the freed bytes are set to the least significant byte of value." The intent is that a stale read sees an obviously wrong pattern instead of plausible old data. Check that it actually bites on the block you care about: on glibc 2.41 (run in a Debian container) a small freed block goes onto glibc's per-thread cache of freed blocks, the tcache, which stores its own bookkeeping in the first 16 bytes of the block, and the fill is skipped for it, so with MALLOC_PERTURB_=165 a freed 32-byte block still held its old contents after those 16 bytes, and the stale call in the example below still ran. A larger block was filled with a5 bytes. Setting GLIBC_TUNABLES=glibc.malloc.tcache_count=0 as well turns the tcache off and the fill then reaches small blocks too (the first 8 bytes still hold the allocator's own list pointer). So treat this lever as a supplement that needs the tunable for small objects, and rely on ASan as the dependable one; (3) if the crash is timing-dependent (one thread frees while another is mid-use), run under ThreadSanitizer (-fsanitize=thread, a compiler instrumentation that detects data races between threads), which reports the free and the use as a race. ASan reports a use that happens after the free; a use that is concurrent with the free is a data race, and that is what TSan is for. Then reduce the input (delta-debugging, or a fuzzer corpus minimizer) until the report still fires with the fewest steps, and keep that input as the regression test.
Assessing security impact. Treat a use-after-free in a high-privilege process as security-relevant, and potentially exploitable, until shown otherwise: freed memory can be reallocated for other data before the stale use, so if an outsider can influence what the program allocates or processes in that window, the stale pointer may read or act on data the outsider shaped. The triage step is to establish, from the code and the ASan report, which code can run between the free() and the stale use, whether any of it handles input an unprivileged party controls (network, files, IPC, plugin data), and whether the stale use is a read, a write or a call through the object. Escalate it to the security team as exploitable until proven otherwise, because proving it is not exploitable needs a bounded argument (the window cannot be reached with outside input) rather than an unsuccessful attempt to trigger it, and the agent's privilege raises the cost of being wrong. Do not build or publish a working exploit to settle the question; the report, the reachability analysis and the fix are what the interview-level answer and a real triage need. Then grep every other code path that holds the same pointer or can run between the two points. A high-privilege agent that processes untrusted input (the scenario given) is exactly the profile where this matters, not just "occasionally crashes".
Worked example. A session object holds a callback pointer. One path frees the session; another still has the old pointer and calls through it.
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
struct session { char user[24]; void (*on_close)(const char *); };
static void say(const char *u) { printf("closing %s\n", u); }
int main(void) {
struct session *s = malloc(sizeof *s); /* 32 bytes requested */
strcpy(s->user, "alice");
s->on_close = say;
free(s); /* path A frees the session ... */
s->on_close(s->user); /* ... path B still calls through */
return 0;
}
Compiled with gcc -g -O0 (no sanitizer; GCC 14 with -Wall already warns pointer 's' used after 'free' for this trivial case, but real bugs span functions and threads, where that warning cannot see them), this does not crash: it prints closing followed by a few garbage bytes that differ from run to run, because nothing has reused the freed 32 bytes yet. glibc's tcache (its per-thread cache of freed blocks) writes its own list bookkeeping into the first 16 bytes of a freed block, which is why the user string comes out as garbage and varies, while the rest of the block, including the callback pointer, is untouched, so the stale call still runs. That is the dangerous case: it looks fine in a quick manual test. Compiled with -fsanitize=address,undefined, the same run reports (abridged: the PID, the stack frames inside the C library and the trailing shadow-memory dump are left out, and the ... stand for elided addresses):
==PID==ERROR: AddressSanitizer: heap-use-after-free on address 0x503000000058
READ of size 8 at 0x503000000058 thread T0
#0 ... in main session.c:14
0x503000000058 is located 24 bytes inside of 32-byte region [...040,...060)
freed by thread T0 here: ... main session.c:13
previously allocated by thread T0 here: ... main session.c:10
pinning the bug to s->on_close(s->user) (the exact source line of the bad read), with the free() site and the original allocation site each reported separately so you can trace the whole lifetime from one report.
Contrast with a different heap bug: if the code instead wrote past the end of a buffer (smashing the allocator's bookkeeping for the neighbouring block), plain glibc typically aborts later, at some unrelated free or malloc, with a message such as "free(): invalid next size" or "double free or corruption", while ASan reports a heap-buffer-overflow immediately at the bad write. The fix is different (bounds-check the write), and the late, distant symptom looks nothing like the UAF's stale read, which is why the first step is to let ASan name the bug class rather than guessing from the crash.
Trade-offs and pitfalls. Prioritizing the fix: a UAF on a high-privilege agent's hot path gets fixed before a cosmetic one even if both "only crash" today, because its security impact can change with unrelated code changes elsewhere in the binary (a new allocation appearing on that path tomorrow). Do not fix it by adding a null check after free without also auditing every other place that holds a copy of the same pointer (an "alias": a second variable holding the same address). The structural fix is to null out every known alias at the single free point, or move to a smart pointer (a wrapper type, such as C++'s std::unique_ptr, that owns the object and destroys it exactly once; it does not make a stale use a compile error, but copying it is rejected at compile time so there is a single owner, and a moved-from or reset pointer is null, so a late use becomes an immediate, easily diagnosed null dereference instead of silent undefined behaviour on freed memory; std::shared_ptr or weak_ptr suit the case where several holders legitimately share the object). Watch for a fix that moves the bug rather than closing it: freeing later (delaying a free to "be safe") can convert a UAF into a leak if the delayed free point is never reached on an error path. Verify the fix by rerunning the same ASan reproduction: a real fix removes the specific ASan report, not just the end-to-end crash.
A C function copies a user-supplied string into a fixed-size buffer and appends a terminator. What checks would you add to prevent off-by-one errors and overflow? Describe the logic in C rather than writing a full program.
Sample Answer
Direct answer
Reserve one byte for the terminator before you copy anything, test the bound before every read and write rather than after, and tell the caller when the input was truncated. Do not rely on strncpy to terminate for you: if the source is as long as the count, it does not (cppreference: "If count is reached before the entire array src was copied, the resulting character array is not null-terminated", https://en.cppreference.com/w/c/string/byte/strncpy).
The checks, in order
- Validate the arguments.
dstnon-NULL anddst_size > 0. A size of 0 has no room even for the terminator, so there is nothing safe to write; return an error and write nothing. - Define the contract in units of the whole buffer.
dst_sizeis the total bytes including the terminator, so the payload room isdst_size - 1. This one line prevents the classic off-by-one, because every later comparison usesroom, notdst_size. - Bound the loop on the destination, not the source. Stop when
i == roomor the source ends, whichever comes first. The source is untrusted and may be any length. - Terminate unconditionally at index
i. Because the loop stops withi <= room,dst[i]is always inside the buffer, and the result is always a valid C string. - Report truncation. If
src[i]is not'\0', input was cut. Whether to reject, truncate silently or log is a policy decision for the caller, but a function that cannot say "truncated" invites silent data corruption (a truncated path or filename can point at a different file).
Worked example
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
/* Copies src into dst (dst_size bytes in total, terminator included).
Returns 0 if everything fit, -1 if src was truncated or an argument was invalid.
When dst_size > 0 and dst is non-NULL, dst is always a valid C string afterwards. */
int copy_str(char *dst, size_t dst_size, const char *src)
{
if (dst == NULL || dst_size == 0)
return -1;
if (src == NULL) {
dst[0] = '\0';
return -1;
}
size_t room = dst_size - 1; /* bytes available for payload; 1 reserved for '\0' */
size_t i = 0;
while (i < room && src[i] != '\0') { /* the bound is tested BEFORE each read and write */
dst[i] = src[i];
i++;
}
dst[i] = '\0'; /* i <= room, so i < dst_size: always in bounds */
return src[i] == '\0' ? 0 : -1; /* anything left in src means truncation */
}
static void try(size_t size, const char *src)
{
char *dst = malloc(size ? size : 1); /* exact-size heap block so ASan sees any overrun */
int rc = copy_str(dst, size, src);
printf("size=%zu src=\"%s\" -> rc=%d dst=\"%s\"\n", size, src, rc, size ? dst : "(none)");
free(dst);
}
#ifdef BUGGY
static void buggy(const char *src)
{
size_t size = 4;
char *dst = malloc(size);
strncpy(dst, src, size); /* copies 4 bytes, no terminator when src is longer */
dst[size] = '\0'; /* off-by-one: index 4 of a 4-byte block */
printf("buggy copy: \"%s\"\n", dst);
free(dst);
}
#endif
int main(void)
{
try(8, "hello"); /* fits with room to spare */
try(6, "hello"); /* exact fit: 5 chars + terminator */
try(5, "hello"); /* one short: truncated to "hell" */
try(1, "hello"); /* only room for the terminator */
try(0, "hello"); /* no room at all */
try(4, ""); /* empty source */
#ifdef BUGGY
buggy("hello");
#endif
return 0;
}
Compiled with gcc -O2 -Wall -Wextra -fsanitize=address,undefined copy.c in a Linux arm64 container (GCC 14.4.0) and run, the output is:
size=8 src="hello" -> rc=0 dst="hello"
size=6 src="hello" -> rc=0 dst="hello"
size=5 src="hello" -> rc=-1 dst="hell"
size=1 src="hello" -> rc=-1 dst=""
size=0 src="hello" -> rc=-1 dst="(none)"
size=4 src="" -> rc=0 dst=""
Each case uses a malloc block of exactly size bytes, so AddressSanitizer would flag any single byte written past the end. The test cases target the boundaries: exact fit (6 bytes for 5 characters), one short, only-terminator room, zero room, empty input.
The bug this prevents
The tempting version is strncpy(dst, src, size); dst[size] = '\0';. Compiling that variant with -DBUGGY (it is in the listing above) at -O2 (GCC 14.4.0, arm64), GCC warned array subscript 4 is outside array bounds of 'char[4]' [-Warray-bounds=] (that warning appears only at -O2; at -O1 GCC instead warns 'strncpy' output truncated copying 4 bytes from a string of length 5 [-Wstringop-truncation], and at -O0 it prints no warning), and with -fsanitize=address the run was reported as a write 0 bytes after a 4-byte region. The correct form of the strncpy idiom writes the terminator at dst[size - 1], but that still pads and scans the full buffer and cannot tell you about truncation, so the explicit loop above is clearer.
Alternatives and when to choose them
snprintf(dst, size, "%s", src): always terminates whensize > 0and returns the length it would have written, soret >= sizemeans truncation. Good default when you are already allowed to usesnprintf; remember a negative return is an error.memcpyafter computingn = strlen(src): only if you already knowsrcis terminated;strlenon an unterminated untrusted buffer is itself an overread. Bound it with a known maximum length first.- Length-carrying types (a pointer plus length, or
std::string_viewin C++) remove the problem where you control the interface.
Trade-offs and pitfalls
- Off-by-one comes from mixing "capacity" and "length" in the same expression; name them separately.
- Check the return value. A bounded copy that nobody checks is a silent truncation bug.
- Validate the content too: a correctly bounded copy of a string with an embedded newline or
../can still be an injection.
How do you securely erase secrets such as keys or passwords from memory in C or C++, and why is that harder than calling memset before freeing the buffer? Include how you decide where such data should live and how many copies to allow.
Sample Answer
Direct answer. Erasing a secret means making sure the bytes that held it are overwritten before the memory is reused, and memset followed by free (or by returning from the function) is not reliable because the compiler is allowed to delete a write that nothing reads afterwards. That optimization is called dead-store elimination: the language only requires the program's observable behaviour to be preserved, and a store to an object that is never read again is not observable, so the compiler may drop it. So you use a function that is specified not to be optimized away, and, because erasure only clears the copy you know about, you also decide where secrets live and keep the number of copies small.
Why memset before free can disappear
Look at what an optimizing compiler produced for a function that wipes a 32-byte key with plain memset. All four variants are built from this file:
#define _DEFAULT_SOURCE
#include <stddef.h>
#include <string.h>
void use_key(const unsigned char *k, size_t n); /* opaque: defined elsewhere */
/* 1. plain memset on a local: the buffer is never read again */
void plain_memset(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset(key, 0, sizeof key);
}
/* 2. explicit_bzero: glibc 2.25+, BSDs */
void with_explicit_bzero(void) {
unsigned char key[32];
use_key(key, sizeof key);
explicit_bzero(key, sizeof key);
}
/* 3. memset through a volatile function pointer: portable fallback */
static void *(*const volatile memset_v)(void *, int, size_t) = memset;
void with_volatile_ptr(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset_v(key, 0, sizeof key);
}
/* 4. GNU/Clang inline-asm barrier after memset */
void with_asm_barrier(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset(key, 0, sizeof key);
__asm__ __volatile__("" : : "r"(key) : "memory");
}
Compiled and disassembled in a Linux container (GCC 14.4) with gcc -O2 -c wipe.c -o wipe.o && objdump -dr --no-show-raw-insn wipe.o. On aarch64 the plain-memset function is as follows (the other three functions in the object file are left out here; the final nop is alignment padding before the next function):
0000000000000000 <plain_memset>:
0: stp x29, x30, [sp, #-48]!
4: mov x1, #0x20 // #32
8: mov x29, sp
c: add x0, sp, #0x10
10: bl 0 <use_key>
10: R_AARCH64_CALL26 use_key
14: ldp x29, x30, [sp], #48
18: ret
1c: nop
How to read this listing (aarch64, the 64-bit Arm instruction set; you do not need to know assembly to check the claim):
stp x29, x30, [sp, #-48]!and the matchingldp ..., [sp], #48at the end save and restore two bookkeeping registers (the frame pointer and the return address) and move the stack pointerspdown 48 bytes and back. That is the function making room for its local variables, including the 32-bytekey.add x0, sp, #0x10puts the address ofkey(16 bytes abovesp) inx0, andmov x1, #0x20puts the length 32 (0x20) inx1. These are the two arguments ofuse_key, passed in registersx0andx1.bl 0 <use_key>is a call (branch with link). TheR_AARCH64_CALL26 use_keyline underneath is only the linker's note that the target address is filled in later.retreturns.
What you are checking is what is missing: between the bl and the ret there is no store to the stack and no call to memset. The memset(key, 0, ...) line in the source produced no instructions at all, so the zeroing is gone. The barrier variant keeps it as two explicit stores of the zero register (xzr always reads as zero; stp stores a pair of registers, so each instruction writes 16 zero bytes), stp xzr, xzr, [sp, #32] and stp xzr, xzr, [sp, #48], after the use_key call and covering all 32 bytes of that function's key; with_explicit_bzero still contains bl explicit_bzero, and with_volatile_ptr an indirect blr x3. The same experiment built for x86-64 (--platform linux/amd64, same compiler and flags) shows the identical outcome: plain_memset has only the call to use_key, with_explicit_bzero calls explicit_bzero, with_volatile_ptr ends in call *%rax, and the barrier version writes zeros with pxor (which clears an SSE register, here %xmm0, to all zero bits) and two movaps stores (each writes the 16 bytes of that register to the stack, 2 x 16 = 32 bytes). This is a property of this compiler, flags and function, which is exactly why the check is to read the optimized disassembly rather than assume.
Why the volatile pointer and the empty asm stop the deletion
Both tricks work by making the compiler unable to prove that the wipe is unobservable.
- Volatile function pointer (
memset_v): the compiler must assume avolatileobject can change at any time, so it has to load the pointer from memory and call whatever address it finds there. It can no longer know the callee ismemset, so it cannot reason that the call only writes to a dead buffer and remove it. The call is kept (blr x3on aarch64,call *%raxon x86-64: a call through a register). - Empty
asmwith a"memory"clobber (__asm__ __volatile__("" : : "r"(key) : "memory")): the quoted string is empty, so it emits no instruction, but it is declared to read the pointerkey(the"r"(key)input) and to possibly read or write any memory (a clobber is a declaration that the asm may change the named thing). The GCC manual saysvolatilestops the compiler deleting the statement or hoisting it out of a loop, though it cautions that the compiler can still move even avolatile asmrelative to other code. What pins the order here is the"memory"clobber together with the"r"(key)input: the asm is declared to readkeyand any memory, so thememsetbefore it might be observed by it, and the zero stores have to stay, which is why the stores reappear in the listing.
Neither is a language guarantee; both are properties of how current GCC and Clang treat these constructs, which is why the table below ranks the standard functions first and why you still read the disassembly.
What to use, in order of preference
| Option | Guarantee | Availability |
|---|---|---|
memset_explicit | C23 standard function; the write is guaranteed to happen | needs a C23 library; check yours |
explicit_bzero | calls "are never optimized away by the compiler" (glibc man page) | glibc 2.25 and later; BSDs |
memset_s | C11 Annex K (an optional bounds-checking extension to the C standard that many libraries do not implement); write guaranteed | not provided by glibc; check __STDC_LIB_EXT1__ |
memset through a volatile function pointer, or memset followed by an empty asm with a "memory" clobber | works in practice, no standard guarantee | portable fallback; verify the disassembly |
The CERT C rule MSC06-C ("Beware of compiler optimizations", one rule in the CERT C Coding Standard, a widely used secure-coding guideline) describes this exact problem, names memset_explicit as the preferred C23 fix, and says to inspect the generated assembly of the optimized release build to confirm the memory is really cleared. Do that in CI for the key-handling functions.
Why even a correct wipe is not enough: where secrets live and how many copies
Zeroing the buffer you hold does not touch other copies. The first three items below are the common failures; the last item is hardening for long-lived keys. Each of these may keep the secret alive:
- Registers and spill slots. A key loaded into registers, or spilled to a stack slot (a temporary stack location where the compiler parks a register value when it runs out of registers), is not part of your buffer.
- Stack frames of callees and temporaries (an inlined copy, a
memcpyinto a local) that are not wiped when the function returns. - Heap copies:
realloccan move a buffer and leave the old block with the secret in it, unwiped. Never grow a secret buffer withrealloc; allocate once at the final size. - Swap, hibernation, core dumps and process forks.
mlockmakes the kernel keep the pages resident so they are not written to swap, but the man page itself warns that laptop suspend saves RAM to disk regardless of locks.madvise(MADV_DONTDUMP)(Linux 3.4+) keeps a range out of core dumps, andMADV_WIPEONFORK(Linux 4.14+) gives a forked child zero-filled pages in that range.mlockcounts againstRLIMIT_MEMLOCKfor unprivileged processes, so size the locked region small.
Decision rules I would commit to:
- Minimize copies: one canonical buffer per secret, passed by pointer, never returned by value and never logged or placed in a string that gets duplicated (
std::stringgrowth, JSON building, error messages). - Choose the storage by lifetime. A short-lived secret (a password during login) can live in a stack buffer that you wipe with
explicit_bzeroormemset_explicitbefore every return path (single exit with a cleanup label); the stack is bounded and not subject torealloc. A long-lived key (a TLS or signing key held for the process lifetime) belongs in one dedicated allocation, ideally page-aligned, locked withmlock, excluded from dumps, and wiped at shutdown. If the platform offers it, a secret-memory facility or hardware (a hardware security module, or keeping the key in a separate process) beats any in-process scheme. - Wipe on every exit path, including error returns, and wipe the intermediate values of the computation, not just the final key.
- Do not rely on wiping alone: assume a memory-disclosure bug (a buffer over-read) exposes whatever is live, so limit how long a secret exists in plaintext.
Trade-offs and pitfalls
mlockcan fail whenRLIMIT_MEMLOCKis low; decide whether that is fatal for your service and check the return value.- Wiping costs a few stores and is irrelevant for performance, so there is no reason to skip it for speed.
- C++ objects: a destructor that wipes helps, but containers copy on growth. Use a fixed buffer owned by one object that is not copyable.
- Compilers can also keep a stale copy in a register or stack slot that only a barrier plus careful design can address; no language-level wipe is a complete defence, which is why the number of copies matters as much as the wipe.
A loop like for (int i = 0; i <= n; ++i) was merged into a C code path where n is the number of valid array elements. The code sometimes works in testing and sometimes crashes in production. What makes this pattern dangerous in C, and how can it turn into a harder-to-debug memory bug than a simple out-of-bounds read?
Sample Answer
Direct answer. In C, an array of n elements has valid indexes 0 through n - 1. The loop for (int i = 0; i <= n; ++i) runs one time too many and touches element n, one past the end. That is undefined behaviour (UB: the language gives no guarantee, so the program may crash, work, or misbehave later). Because it often lands on memory that is legitimately owned by something else, the result depends on layout, compiler options and the allocator, which is why it passes tests and fails in production. It is worse than an out-of-bounds read because a write silently corrupts a neighbour: another field, a saved register (a value the CPU parked on the stack to restore when the function returns), or the allocator's own bookkeeping (the chunk headers described in layout 3 below), and the failure appears later, somewhere unrelated.
Why the symptoms vary
Three real layouts, each tested below, show the same bug looking completely different.
1. Neighbour inside the same object. If the array is followed by another field, the extra write lands on it, and nothing is "out of bounds" as far as the hardware or the heap is concerned.
#include <stdio.h>
#include <stdlib.h>
/* Neighbour inside the same object: the overrun lands on a field that is "valid memory". */
struct stats {
int samples[4];
int limit; /* sits right after the array */
};
int main(void) {
struct stats s = { .samples = {0}, .limit = 100 };
int n = 4; /* 4 valid elements: indexes 0..3 */
for (int i = 0; i <= n; ++i) /* BUG: writes samples[4] */
s.samples[i] = -1;
printf("limit after loop = %d (expected 100)\n", s.limit);
return 0;
}
Built and run in a Linux container (GCC 14.4, aarch64):
$ gcc -O0 overrun_struct.c -o a0 && ./a0
limit after loop = -1 (expected 100)
$ gcc -O2 overrun_struct.c -o a2 && ./a2
overrun_struct.c: In function 'main':
overrun_struct.c:14:22: warning: iteration 4 invokes undefined behavior [-Waggressive-loop-optimizations]
14 | s.samples[i] = -1;
| ~~~~~~~~~~~~~^~~~
overrun_struct.c:13:23: note: within this loop
13 | for (int i = 0; i <= n; ++i) /* BUG: writes samples[4] */
| ~~^~~~
limit after loop = 100 (expected 100)
$ gcc -O2 -g -fsanitize=address,undefined overrun_struct.c -o aa && ./aa
overrun_struct.c:14:18: runtime error: index 4 out of bounds for type 'int [4]'
limit after loop = 100 (expected 100)
The same source corrupts limit at -O0 and appears correct at -O2. At -O2 the compiler assumes the store to samples[4] cannot be a legal access to limit and may assume the undefined iteration never happens (GCC says so in its warning); in the generated aarch64 code the printf argument is simply the constant 100 (mov w1, 100), so limit is never even reloaded from the corrupted memory. A test build and a release build can therefore disagree. Note that UBSan (UndefinedBehaviorSanitizer, which checks array bounds when it knows the array size) caught this one, while ASan, which watches heap, stack-frame and global boundaries and not fields inside one object, did not.
2. Slack in the heap block. malloc rounds up. Measured on glibc 2.41 (the Debian 13 container, both aarch64 and x86-64):
#include <malloc.h>
#include <stdio.h>
#include <stdlib.h>
int main(void) {
int n = 4;
for (int bytes = 16; bytes <= 40; bytes += 8) {
void *p = malloc((size_t)bytes);
printf("malloc(%d) usable size = %zu\n", bytes, malloc_usable_size(p));
free(p);
}
(void)n;
return 0;
}
malloc(16) usable size = 24
malloc(24) usable size = 24
malloc(32) usable size = 40
malloc(40) usable size = 40
(gcc -O2 usable_size.c -o c && ./c.) An array of four int (16 bytes) has 8 spare bytes, so writing heap[4] lands in slack and nothing visibly breaks. That is the "works in testing" case.
3. No slack: the allocator's own header. glibc hands out memory in chunks: each block you get from malloc is preceded by a small header (the allocator's bookkeeping) holding, among other things, the block's size, which free reads to know what it is releasing. Laid out in memory: [header: size][your 24 bytes][next chunk's header: size][next block ...]. With six int (24 bytes, exactly the usable size), element 6 overwrites the first bytes of the next chunk's header:
#include <stdio.h>
#include <stdlib.h>
int main(void) {
int n = 6; /* 6 ints = 24 bytes = the whole usable size */
int *a = malloc(n * sizeof *a);
int *b = malloc(n * sizeof *b);
if (!a || !b) return 1;
for (int i = 0; i <= n; ++i) /* BUG: a[6] lands on the next chunk's size field */
a[i] = -1;
printf("loop finished, nothing crashed yet\n");
fflush(stdout);
free(a);
printf("free(a) returned\n");
fflush(stdout);
free(b);
printf("freed\n");
return 0;
}
$ gcc -O0 overrun_heap.c -o d && ./d
loop finished, nothing crashed yet
free(a) returned
munmap_chunk(): invalid pointer
Aborted
$ gcc -O0 -g -fsanitize=address overrun_heap.c -o da && ./da
==18==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x503000000058 at pc ...
WRITE of size 4 at 0x503000000058 thread T0
#0 0x400ac4 in main /w/overrun_heap.c:10
The loop and free(a) both succeed. The process dies at free(b), because the -1 clobbered the size field that free trusts. munmap_chunk(): invalid pointer is glibc's message for a block whose recorded size and flags made it take the wrong release path and then fail its own sanity check; the text means "the header of this block is not what I wrote", and it says nothing about which line damaged it. The crash is in a different statement from the bug, in code that looks correct, and the message is glibc-version specific. This is why the bug is harder to diagnose than a bad read: a read returns a wrong value at the faulty line, a write breaks something that is used later. On the stack the same overrun can hit a neighbouring local, a saved frame pointer (the caller's frame address, kept on the stack) or a return address, which is why MITRE's CWE-193 (an entry in the Common Weakness Enumeration, the public catalogue of software weakness types; this one is the off-by-one error) notes that it can lead to memory corruption and sometimes arbitrary code execution, not just a crash.
How to find it and prevent it
- Compile with
-Wall -Wextra -O2and read warnings; the-Waggressive-loop-optimizationsmessage above names the exact iteration. - Run tests under ASan and UBSan (
-fsanitize=address,undefined), and make sure at least one test runs with a full-size input, since the bug needs the last iteration to execute. - Reproduce in the debugger with a hardware watchpoint (a CPU feature that halts the program the moment a chosen address is written) on the corrupted field (
watch s.limitin gdb orwatchpoint set variablein lldb) to see which line writes it. - Fix the bound,
i < n, and prefersize_tfor indexes and counts so a negativencannot slip through a signed comparison. Derivenfromsizeof a / sizeof a[0]when the array is local, so the count cannot drift from the declaration.
Trade-offs and pitfalls
A condition written <= n is correct when n is the last valid index and wrong when it is the count, so name variables count or last and the loop reads correctly. Adding a canary or padding to hide the symptom is not a fix: it only moves where the write lands. Treat "works on my machine, at this optimization level" as no evidence at all for code with UB.
Unlock Full Question Bank
Get access to all 8 Language-Level Memory Management (C/C++/Rust) interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.