Direct answer. Erasing a secret means making sure the bytes that held it are overwritten before the memory is reused, and memset followed by free (or by returning from the function) is not reliable because the compiler is allowed to delete a write that nothing reads afterwards. That optimization is called dead-store elimination: the language only requires the program's observable behaviour to be preserved, and a store to an object that is never read again is not observable, so the compiler may drop it. So you use a function that is specified not to be optimized away, and, because erasure only clears the copy you know about, you also decide where secrets live and keep the number of copies small.
Why memset before free can disappear
Look at what an optimizing compiler produced for a function that wipes a 32-byte key with plain memset. All four variants are built from this file:
c
#define _DEFAULT_SOURCE
#include <stddef.h>
#include <string.h>
void use_key(const unsigned char *k, size_t n); /* opaque: defined elsewhere */
/* 1. plain memset on a local: the buffer is never read again */
void plain_memset(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset(key, 0, sizeof key);
}
/* 2. explicit_bzero: glibc 2.25+, BSDs */
void with_explicit_bzero(void) {
unsigned char key[32];
use_key(key, sizeof key);
explicit_bzero(key, sizeof key);
}
/* 3. memset through a volatile function pointer: portable fallback */
static void *(*const volatile memset_v)(void *, int, size_t) = memset;
void with_volatile_ptr(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset_v(key, 0, sizeof key);
}
/* 4. GNU/Clang inline-asm barrier after memset */
void with_asm_barrier(void) {
unsigned char key[32];
use_key(key, sizeof key);
memset(key, 0, sizeof key);
__asm__ __volatile__("" : : "r"(key) : "memory");
}
Compiled and disassembled in a Linux container (GCC 14.4) with gcc -O2 -c wipe.c -o wipe.o && objdump -dr --no-show-raw-insn wipe.o. On aarch64 the plain-memset function is as follows (the other three functions in the object file are left out here; the final nop is alignment padding before the next function):
0000000000000000 <plain_memset>:
0: stp x29, x30, [sp, #-48]!
4: mov x1, #0x20 // #32
8: mov x29, sp
c: add x0, sp, #0x10
10: bl 0 <use_key>
10: R_AARCH64_CALL26 use_key
14: ldp x29, x30, [sp], #48
18: ret
1c: nop
How to read this listing (aarch64, the 64-bit Arm instruction set; you do not need to know assembly to check the claim):
stp x29, x30, [sp, #-48]! and the matching ldp ..., [sp], #48 at the end save and restore two bookkeeping registers (the frame pointer and the return address) and move the stack pointer sp down 48 bytes and back. That is the function making room for its local variables, including the 32-byte key.
add x0, sp, #0x10 puts the address of key (16 bytes above sp) in x0, and mov x1, #0x20 puts the length 32 (0x20) in x1. These are the two arguments of use_key, passed in registers x0 and x1.
bl 0 <use_key> is a call (branch with link). The R_AARCH64_CALL26 use_key line underneath is only the linker's note that the target address is filled in later.
ret returns.
What you are checking is what is missing: between the bl and the ret there is no store to the stack and no call to memset. The memset(key, 0, ...) line in the source produced no instructions at all, so the zeroing is gone. The barrier variant keeps it as two explicit stores of the zero register (xzr always reads as zero; stp stores a pair of registers, so each instruction writes 16 zero bytes), stp xzr, xzr, [sp, #32] and stp xzr, xzr, [sp, #48], after the use_key call and covering all 32 bytes of that function's key; with_explicit_bzero still contains bl explicit_bzero, and with_volatile_ptr an indirect blr x3. The same experiment built for x86-64 (--platform linux/amd64, same compiler and flags) shows the identical outcome: plain_memset has only the call to use_key, with_explicit_bzero calls explicit_bzero, with_volatile_ptr ends in call *%rax, and the barrier version writes zeros with pxor (which clears an SSE register, here %xmm0, to all zero bits) and two movaps stores (each writes the 16 bytes of that register to the stack, 2 x 16 = 32 bytes). This is a property of this compiler, flags and function, which is exactly why the check is to read the optimized disassembly rather than assume.
Why the volatile pointer and the empty asm stop the deletion
Both tricks work by making the compiler unable to prove that the wipe is unobservable.
- Volatile function pointer (
memset_v): the compiler must assume a volatile object can change at any time, so it has to load the pointer from memory and call whatever address it finds there. It can no longer know the callee is memset, so it cannot reason that the call only writes to a dead buffer and remove it. The call is kept (blr x3 on aarch64, call *%rax on x86-64: a call through a register).
- Empty
asm with a "memory" clobber (__asm__ __volatile__("" : : "r"(key) : "memory")): the quoted string is empty, so it emits no instruction, but it is declared to read the pointer key (the "r"(key) input) and to possibly read or write any memory (a clobber is a declaration that the asm may change the named thing). The GCC manual says volatile stops the compiler deleting the statement or hoisting it out of a loop, though it cautions that the compiler can still move even a volatile asm relative to other code. What pins the order here is the "memory" clobber together with the "r"(key) input: the asm is declared to read key and any memory, so the memset before it might be observed by it, and the zero stores have to stay, which is why the stores reappear in the listing.
Neither is a language guarantee; both are properties of how current GCC and Clang treat these constructs, which is why the table below ranks the standard functions first and why you still read the disassembly.
What to use, in order of preference
| Option | Guarantee | Availability |
|---|
memset_explicit | C23 standard function; the write is guaranteed to happen | needs a C23 library; check yours |
explicit_bzero | calls "are never optimized away by the compiler" (glibc man page) | glibc 2.25 and later; BSDs |
memset_s | C11 Annex K (an optional bounds-checking extension to the C standard that many libraries do not implement); write guaranteed | not provided by glibc; check __STDC_LIB_EXT1__ |
memset through a volatile function pointer, or memset followed by an empty asm with a "memory" clobber | works in practice, no standard guarantee | portable fallback; verify the disassembly |
The CERT C rule MSC06-C ("Beware of compiler optimizations", one rule in the CERT C Coding Standard, a widely used secure-coding guideline) describes this exact problem, names memset_explicit as the preferred C23 fix, and says to inspect the generated assembly of the optimized release build to confirm the memory is really cleared. Do that in CI for the key-handling functions.
Why even a correct wipe is not enough: where secrets live and how many copies
Zeroing the buffer you hold does not touch other copies. The first three items below are the common failures; the last item is hardening for long-lived keys. Each of these may keep the secret alive:
- Registers and spill slots. A key loaded into registers, or spilled to a stack slot (a temporary stack location where the compiler parks a register value when it runs out of registers), is not part of your buffer.
- Stack frames of callees and temporaries (an inlined copy, a
memcpy into a local) that are not wiped when the function returns.
- Heap copies:
realloc can move a buffer and leave the old block with the secret in it, unwiped. Never grow a secret buffer with realloc; allocate once at the final size.
- Swap, hibernation, core dumps and process forks.
mlock makes the kernel keep the pages resident so they are not written to swap, but the man page itself warns that laptop suspend saves RAM to disk regardless of locks. madvise(MADV_DONTDUMP) (Linux 3.4+) keeps a range out of core dumps, and MADV_WIPEONFORK (Linux 4.14+) gives a forked child zero-filled pages in that range. mlock counts against RLIMIT_MEMLOCK for unprivileged processes, so size the locked region small.
Decision rules I would commit to:
- Minimize copies: one canonical buffer per secret, passed by pointer, never returned by value and never logged or placed in a string that gets duplicated (
std::string growth, JSON building, error messages).
- Choose the storage by lifetime. A short-lived secret (a password during login) can live in a stack buffer that you wipe with
explicit_bzero or memset_explicit before every return path (single exit with a cleanup label); the stack is bounded and not subject to realloc. A long-lived key (a TLS or signing key held for the process lifetime) belongs in one dedicated allocation, ideally page-aligned, locked with mlock, excluded from dumps, and wiped at shutdown. If the platform offers it, a secret-memory facility or hardware (a hardware security module, or keeping the key in a separate process) beats any in-process scheme.
- Wipe on every exit path, including error returns, and wipe the intermediate values of the computation, not just the final key.
- Do not rely on wiping alone: assume a memory-disclosure bug (a buffer over-read) exposes whatever is live, so limit how long a secret exists in plaintext.
Trade-offs and pitfalls
mlock can fail when RLIMIT_MEMLOCK is low; decide whether that is fatal for your service and check the return value.
- Wiping costs a few stores and is irrelevant for performance, so there is no reason to skip it for speed.
- C++ objects: a destructor that wipes helps, but containers copy on growth. Use a fixed buffer owned by one object that is not copyable.
- Compilers can also keep a stale copy in a register or stack slot that only a barrier plus careful design can address; no language-level wipe is a complete defence, which is why the number of copies matters as much as the wipe.