Language-Level Memory Management (C/C++/Rust) Questions
How C and C++ expose and manage memory at the language level: pointers and pointer arithmetic, pointers to pointers and function pointers, arrays versus pointers and array decay, stack versus heap storage, manual allocation and freeing in C (malloc, realloc, free), new and delete versus malloc and free, placement new, RAII and owning smart pointers (unique_ptr, including custom deleters for C resources), shallow versus deep copies and move semantics in C++. Covers reasoning about who owns a buffer and designing ownership and lifetime contracts across function, library and plugin boundaries; dangling and uninitialized pointers, leaks, double frees, use-after-free, off-by-one errors and buffer overruns and how to prevent them; undefined behavior, strict aliasing and safe byte reinterpretation, endianness; struct layout, alignment and padding as the language defines them; and finding and diagnosing memory bugs with sanitizers, Valgrind-style tools, tracing allocators and heap-corruption triage, including allocator design and heap fragmentation at the language level. Boundary: garbage-collector behavior and tuning, embedded memory budgets, register access and packing structs to hardware layouts, OS virtual memory and paging, and lock-based concurrency are covered elsewhere.
What are the main techniques to prevent and detect buffer overflows in C beyond swapping in safer-looking library functions? Cover both coding practice and build-time or runtime defenses.
Sample Answer
Direct answer
Defense for buffer overflows has four layers: write code where the length always travels with the pointer and is checked before every copy; get the compiler to warn and to insert checks (-Wall -Wextra, -D_FORTIFY_SOURCE, -fstack-protector-strong); catch the bugs you did not see in review with AddressSanitizer plus fuzzing before release; and keep OS-level mitigations on as a last line that turns many exploits into crashes. (A fuzzer is a program that feeds a target huge numbers of generated or mutated inputs looking for crashes; fuzzing is running one.) No single layer is enough, and the safe-looking library function swap is the weakest of them.
Why "safer functions" alone fall short
strncpy does not terminate when the source fills the count (https://en.cppreference.com/w/c/string/byte/strncpy). Any bounded function still needs the correct bound, and the common bug is passing the wrong number (the pointer's size, the source length, or the buffer size without the terminator). The fix is structural: make the size impossible to forget.
Layer 1: coding practice
- Pass pointer and length together (a small
struct buf { uint8_t *p; size_t len; }), and checklenbefore copying, not after. - Use
size_tfor sizes and check arithmetic before it happens:if (n > cap - used) return ERR;rather thanused + n > cap, which can wrap. - Prefer functions that report truncation (
snprintfreturns the length it wanted), and check the result. - Allocate with the size computed from the same constant used to bound the copy, so the two cannot drift apart.
- Parse by validating the length field against what you actually received, then copy.
Layer 2: build-time defenses
-Wall -Wextrafor static diagnostics. They are not complete: in the example below, a plainstrcpyinto an 8-byte array fromargv[1]compiled with no warning.-fstack-protector-strongputs a canary (a guard value) between local arrays and the saved return address and checks it on function exit. Picture the stack frame asbuf[8], then the canary, then the saved return address (the place the function jumps back to when it ends): an overflow that runs pastbuftoward the return address must overwrite the canary first, and the check at exit sees the changed value and aborts instead of returning to an attacker-chosen address. GCC defines it as-fstack-protectorplus functions that have local arrays or reference local frame addresses (GCC manual: https://gcc.gnu.org/onlinedocs/gcc/Instrumentation-Options.html).-D_FORTIFY_SOURCE: glibc and the compiler add lightweight checks to string and memory functions when the destination size is known. The recommendation below uses level 2. Level 1 needs-O1or higher, level 2 adds more checks (some conforming programs can fail), and level 3 adds checks for buffers whose size is only known at run time (for example frommalloc(n)) and needs GCC 12 or later with glibc 2.33 or later (https://man7.org/linux/man-pages/man7/feature_test_macros.7.html).
Layer 3: test-time detection
-fsanitize=address instruments memory accesses to detect out-of-bounds and use-after-free bugs. Combine with -fsanitize=undefined, then run unit tests and a fuzzer so the sanitizer sees hostile inputs. This is the layer that finds the bug rather than merely containing it.
Worked example: one overflow, four builds
#include <stdio.h>
#include <string.h>
int main(int argc, char **argv)
{
char buf[8];
if (argc < 2) return 1;
strcpy(buf, argv[1]); /* no length check: overflows for inputs of 8 or more chars */
printf("copied: %s\n", buf);
return 0;
}
Run with a 32-character argument of A in a Linux arm64 container (GCC 14.4.0):
gcc -O0 -fno-stack-protector -> copied: AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (exit 0, silent corruption)
gcc -O0 -fstack-protector-strong -> *** stack smashing detected ***: terminated (Aborted)
gcc -O2 -D_FORTIFY_SOURCE=2 -fno-stack-protector -> *** buffer overflow detected ***: terminated (Aborted)
gcc -O1 -g -fsanitize=address -> ERROR: AddressSanitizer: stack-buffer-overflow ... WRITE of size 33
The unprotected build overflowed an 8-byte array with 33 bytes (the 32 A characters plus the terminating zero byte that strcpy also writes, which is where WRITE of size 33 in the sanitizer line comes from), printed success and exited 0: this is why overflows are dangerous, since nothing signalled the corruption. The other three builds stopped it, at different costs and with different information (the sanitizer names the access, the production flags just abort).
Layer 4: OS and platform mitigations
Two mitigations are normally on by default. A non-executable stack marks stack memory as data, so an attacker who injects machine code into an overflowed buffer cannot simply run it. Address-space randomization (ASLR) loads the program, libraries, stack and heap at different addresses on each run, so an attacker cannot hard-code the address of the code they want to jump to. They make exploiting a surviving overflow harder; they are not a fix and can be bypassed (for example, an information leak reveals the randomized addresses), so treat them as damage limiters.
Recommendation
For a new C codebase handling untrusted data: length-carrying buffer types and checked arithmetic in the code, -Wall -Wextra -fstack-protector-strong -D_FORTIFY_SOURCE=2 in every release build, and a CI job that runs tests and a fuzzer under -fsanitize=address,undefined. If it is a hot, high-exposure parser, consider a memory-safe language for that component. What would change this: on a tiny embedded target without room for canaries, lean harder on coding rules, static analysis and host-side sanitizer testing of the same source.
Pitfalls
- Hardening flags are not tests: a program that aborts on overflow has still got the bug.
- Sanitizers slow programs and use more memory, so they belong in test builds, not production.
- Heap overflows, off-by-one writes inside a struct, and overruns that stay within one allocation may not trip a stack canary at all.
Describe the lifecycle of memory allocated with malloc and released with free. What ownership rules would you document for a team to prevent leaks, double frees, and accidental sharing of mutable buffers?
Sample Answer
Direct answer. malloc(n) asks the allocator for a block of at least n bytes and returns its address, or NULL on failure. The bytes are uninitialized, the block is aligned for any object type (its address suits the strictest alignment requirement of any ordinary C type, so you can store an int, a double or a struct there), and it stays valid until you pass the same address once to free or realloc. After that the pointer is dangling. The rule set that prevents leaks, double frees and accidental sharing is: one named owner per allocation, ownership transfers stated in the API, frees paired with allocations in the same module, owner pointers nulled after free, realloc through a temporary, and borrowed pointers const unless mutation is the point.
Lifecycle
- Allocate.
p = malloc(count * sizeof *p). Check for NULL.calloc(count, size)zero-fills. Either way, rejectcount > SIZE_MAX / sizeof *pbefore multiplying, because an overflowed size silently allocates too little. Worked numbers for 4-byteinton a 64-bit system, wheresize_twraps at 2^64: withcount = 2^62 + 1 = 4611686018427387905, the productcount * 4is 2^64 + 4, which wraps to4.malloc(4)succeeds, and then a loop that writescountints runs far past the 4-byte block. The guard catches it:SIZE_MAX / 4is 4611686018427387903 andcountis larger, so the request is rejected. The program under 'Checking the wrap arithmetic' below prints these numbers. - Initialize.
mallocmemory has indeterminate contents (whatever bytes the allocator had lying around, not zeros). Reading before writing is UB. - Use. Stay inside
[p, p + count). Keep the size next to the pointer. - Resize (optional).
realloc(p, n)may move the block; the old pointer is then invalid. On failure it returns NULL and the old block is still valid and still yours. - Free exactly once.
free(p). Afterwards no use ofpor any copy of it.free(NULL)does nothing.
Checking the wrap arithmetic
#include <stdio.h>
#include <stdint.h>
#include <stdlib.h>
int main(void) {
size_t count = ((size_t)1 << 62) + 1;
size_t bytes = count * sizeof(int); /* wraps modulo 2^64 */
printf("count=%zu bytes after wrap=%zu\n", count, bytes);
printf("SIZE_MAX/4=%zu guard rejects it: %d\n",
(size_t)(SIZE_MAX / sizeof(int)), count > SIZE_MAX / sizeof(int));
return 0;
}
Built with gcc -Wall -Wextra on 64-bit Linux (GCC 14) it prints:
count=4611686018427387905 bytes after wrap=4
SIZE_MAX/4=4611686018427387903 guard rejects it: 1
What goes wrong (each run under AddressSanitizer)
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
int main(int argc, char **argv) {
int mode = argc > 1 ? atoi(argv[1]) : 0;
char *buf = malloc(32);
if (!buf) return 1;
strcpy(buf, "owned by main");
if (mode == 1) { /* leak: owner forgets */
buf = NULL;
} else if (mode == 2) { /* double free */
free(buf); free(buf);
} else if (mode == 3) { /* realloc overwrite-the-only-pointer bug */
buf = realloc(buf, (size_t)1 << 62); /* fails and returns NULL */
printf("realloc returned %p; the old block is now unreachable\n", (void *)buf);
fflush(stdout);
} else {
char *tmp = realloc(buf, 64); /* safe pattern: use a temporary */
if (!tmp) { free(buf); return 1; }
buf = tmp;
printf("grown buffer still says: %s\n", buf);
free(buf);
}
return 0;
}
Built with gcc -g -Wall -Wextra -fsanitize=address s7.c -o s7 (GCC warns about the deliberate double free):
./s7 0printsgrown buffer still says: owned by main(the safe path with a temporary)../s7 1(owner forgets):ERROR: LeakSanitizer: detected memory leakswithDirect leak of 32 byte(s) in 1 object(s) allocated from: ... main /w/s7.c:7../s7 2(double free):ERROR: AddressSanitizer: attempting double-free on 0x503000000040 in thread T0:plus the stack of the first free and of the allocation../s7 3withASAN_OPTIONS=allocator_may_return_null=1(otherwise ASan aborts on the absurd 2^62-byte request): it printsrealloc returned (nil); the old block is now unreachableand then a LeakSanitizer report of the same 32 bytes. That is the "p = realloc(p, n)" leak: the only pointer to the old block was overwritten by NULL.
Ownership rules to document (six)
- One owner. Every allocation has exactly one variable or object responsible for freeing it. Write it in the header comment: "returns an owned buffer; caller frees with
x_free". - Transfer is explicit. Name functions so direction is visible (
x_new/x_free,take_for "I now own it",borrow_or aconstparameter for "you may look"). A function that stores a pointer past its return must document who frees it. - Allocate and free in the same layer. Library allocates, library frees (
lib_free). Never free memory from a different allocator or a different runtime'smalloc. - Free paths are single. Use one cleanup label (
goto out;) per function so every error path frees what was allocated:
#include <stdio.h>
#include <stdlib.h>
/* fail_b simulates the second allocation failing, so the error path can be exercised. */
static int build(size_t n, int fail_b) {
int rc = -1;
char *a = NULL, *b = NULL;
a = malloc(n);
if (!a) goto out;
b = fail_b ? NULL : malloc(n);
if (!b) goto out; /* a is still freed below */
rc = 0;
out:
free(b); /* free(NULL) is a no-op */
free(a);
return rc;
}
int main(void) {
printf("success path: rc = %d\n", build(8, 0));
printf("second allocation fails: rc = %d\n", build(8, 1));
return 0;
}
Both pointers start as NULL, so jumping to out from any point frees exactly what exists. Built with gcc -g -Wall -Wextra -fsanitize=address,undefined and run in a gcc:14 container, it prints the two lines below with no sanitizer report; deleting the free(a); line makes LeakSanitizer report an 8-byte leak, so the check can fail.
success path: rc = 0
second allocation fails: rc = -1
- Null after free, and
reallocvia a temporary.p = NULLright afterfree(p)on the owning field;tmp = realloc(p, n); if (!tmp) { ...handle...; } else p = tmp;. - No silent mutable sharing. Pass
const T *to readers. If two components need to write the same buffer, either copy on handoff or make one the owner and give the other a documented window with a length. Do not stash a borrowed pointer in a struct that outlives the buffer.
Trade-offs and pitfalls
- Rule enforcement is by convention in C; back it with tests under AddressSanitizer and LeakSanitizer and a periodic review of every
mallocagainst itsfree. - On a small embedded target the answer changes: many teams avoid the general heap after startup and use fixed pools (a preallocated set of equal-sized slots handed out and returned), because fragmentation (free memory split into pieces too small to satisfy a request, even when the total is enough) and failure under load are hard to test. That is a project decision to state, not a default.
- Reference counting or an arena (one block freed in one call) are alternatives when many small objects share a lifetime: an arena trades per-object frees for one free at the end:
one malloc'd block: [ obj A ][ obj B ][ obj C ][ obj D ] ...unused...
^ next allocation goes here
free(block) once -> A, B, C and D are all released together
Write a small C function that reports whether the host is little-endian, and one that converts a 32-bit value to big-endian. What portability and strict-aliasing pitfalls should you avoid, and what bugs appear when code assumes host byte order matches the data format?
Sample Answer
Direct answer. Endianness is the order in which the bytes of a multi-byte number sit in memory: little-endian stores the least significant byte first (x86-64 and most ARM systems), big-endian stores it last (network byte order, the big-endian order that internet protocols use for numbers on the wire, and many file formats). To test the host, store a known value such as 1 and look at its first byte through a character type or memcpy. To convert a 32-bit value to big-endian, either swap bytes arithmetically or, better for parsing, build and take apart values with shifts at fixed byte positions, which is correct on every host. The pitfalls are casting a char * or packet buffer to uint32_t * (a strict-aliasing violation and a possible misaligned access, both undefined behaviour) and treating the packet's byte order as the host's byte order.
The two functions, plus the parsing helpers
#include <stdalign.h>
#include <stdint.h>
#include <stdio.h>
#include <string.h>
/* Runtime check: inspect the first byte of a known value through memcpy. */
static int host_is_little_endian(void) {
const uint16_t one = 1;
unsigned char first;
memcpy(&first, &one, 1);
return first == 1;
}
/* Reverse the byte order of a value. Pure arithmetic, so it behaves the same on every host. */
static uint32_t swap32(uint32_t v) {
return ((v & 0x000000FFu) << 24) | ((v & 0x0000FF00u) << 8) |
((v & 0x00FF0000u) >> 8) | ((v & 0xFF000000u) >> 24);
}
/* Host value -> value whose in-memory bytes are big-endian: swap on a little-endian host, identity on a big-endian one. */
static uint32_t host_to_be32(uint32_t v) {
return host_is_little_endian() ? swap32(v) : v;
}
/* Wire fields: byte positions are fixed by the format, so host order never matters. */
static uint32_t load_be32(const unsigned char *p) {
return ((uint32_t)p[0] << 24) | ((uint32_t)p[1] << 16) | ((uint32_t)p[2] << 8) | (uint32_t)p[3];
}
static void store_be32(unsigned char *p, uint32_t v) {
p[0] = (unsigned char)(v >> 24); p[1] = (unsigned char)(v >> 16);
p[2] = (unsigned char)(v >> 8); p[3] = (unsigned char)v;
}
static uint32_t load_le32(const unsigned char *p) {
return (uint32_t)p[0] | ((uint32_t)p[1] << 8) | ((uint32_t)p[2] << 16) | ((uint32_t)p[3] << 24);
}
/* The tempting wrong way: reinterpret packet bytes as a uint32_t. */
static uint32_t bad_load_be32(const unsigned char *p) {
return *(const uint32_t *)p;
}
int main(void) {
printf("host is little-endian: %d\n", host_is_little_endian());
uint32_t be = host_to_be32(0x11223344u);
unsigned char mem[4];
memcpy(mem, &be, 4);
printf("host_to_be32(0x11223344) = 0x%08x, bytes in memory: %02x %02x %02x %02x\n", be, mem[0], mem[1], mem[2], mem[3]);
alignas(4) unsigned char pkt[12] = {0xFF, 0, 0, 0, 0x10, 0, 0, 0, 0, 0, 0, 0x2A};
printf("load_be32(pkt+1) = %u (bytes 00 00 00 10, expect 16)\n", load_be32(pkt + 1));
printf("load_le32(pkt+1) = %u (same bytes read little-endian)\n", load_le32(pkt + 1));
printf("load_be32(pkt+8) = %u (bytes 00 00 00 2a, expect 42)\n", load_be32(pkt + 8));
printf("bad_load_be32(pkt+8) = %u (aligned, but host order)\n", bad_load_be32(pkt + 8));
unsigned char out[4];
store_be32(out, 0x11223344u);
printf("store_be32 bytes: %02x %02x %02x %02x\n", out[0], out[1], out[2], out[3]);
printf("bad_load_be32(pkt+1) = %u (misaligned read)\n", bad_load_be32(pkt + 1));
return 0;
}
Built and run in a Linux container (GCC 14.4, aarch64 little-endian host) with gcc -O2 -Wall -Wextra -fsanitize=address,undefined endian.c -o endian && ./endian:
endian.c:39:12: runtime error: load of misaligned address 0xffff90800041 for type 'const uint32_t', which requires 4 byte alignment
host is little-endian: 1
host_to_be32(0x11223344) = 0x44332211, bytes in memory: 11 22 33 44
load_be32(pkt+1) = 16 (bytes 00 00 00 10, expect 16)
load_le32(pkt+1) = 268435456 (same bytes read little-endian)
load_be32(pkt+8) = 42 (bytes 00 00 00 2a, expect 42)
bad_load_be32(pkt+8) = 704643072 (aligned, but host order)
store_be32 bytes: 11 22 33 44
bad_load_be32(pkt+1) = 268435456 (misaligned read)
(The first line is UBSan's message on stderr, so it appears before the buffered output; the 0xffff... address differs between runs. UBSan also prints a note: pointer points here line and a byte dump after that message, omitted here.) The same source compiled and run on a big-endian machine (an s390x Alpine container, s390x being a big-endian IBM mainframe CPU, run under QEMU, a CPU emulator, gcc -O2 -Wall -Wextra endian.c -o endians) printed:
host is little-endian: 0
host_to_be32(0x11223344) = 0x11223344, bytes in memory: 11 22 33 44
load_be32(pkt+1) = 16 (bytes 00 00 00 10, expect 16)
load_le32(pkt+1) = 268435456 (same bytes read little-endian)
load_be32(pkt+8) = 42 (bytes 00 00 00 2a, expect 42)
bad_load_be32(pkt+8) = 42 (aligned, but host order)
store_be32 bytes: 11 22 33 44
bad_load_be32(pkt+1) = 16 (misaligned read)
Reading the little-endian run: host_to_be32 swapped the bytes of 0x11223344, so the integer it returned prints as 0x44332211, but what matters is the memory behind it. A little-endian CPU stores the least significant byte first, so the integer 0x44332211 sits in memory as 11 22 33 44, which is the big-endian layout of the original number. That is why the printed integer looks "wrong" while the bytes are right. On the big-endian machine the value was already in that order, so nothing was swapped and it printed 0x11223344 with the same bytes. The load_be32, store_be32 and the in-memory bytes of host_to_be32 are identical on both machines, which is the point: they describe the format. bad_load_be32 gives 42 on the big-endian host and 704643072 on the little-endian one, so code using it passes on one machine and breaks on the other.
Reasoning through the pitfalls
- Why
host_to_be32must check the host. The byte swapswap32is pure arithmetic and always reverses the bytes. On a big-endian host the value is already in big-endian order, so the conversion must do nothing, and a function that always swaps is wrong there. In practice callhtonl(<arpa/inet.h>, POSIX; "host to network long": converts a 32-bit value from host order to big-endian network order), or on GCC/Clang__builtin_bswap32(a compiler built-in that reverses the four bytes) guarded by the compiler's byte-order macro, and let the library handle the host difference. - Strict aliasing. C lets the compiler assume that an object of one type is not accessed through a pointer to an unrelated type (the strict-aliasing rule; character types are the exception).
*(const uint32_t *)pon a byte buffer breaks that rule. Usememcpyinto auint32_t(compilers turn it into one load) or build the value with shifts; both are valid in C and C++. - Alignment. A
uint32_t *pointing at an odd address is misaligned: UBSan flagged thepkt+1read above, and some CPUs fault on it. Parsing with byte loads has no alignment requirement. - Shifting signed values. Shift unsigned operands, and cast each byte to
uint32_tbefore<<(auint8_tpromotes to a signedint: for the byte0x80,0x80 << 24is0x80000000, which does not fit in a signed 32-bitintwhose maximum is0x7FFFFFFF, and in C that overflow in a shift is undefined). - Compiler quality. The shift-and-or
load_be32compiles at-O2toldr w0, [x0]plusrev w0, w0(AArch64: load four bytes, then reverse their order) andmovl (%rdi), %eaxplusbswap %eax(x86-64: the same pair, load then byte-reverse), so the portable code costs no more than the cast.
Bugs that appear when host order is assumed
Parsing a network packet or a file header by overlaying a struct or casting works on the developer's little-endian laptop only if the format happens to be little-endian. Typical failures: a TCP port or a length field read byte-reversed (80 reads as 20480, 0x50 swapped to 0x5000); a magic number that matches on one build and never on another; a loop that trusts a byte-reversed length and over-reads (a security bug, since an attacker chooses the bytes); a file written on one CPU and unreadable on another. The defence is the rule in the code: say the byte order in the format specification, read and write each field with explicit byte positions, and test on a big-endian target (an emulator is enough).
Trade-offs
Use the portable shift form by default. Reach for memcpy plus a swap when profiling shows the byte-wise form matters on a compiler that does not recognize the pattern, and for bulk data convert in a loop the compiler can vectorize (compile into instructions that handle several values at once). Do not gate behaviour on a runtime endianness probe in hot paths: it is constant for a build, so use compile-time macros there.
You are designing a C event-loop API that stores a callback and a user-data pointer for later invocation. The callback may run long after registration, and callers can unregister at any time. What rules or safeguards would you put in place so the callback does not dereference freed memory?
Sample Answer
Direct answer
Give the callback registration a lifetime that is longer than any pending call, and make the owner of user_data explicit. Reference counting means keeping a small integer inside the object that says how many holders still use it; each holder increments it when it starts using the object and decrements it when done, and whoever brings it to zero frees the object. Practically: reference-count the registration object, let every queued or running invocation hold its own reference, make unregister only mark the entry inactive and drop the registration's reference (never free directly), check the active flag immediately before each call, and free the user data through a release callback exactly once when the count reaches zero. That removes the three failure modes: a callback that fires after unregister, a callback that unregisters itself or a neighbour while running, and a shared object freed while another callback still uses it.
Terms: an event loop is the central loop that waits for events (a timer, a socket becoming readable) and calls the callback registered for each; "re-enter" means a callback calling back into the loop (for example to unregister) while the loop is in the middle of dispatching; a deferred operation is one that records the request now and does the work later. The core mechanism is the reference count plus the active flag; the handle table and the deferred free queue under "Alternatives" are other ways to reach the same safety.
Rules to put in the API contract
- Registration returns a handle (the watcher). The caller never frees
user_datadirectly after registering; it hands ownership to the loop with arelease(user)function, or documents a different rule. unregisteris deferred, and it is called once per registration. It setsactive = 0and drops the registration reference. No further callbacks run for that watcher, including events already queued. The registration reference was the caller's only claim on the watcher, so when nothing is queued the watcher is freed right there: afterunregisterthe caller must drop its handle and must not callunregister(orpost) on it again. A second call would read freed memory (run exactly that way under AddressSanitizer, it reportsheap-use-after-free,READ of size 4). If you want a second call to be a harmless no-op, give the caller its own reference to the watcher, or use generation-numbered handles (below).- Every pending invocation owns a reference.
postincrements the count; the loop decrements it after the call or after skipping the call. This makes "the callback may run long after registration" safe, since the watcher outlives the queue entry. - Re-check
activeat call time, not at post time. - The running callback is protected too: the queue entry's reference is released only after the call returns, so a callback may unregister itself and still read its own context.
- One thread rule. The loop above is single-threaded (plain
unsignedcounts). Ifunregistercan be called from another thread, the count must be atomic and the active flag needs synchronization; say which in the API docs. - Shared objects (one callback frees what another uses): do not free by hand inside callbacks. Give shared state its own reference count (or have one owner and pass non-owning handles that are validated through the owner), so the last user frees it.
Worked example
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
typedef void (*cb_fn)(void *user, int event);
typedef void (*release_fn)(void *user);
typedef struct watcher {
cb_fn fn;
void *user;
release_fn release; /* frees user; called exactly once, when the last reference drops */
unsigned refs; /* 1 for the loop's registration + 1 per queued event (that entry's reference also covers the call while it runs) */
int active; /* cleared by unregister; checked before every call */
} watcher;
#define MAX_QUEUED 16
typedef struct { watcher *w; int event; } queued;
typedef struct { queued q[MAX_QUEUED]; size_t n; } loop;
static void w_unref(watcher *w)
{
if (--w->refs == 0) {
if (w->release) w->release(w->user);
free(w);
}
}
watcher *loop_register(cb_fn fn, void *user, release_fn release)
{
if (fn == NULL) return NULL; /* reject a NULL callback at registration, not at call time */
watcher *w = calloc(1, sizeof *w);
if (w == NULL) return NULL;
w->fn = fn; w->user = user; w->release = release; w->refs = 1; w->active = 1;
return w;
}
void loop_unregister(watcher *w)
{
if (w->active) {
w->active = 0; /* no further callbacks, even for events already queued */
w_unref(w); /* drop the registration reference; queued events keep w alive */
}
}
int loop_post(loop *lp, watcher *w, int event)
{
if (!w->active || lp->n == MAX_QUEUED) return -1;
w->refs++; /* the queue entry owns a reference until it is processed */
lp->q[lp->n++] = (queued){ w, event };
return 0;
}
void loop_run(loop *lp)
{
for (size_t i = 0; i < lp->n; i++) { /* callbacks may post more events: n can grow */
queued e = lp->q[i];
if (e.w->active)
e.w->fn(e.w->user, e.event);
w_unref(e.w); /* release the queue's reference AFTER the call */
}
lp->n = 0;
}
/* ---- demo ---- */
typedef struct { char name[16]; watcher *self; watcher *other; } session;
static void session_release(void *u) { session *s = u; printf(" free session %s\n", s->name); free(s); }
static session *new_session(const char *name)
{
session *s = calloc(1, sizeof *s);
strncpy(s->name, name, sizeof s->name - 1);
return s;
}
static void on_a(void *u, int ev)
{
session *s = u;
printf("A(%s) got event %d; unregistering B, then itself\n", s->name, ev);
loop_unregister(s->other); /* B's queued event still holds a reference, so B stays valid until that entry is skipped */
loop_unregister(s->self); /* A drops its own registration; the running call still holds a reference */
printf(" A still reads its own session after unregistering: %s\n", s->name);
}
static void on_b(void *u, int ev) { session *s = u; printf("B(%s) got event %d (must not print)\n", s->name, ev); }
int main(void)
{
loop lp = { .n = 0 };
session *sa = new_session("alpha"), *sb = new_session("beta");
watcher *wa = loop_register(on_a, sa, session_release);
watcher *wb = loop_register(on_b, sb, session_release);
sa->self = wa; sa->other = wb;
loop_post(&lp, wa, 1); /* A runs first and unregisters B */
loop_post(&lp, wb, 2); /* B is already queued when A unregisters it: its call must be skipped */
loop_run(&lp);
printf("done\n");
return 0;
}
Reading the code: loop_register creates a watcher with refs = 1 (the registration's own reference) and active = 1. loop_post adds a queue entry and bumps refs. loop_run takes each entry, calls the callback only if active is still set, and then drops the entry's reference. loop_unregister clears active and drops the registration reference. w_unref is the only place that frees, and it calls release first so the user data goes with it.
Compiled with gcc -O1 -g -Wall -Wextra -fsanitize=address,undefined loop.c (GCC 14.4.0, Linux arm64) and run, the program prints:
A(alpha) got event 1; unregistering B, then itself
A still reads its own session after unregistering: alpha
free session alpha
free session beta
done
Walk through the counts. Each watcher starts at 1 (registration); post raises each to 2. When A runs, it unregisters B (B: 2 to 1, inactive) and itself (A: 2 to 1). Reading s->name afterwards is still safe because the queue entry holds A's second reference. After the call returns, the loop drops it (A: 1 to 0) and the release function frees alpha. B's queued event is then skipped because B is inactive, and the loop drops its reference (B: 1 to 0), freeing beta. AddressSanitizer reported no errors and no leaks. ("Inactive means never called" is part of the contract: if the application unregistered A before the loop ran, A's callback would be skipped and would never get to unregister B, so cleanup must not depend on a callback running. LeakSanitizer checks that cleanup path too.)
Why not just NULL the pointer?
Setting a field to NULL on unregister only helps callers that look at that field. A copy of the pointer already in a queue still points at the freed block. Reference counting (or a handle table with generation numbers) is what makes stale references detectable or harmless. A handle table is an array of slots that callers refer to by index; a generation number is a counter stored in each slot and in the handle, bumped whenever the slot is reused, so a stale handle with an old generation is recognized as dead.
Alternatives and when to choose them
- Handle indirection: callers hold an integer ID and the loop looks it up at call time; a stale ID fails the lookup instead of dereferencing freed memory. Good when handles cross an API boundary or a thread.
- Deferred free queue: unregister marks the entry, and the loop frees marked entries only between dispatch rounds. Cheaper than counting, but you must guarantee no pending invocation can outlive a round.
- Choose reference counting when invocations can be queued across rounds, as here.
Trade-offs and pitfalls
- Reference cycles leak (a callback's context holding the watcher that holds the context): break them by unregistering explicitly.
- A release function that itself calls back into the loop can re-enter; forbid it in the contract.
- Do not decrement before the call returns; releasing early reproduces the bug.
- The caller's handle is dead once
unregisterreturns: callingunregistertwice on the same pointer is a use-after-free in this design (see rule 2).
Legacy C code reads a float's bits by casting a float pointer to a uint32_t pointer. Why is that undefined behavior, and what are two correct ways to reinterpret the value in both directions?
Sample Answer
Direct answer. Casting a float* to a uint32_t* and dereferencing it violates C's strict aliasing rule: an object of one type accessed through a pointer to an unrelated type is undefined behavior, and an optimizing compiler is allowed to assume it never happens, which can produce a wrong answer, not just a theoretical violation. The two correct fixes are memcpy between the two types, and (in C specifically) a union with both members, both of which reinterpret the bits without the compiler ever assuming non-aliasing across the boundary.
Why it is undefined, precisely. The C standard restricts which pointer type may be used to access a given object's stored value (the "effective type" rule: every object has one type it was last written as, and only an lvalue, an expression that names a specific piece of storage, of a compatible type may read or write it); accessing a float object through a uint32_t lvalue is not one of the permitted exceptions (which cover things like unsigned char*). Because this is undefined rather than merely "implementation defined," the compiler's optimizer is entitled to assume a float* and a uint32_t* never alias (two pointers alias when they point at the same memory) the same memory, and to reorder, cache, or eliminate loads/stores on that assumption. This is exactly how a miscompile happens: the bug is not in some obscure corner case, it is in the normal operation of optimizations like load/store reordering and redundant-load elimination that strict-aliasing analysis enables.
A bit-level picture first. "Reinterpreting bits" means treating the same 4 bytes once as an IEEE-754 float (the standard binary format for floating-point numbers) and once as a plain 32-bit integer, with no conversion, just a different reading of identical bits. The float 1.0f is stored as the bit pattern 0x3f800000; that is not a coincidence to memorize, it is IEEE-754's encoding (1 sign bit, 8 exponent bits, 23 mantissa bits) producing that exact pattern for the value 1.0. float_to_bits(1.0f) below returns exactly that pattern, confirming the two views are of the same bytes.
Worked miscompile, built and run at two optimization levels. A function writes through an int* alias and a float* alias to the same object, then reads back through the int*:
#include <stdio.h>
#include <stdint.h>
#include <string.h>
__attribute__((noinline))
int bad_alias(int *i, float *f) {
*i = 1;
*f = 2.0f; /* the optimizer assumes this cannot touch the int object *i points to */
return *i; /* so it may return the earlier value, 1, without reloading */
}
static uint32_t float_to_bits(float x) { uint32_t u; memcpy(&u, &x, sizeof u); return u; }
static float bits_to_float(uint32_t u) { float x; memcpy(&x, &u, sizeof x); return x; }
static uint32_t float_to_bits_union(float x) { union { float f; uint32_t u; } v = { .f = x }; return v.u; }
int main(void) {
union { int i; float f; } u; /* both pointers really alias these 4 bytes */
int r = bad_alias(&u.i, &u.f);
printf("bad_alias returned 0x%08x, the object holds 0x%08x\n", (unsigned)r, float_to_bits(u.f));
printf("%08x %g %08x\n", float_to_bits(1.0f), bits_to_float(0x40490fdb), float_to_bits_union(1.0f));
float x = 1.0f;
unsigned legacy = *(unsigned *)&x; /* the raw cast from the question */
printf("%08x\n", legacy);
return 0;
}
Built with gcc -Wall -Wextra (GCC 14.4, aarch64 Linux container), this prints at -O0:
bad_alias returned 0x40000000, the object holds 0x40000000
3f800000 3.14159 3f800000
3f800000
and at -O2 the first line becomes bad_alias returned 0x00000001, the object holds 0x40000000 (the other lines are unchanged), together with a compile-time warning: dereferencing type-punned pointer will break strict-aliasing rules [-Wstrict-aliasing] on the *(unsigned *)&x line; adding -fno-strict-aliasing at -O2 restores 0x40000000.
Called on a union { int i; float f; } (so the two pointers really do alias the same four bytes, as they would through a raw type-punning cast): at -O0, bad_alias returns 0x40000000 (the bit pattern of 2.0f, the value that was actually last written), because with optimizations off the compiler reloads *i rather than reusing a cached value. At -O2, the exact same source and the exact same object returns 0x00000001, the stale value from before the float write, because the optimizer's strict-aliasing analysis lets it skip reloading *i after a write through a pointer it has proven (under the aliasing rule) cannot touch the same storage. Step by step: the compiler sees *i = 1 (writes 0x00000001 into the shared bytes), keeps that value 1 in a register as a cached copy of *i, sees *f = 2.0f but, because float* and int* are assumed not to alias, does not invalidate its cached *i, and then return *i hands back the cached register value 1 instead of re-reading memory, where the real bytes now hold 2.0f's pattern. The object genuinely holds 0x40000000 in both builds (confirmed by reading it back through a memcpy-based accessor), so the -O2 build's return value is simply wrong: a silent miscompile, not a crash, and reproducible deterministically at that optimization level with that compiler. Compiling with -fno-strict-aliasing restores the -O0 answer at -O2, confirming the aliasing assumption is exactly what changed the result; at -O2, -Wstrict-aliasing=2 (GCC 14) flags a raw type-punning cast like *(unsigned *)&x at compile time with "dereferencing type-punned pointer will break strict-aliasing rules"; the warning only works while -fstrict-aliasing is active (on by default from -O2), so at -O0 the same cast compiles with no warning at all and a green -O0 build proves nothing.
Fix 1: memcpy. memcpy(&dst, &src, sizeof dst) between the two types copies the bytes through unsigned char-equivalent access, which the standard explicitly permits regardless of the source and destination's declared types, and which optimizing compilers recognize and compile down to a plain register move or load/store when the size is a compile-time constant (no real function-call overhead in practice):
static uint32_t float_to_bits(float x) { uint32_t u; memcpy(&u, &x, sizeof u); return u; }
static float bits_to_float(uint32_t u) { float x; memcpy(&x, &u, sizeof x); return x; }
Running float_to_bits(1.0f) returns 0x3f800000 (the correct IEEE-754 bit pattern for 1.0, matching the bit picture above), and bits_to_float(0x40490fdb) returns 3.14159 (the correct round-trip of pi's bit pattern), matching bit-for-bit between -O0 and -O2 builds, because memcpy never exposes the aliasing hazard to the optimizer in the first place.
Fix 2: union (C only). C explicitly permits reading a union member other than the one last written ("type punning": reading the stored bytes through a different member's type than the one used to write them), as long as the access goes through the union type itself, not through a cast pointer: union { float f; uint32_t u; } v = { .f = x }; return v.u; is well-defined in C and gives the identical bit pattern as the memcpy version. This is a C-specific guarantee; in ISO C++ the same union access is formally undefined behavior (GCC supports it as a non-standard extension there), so the memcpy form, or std::bit_cast (C++20, a standard library function that performs exactly this bit reinterpretation; requires both types be the same size and trivially copyable, meaning the type can be copied by copying its raw bytes with no special constructor logic), is the portable choice in C++.
Trade-offs and pitfalls. Do not "fix" this by disabling -fstrict-aliasing project-wide: it silences this specific miscompile but gives up a real optimization across the entire codebase and hides every other latent aliasing bug instead of fixing them, which is a maintenance trap for the next person who re-enables it. Prefer memcpy over the union when you need C++ portability or simply want one idiom that works in both languages; prefer the union inside pure C code where its clarity at the call site is valuable and the standard guarantee is unambiguous. Either fix costs nothing at a reasonable optimization level once the compiler recognizes the pattern, so there is no real performance argument for keeping the raw cast.
Unlock Full Question Bank
Get access to all 32 Language-Level Memory Management (C/C++/Rust) interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.