Language-Level Memory Management (C/C++/Rust) Questions
How C and C++ expose and manage memory at the language level: pointers and pointer arithmetic, pointers to pointers and function pointers, arrays versus pointers and array decay, stack versus heap storage, manual allocation and freeing in C (malloc, realloc, free), new and delete versus malloc and free, placement new, RAII and owning smart pointers (unique_ptr, including custom deleters for C resources), shallow versus deep copies and move semantics in C++. Covers reasoning about who owns a buffer and designing ownership and lifetime contracts across function, library and plugin boundaries; dangling and uninitialized pointers, leaks, double frees, use-after-free, off-by-one errors and buffer overruns and how to prevent them; undefined behavior, strict aliasing and safe byte reinterpretation, endianness; struct layout, alignment and padding as the language defines them; and finding and diagnosing memory bugs with sanitizers, Valgrind-style tools, tracing allocators and heap-corruption triage, including allocator design and heap fragmentation at the language level. Boundary: garbage-collector behavior and tuning, embedded memory budgets, register access and packing structs to hardware layouts, OS virtual memory and paging, and lock-based concurrency are covered elsewhere.
Write or describe a C function that safely copies up to n bytes from an untrusted input buffer into a destination buffer, guarantees null termination when appropriate, and returns an error if truncation occurs. How would you test edge cases such as zero-length buffers and exact fits?
Sample Answer
Direct answer. The safe shape is a function that takes the destination capacity and the source length as separate numbers, reserves one byte of the capacity for the terminating '\0' (NUL, the zero byte that ends a C string), copies at most the remaining room with memcpy, always writes the terminator, and returns a status that distinguishes "everything fit" from "I cut it short" and from "you gave me bad arguments". The destination is only ever written after the arguments are checked, and the length arithmetic is done so it cannot underflow. Then I test the boundaries (capacity 0, 1, exact fit, one too long) with guard bytes around the destination and run the tests under AddressSanitizer (ASan, a compiler feature that traps out-of-bounds accesses).
Design decisions
- Why not
strcpyorstrncpy.strcpyhas no limit, so a long input overruns the destination.strncpyis bounded, but it does not terminate the result when the source is as long as the limit, and it pads the rest with zeros (wasted work on a large buffer), so callers forget the terminator. Neither reports truncation. - Explicit
src_len. Untrusted input may have no terminator at all, or an embedded NUL, so the function takes a byte count and does not callstrlenon the source. If the source is meant to be a C string with an upper bound, compute the count withstrnlen(src, max)first (POSIX; likestrlenbut it stops after at mostmaxbytes, so a buffer with no terminator is never read past its end). - Reserve the terminator first.
room = dst_size - 1is safe only becausedst_size == 0was rejected before it: with an unsigned type,0 - 1wraps to the largestsize_t, which would turn a bounds check into a huge copy. - Truncation is a result, not a silent success. Silently chopping a path, a hostname or a command line changes its meaning (a truncated path can name a different file), so the caller must be able to refuse. On truncation the destination still holds a valid, terminated prefix, so nothing downstream reads garbage if the caller ignores the status.
- Zero-capacity destination. There is no room even for the terminator, so the function writes nothing and returns an argument error. It cannot return success, because it could not produce a string.
- Empty copy.
src_len == 0is valid, including a NULL source, and yields an empty string;memcpywith a NULL pointer is undefined even for length 0, so the code skips it whenn == 0.
Implementation with tests
#include <assert.h>
#include <stddef.h>
#include <stdio.h>
#include <string.h>
enum copy_status { COPY_OK = 0, COPY_TRUNCATED = 1, COPY_BAD_ARGS = -1 };
/* Copy up to src_len bytes of src into dst (capacity dst_size) as a C string.
src is untrusted bytes: it may lack a terminator and may contain embedded NULs.
On COPY_OK and COPY_TRUNCATED, dst is NUL-terminated and holds a prefix of src.
Returns COPY_BAD_ARGS (and touches nothing) if dst is NULL, dst_size is 0, or src is NULL with src_len > 0. */
enum copy_status bounded_copy(char *dst, size_t dst_size, const char *src, size_t src_len) {
if (dst == NULL || dst_size == 0 || (src == NULL && src_len > 0))
return COPY_BAD_ARGS;
size_t room = dst_size - 1; /* reserve one byte for the terminator */
size_t n = src_len < room ? src_len : room;
if (n > 0) memcpy(dst, src, n);
dst[n] = '\0';
return n < src_len ? COPY_TRUNCATED : COPY_OK;
}
static void check(const char *name, int cond) { printf("%-34s %s\n", name, cond ? "pass" : "FAIL"); assert(cond); }
int main(void) {
char guard_buf[16]; /* canary layout: guard | dst | guard */
memset(guard_buf, 'G', sizeof guard_buf);
char *dst = guard_buf + 4; /* 8 usable bytes in the middle */
#define GUARDS_INTACT() (memcmp(guard_buf, "GGGG", 4) == 0 && memcmp(guard_buf + 12, "GGGG", 4) == 0)
memset(dst, 'x', 8);
check("dst_size==0 -> BAD_ARGS, untouched", bounded_copy(dst, 0, "abc", 3) == COPY_BAD_ARGS && dst[0] == 'x');
check("dst==NULL -> BAD_ARGS", bounded_copy(NULL, 8, "abc", 3) == COPY_BAD_ARGS);
check("src==NULL,len>0 -> BAD_ARGS", bounded_copy(dst, 8, NULL, 3) == COPY_BAD_ARGS);
check("src==NULL,len==0 -> OK, empty", bounded_copy(dst, 8, NULL, 0) == COPY_OK && dst[0] == '\0');
check("zero-length src -> OK, empty", bounded_copy(dst, 8, "abc", 0) == COPY_OK && strlen(dst) == 0);
check("dst_size==1 any src -> TRUNC, empty", bounded_copy(dst, 1, "abc", 3) == COPY_TRUNCATED && dst[0] == '\0');
check("dst_size==1, src_len==0 -> OK", bounded_copy(dst, 1, "", 0) == COPY_OK && dst[0] == '\0');
check("short copy -> OK", bounded_copy(dst, 8, "abc", 3) == COPY_OK && strcmp(dst, "abc") == 0);
check("exact fit 7 chars+NUL -> OK", bounded_copy(dst, 8, "1234567", 7) == COPY_OK && strcmp(dst, "1234567") == 0);
check("one too long (8 chars) -> TRUNC", bounded_copy(dst, 8, "12345678", 8) == COPY_TRUNCATED && strcmp(dst, "1234567") == 0);
check("much too long -> TRUNC, prefix", bounded_copy(dst, 8, "ABCDEFGHIJKLMNOP", 16) == COPY_TRUNCATED && strcmp(dst, "ABCDEFG") == 0);
{
const char raw[5] = {'a', 'b', 'c', 'd', 'e'}; /* no terminator at all */
check("unterminated src -> OK, terminated", bounded_copy(dst, 8, raw, sizeof raw) == COPY_OK && strcmp(dst, "abcde") == 0);
}
{
const char emb[4] = {'a', '\0', 'b', 'c'}; /* embedded NUL */
check("embedded NUL copied as bytes", bounded_copy(dst, 8, emb, 4) == COPY_OK && memcmp(dst, "a\0bc\0", 5) == 0);
}
check("guard bytes around dst intact", GUARDS_INTACT());
puts("all edge cases passed");
return 0;
}
Built and run in a Linux container (GCC 14.4, aarch64) with gcc -O2 -Wall -Wextra -fsanitize=address,undefined bounded_copy.c -o bounded_copy && ./bounded_copy:
dst_size==0 -> BAD_ARGS, untouched pass
dst==NULL -> BAD_ARGS pass
src==NULL,len>0 -> BAD_ARGS pass
src==NULL,len==0 -> OK, empty pass
zero-length src -> OK, empty pass
dst_size==1 any src -> TRUNC, empty pass
dst_size==1, src_len==0 -> OK pass
short copy -> OK pass
exact fit 7 chars+NUL -> OK pass
one too long (8 chars) -> TRUNC pass
much too long -> TRUNC, prefix pass
unterminated src -> OK, terminated pass
embedded NUL copied as bytes pass
guard bytes around dst intact pass
all edge cases passed
How the edge cases are chosen
The interesting inputs sit on the boundaries of one formula, room = dst_size - 1:
| Case | Why it matters | Expected |
|---|---|---|
dst_size == 0 | no room for a terminator; dst_size - 1 would wrap | argument error, destination untouched |
dst_size == 1 | room for the terminator only | empty string; truncated if src_len > 0, OK if 0 |
src_len == dst_size - 1 (exact fit: 7 bytes into 8) | the largest copy that fits | OK, terminated |
src_len == dst_size (8 bytes into 8) | one byte too many | truncated, 7 bytes kept |
| unterminated or embedded-NUL source | the function must not rely on strlen of the source | bytes copied by count |
The destination sits in the middle of a 16-byte array with 'G' guard bytes on each side, so an off-by-one write is caught by the final guard check even when ASan cannot see it (writes inside one array are invisible to ASan, which is why the guards are there). The tests can fail: changing room = dst_size - 1 to room = dst_size and rebuilding makes the run abort on an assertion failure instead of printing all edge cases passed.
Trade-offs and pitfalls
strlcpy(a BSD function that takes the destination size, always terminates the result and reports truncation; glibc added it in 2.38 per its man page) has similar semantics and returns the length of the string it tried to create, but it takes a NUL-terminated source and so reads until the terminator, which an unterminated untrusted buffer does not have.bounded_copytakes a length instead.- Returning a status is only useful if callers check it: mark the function
__attribute__((warn_unused_result))(a GCC/Clang attribute that makes the compiler warn whenever a caller throws the return value away) so ignoring it warns. - Terminated output does not make the content safe. A copied string with an embedded NUL looks shorter to
strlenthan the bytes copied, so decide whether embedded NULs are an error for your protocol and reject them if so. - Complexity is O(n) in the number of bytes copied, constant extra space. For fuzzing, feed random lengths and capacities to the function under ASan and UBSan (UndefinedBehaviorSanitizer) and assert the invariants: output terminated, prefix of the source, guards intact.
Write a small C function that reports whether the host is little-endian, and one that converts a 32-bit value to big-endian. What portability and strict-aliasing pitfalls should you avoid, and what bugs appear when code assumes host byte order matches the data format?
Sample Answer
Direct answer. Endianness is the order in which the bytes of a multi-byte number sit in memory: little-endian stores the least significant byte first (x86-64 and most ARM systems), big-endian stores it last (network byte order, the big-endian order that internet protocols use for numbers on the wire, and many file formats). To test the host, store a known value such as 1 and look at its first byte through a character type or memcpy. To convert a 32-bit value to big-endian, either swap bytes arithmetically or, better for parsing, build and take apart values with shifts at fixed byte positions, which is correct on every host. The pitfalls are casting a char * or packet buffer to uint32_t * (a strict-aliasing violation and a possible misaligned access, both undefined behaviour) and treating the packet's byte order as the host's byte order.
The two functions, plus the parsing helpers
#include <stdalign.h>
#include <stdint.h>
#include <stdio.h>
#include <string.h>
/* Runtime check: inspect the first byte of a known value through memcpy. */
static int host_is_little_endian(void) {
const uint16_t one = 1;
unsigned char first;
memcpy(&first, &one, 1);
return first == 1;
}
/* Reverse the byte order of a value. Pure arithmetic, so it behaves the same on every host. */
static uint32_t swap32(uint32_t v) {
return ((v & 0x000000FFu) << 24) | ((v & 0x0000FF00u) << 8) |
((v & 0x00FF0000u) >> 8) | ((v & 0xFF000000u) >> 24);
}
/* Host value -> value whose in-memory bytes are big-endian: swap on a little-endian host, identity on a big-endian one. */
static uint32_t host_to_be32(uint32_t v) {
return host_is_little_endian() ? swap32(v) : v;
}
/* Wire fields: byte positions are fixed by the format, so host order never matters. */
static uint32_t load_be32(const unsigned char *p) {
return ((uint32_t)p[0] << 24) | ((uint32_t)p[1] << 16) | ((uint32_t)p[2] << 8) | (uint32_t)p[3];
}
static void store_be32(unsigned char *p, uint32_t v) {
p[0] = (unsigned char)(v >> 24); p[1] = (unsigned char)(v >> 16);
p[2] = (unsigned char)(v >> 8); p[3] = (unsigned char)v;
}
static uint32_t load_le32(const unsigned char *p) {
return (uint32_t)p[0] | ((uint32_t)p[1] << 8) | ((uint32_t)p[2] << 16) | ((uint32_t)p[3] << 24);
}
/* The tempting wrong way: reinterpret packet bytes as a uint32_t. */
static uint32_t bad_load_be32(const unsigned char *p) {
return *(const uint32_t *)p;
}
int main(void) {
printf("host is little-endian: %d\n", host_is_little_endian());
uint32_t be = host_to_be32(0x11223344u);
unsigned char mem[4];
memcpy(mem, &be, 4);
printf("host_to_be32(0x11223344) = 0x%08x, bytes in memory: %02x %02x %02x %02x\n", be, mem[0], mem[1], mem[2], mem[3]);
alignas(4) unsigned char pkt[12] = {0xFF, 0, 0, 0, 0x10, 0, 0, 0, 0, 0, 0, 0x2A};
printf("load_be32(pkt+1) = %u (bytes 00 00 00 10, expect 16)\n", load_be32(pkt + 1));
printf("load_le32(pkt+1) = %u (same bytes read little-endian)\n", load_le32(pkt + 1));
printf("load_be32(pkt+8) = %u (bytes 00 00 00 2a, expect 42)\n", load_be32(pkt + 8));
printf("bad_load_be32(pkt+8) = %u (aligned, but host order)\n", bad_load_be32(pkt + 8));
unsigned char out[4];
store_be32(out, 0x11223344u);
printf("store_be32 bytes: %02x %02x %02x %02x\n", out[0], out[1], out[2], out[3]);
printf("bad_load_be32(pkt+1) = %u (misaligned read)\n", bad_load_be32(pkt + 1));
return 0;
}
Built and run in a Linux container (GCC 14.4, aarch64 little-endian host) with gcc -O2 -Wall -Wextra -fsanitize=address,undefined endian.c -o endian && ./endian:
endian.c:39:12: runtime error: load of misaligned address 0xffff90800041 for type 'const uint32_t', which requires 4 byte alignment
host is little-endian: 1
host_to_be32(0x11223344) = 0x44332211, bytes in memory: 11 22 33 44
load_be32(pkt+1) = 16 (bytes 00 00 00 10, expect 16)
load_le32(pkt+1) = 268435456 (same bytes read little-endian)
load_be32(pkt+8) = 42 (bytes 00 00 00 2a, expect 42)
bad_load_be32(pkt+8) = 704643072 (aligned, but host order)
store_be32 bytes: 11 22 33 44
bad_load_be32(pkt+1) = 268435456 (misaligned read)
(The first line is UBSan's message on stderr, so it appears before the buffered output; the 0xffff... address differs between runs. UBSan also prints a note: pointer points here line and a byte dump after that message, omitted here.) The same source compiled and run on a big-endian machine (an s390x Alpine container, s390x being a big-endian IBM mainframe CPU, run under QEMU, a CPU emulator, gcc -O2 -Wall -Wextra endian.c -o endians) printed:
host is little-endian: 0
host_to_be32(0x11223344) = 0x11223344, bytes in memory: 11 22 33 44
load_be32(pkt+1) = 16 (bytes 00 00 00 10, expect 16)
load_le32(pkt+1) = 268435456 (same bytes read little-endian)
load_be32(pkt+8) = 42 (bytes 00 00 00 2a, expect 42)
bad_load_be32(pkt+8) = 42 (aligned, but host order)
store_be32 bytes: 11 22 33 44
bad_load_be32(pkt+1) = 16 (misaligned read)
Reading the little-endian run: host_to_be32 swapped the bytes of 0x11223344, so the integer it returned prints as 0x44332211, but what matters is the memory behind it. A little-endian CPU stores the least significant byte first, so the integer 0x44332211 sits in memory as 11 22 33 44, which is the big-endian layout of the original number. That is why the printed integer looks "wrong" while the bytes are right. On the big-endian machine the value was already in that order, so nothing was swapped and it printed 0x11223344 with the same bytes. The load_be32, store_be32 and the in-memory bytes of host_to_be32 are identical on both machines, which is the point: they describe the format. bad_load_be32 gives 42 on the big-endian host and 704643072 on the little-endian one, so code using it passes on one machine and breaks on the other.
Reasoning through the pitfalls
- Why
host_to_be32must check the host. The byte swapswap32is pure arithmetic and always reverses the bytes. On a big-endian host the value is already in big-endian order, so the conversion must do nothing, and a function that always swaps is wrong there. In practice callhtonl(<arpa/inet.h>, POSIX; "host to network long": converts a 32-bit value from host order to big-endian network order), or on GCC/Clang__builtin_bswap32(a compiler built-in that reverses the four bytes) guarded by the compiler's byte-order macro, and let the library handle the host difference. - Strict aliasing. C lets the compiler assume that an object of one type is not accessed through a pointer to an unrelated type (the strict-aliasing rule; character types are the exception).
*(const uint32_t *)pon a byte buffer breaks that rule. Usememcpyinto auint32_t(compilers turn it into one load) or build the value with shifts; both are valid in C and C++. - Alignment. A
uint32_t *pointing at an odd address is misaligned: UBSan flagged thepkt+1read above, and some CPUs fault on it. Parsing with byte loads has no alignment requirement. - Shifting signed values. Shift unsigned operands, and cast each byte to
uint32_tbefore<<(auint8_tpromotes to a signedint: for the byte0x80,0x80 << 24is0x80000000, which does not fit in a signed 32-bitintwhose maximum is0x7FFFFFFF, and in C that overflow in a shift is undefined). - Compiler quality. The shift-and-or
load_be32compiles at-O2toldr w0, [x0]plusrev w0, w0(AArch64: load four bytes, then reverse their order) andmovl (%rdi), %eaxplusbswap %eax(x86-64: the same pair, load then byte-reverse), so the portable code costs no more than the cast.
Bugs that appear when host order is assumed
Parsing a network packet or a file header by overlaying a struct or casting works on the developer's little-endian laptop only if the format happens to be little-endian. Typical failures: a TCP port or a length field read byte-reversed (80 reads as 20480, 0x50 swapped to 0x5000); a magic number that matches on one build and never on another; a loop that trusts a byte-reversed length and over-reads (a security bug, since an attacker chooses the bytes); a file written on one CPU and unreadable on another. The defence is the rule in the code: say the byte order in the format specification, read and write each field with explicit byte positions, and test on a big-endian target (an emulator is enough).
Trade-offs
Use the portable shift form by default. Reach for memcpy plus a swap when profiling shows the byte-wise form matters on a compiler that does not recognize the pattern, and for bulk data convert in a loop the compiler can vectorize (compile into instructions that handle several values at once). Do not gate behaviour on a runtime endianness probe in hot paths: it is constant for a build, so use compile-time macros there.
A C library works at -O0 but fails only in optimized builds, and the failure vanishes under the debugger. How would you investigate whether undefined behavior from pointer arithmetic, out-of-bounds access, or misaligned access is the cause, and how would you prove your fix is correct?
Sample Answer
Why this pattern points at undefined behaviour (UB). UB means the language places no requirement on what the program does; the optimizer is allowed to assume it never happens and to transform code on that assumption. At -O0 (no optimization) the compiler emits a near-literal translation, so an out-of-bounds read often "works" by reading whatever is next in memory. At -O2 (heavy optimization) the same code can be rewritten (checks deleted, loops changed). A debugger changes memory layout and usually needs -O0 or -Og (light optimization that keeps debugging usable), so the failure vanishes. Failure that depends on optimization level and disappears under the debugger is a strong hint, though not proof: uninitialized variables, data races and compiler bugs show the same pattern.
A demonstration run in a gcc:14 container (GCC 14.4.0, aarch64). The loop bound is one too large, so it reads table[8]:
#include <stdio.h>
/* search a table for a key, but the loop bound is one too large */
static int table[8] = {5, 3, 9, 1, 7, 2, 8, 4};
static int index_of(int key) {
int i;
for (i = 0; i <= 8; i++) /* table[8] is out of bounds */
if (table[i] == key) return i;
return -1;
}
int main(void) {
printf("found 4 at %d\n", index_of(4));
printf("missing key: %d\n", index_of(99));
return 0;
}
gcc -O2 -Wall table_search.c && ./table_search crashed with a segmentation fault (exit 139), and printed nothing, not even the first line. The first printf ran, but when stdout is not a terminal the C library keeps its text in a buffer and writes it out later, and a crash kills the process before that happens, so the line was lost. gcc -O0 printed found 4 at 7 and missing key: -1; so did -O2 -fno-aggressive-loop-optimizations. No warning was printed at -O2 -Wall for this file, so do not count on warnings to find it. Reading table[8] is UB, so GCC may assume the loop never reaches i == 8 with a miss, which means the only way out is the return inside the loop. The i <= 8 test is then redundant, and the -O2 assembly shows it gone: for index_of(99) the loop body is a little index arithmetic, then ldr w2, [x2, -4] (load the next table element), cmp w2, 99 (compare it with the key) and bne .L3 (branch back if different), with no comparison of the counter against 8 anywhere. The key 99 is not in the table, so the loop keeps stepping through memory past the end of table until it reaches an address the process may not read, and the segmentation fault follows. With -O0 the test is kept, the loop stops after table[8], and the program prints -1.
Ordered investigation. Steps 1 to 3 find most cases; steps 4 to 6 rule out look-alikes and a related UB case.
- Confirm it is optimization-dependent and find the boundary. Build with
-O0,-O1,-O2,-O3and compare; then bisect individual flags and, in a big library, individual files (compile half the objects at-O0, half at-O2, and narrow). Per file,__attribute__((optimize("O0")))(a GCC attribute that compiles one function at-O0) or a per-file flag makes the culprit translation unit (one.cfile with the headers it includes, which the compiler optimizes as a whole) obvious. - Turn on the compiler's own warnings:
-Wall -Wextra -Warray-bounds -Wstrict-aliasing(warnings for out-of-range array indexes and for pointer casts that break the aliasing rule), and read the warnings that only appear at-O2(they need the optimizer's analysis). - Run the sanitizers at the failing optimization level, because they instrument the code and report the exact line: UBSan (
-fsanitize=undefined) for signed overflow, bad shifts, misaligned access and out-of-bounds on arrays with known size; ASan (-fsanitize=address) for heap, stack and global out-of-bounds, use-after-free. Fortable_search.c,-O2 -fsanitize=address,undefinedprintedindex 8 out of bounds for type 'int [8]'andglobal-buffer-overflow, with the line. - Rule out the look-alikes: MemorySanitizer (a Clang sanitizer that reports reads of uninitialized memory) or
-ftrivial-auto-var-init=pattern(a GCC/Clang option that fills uninitialized locals with a fixed byte pattern so the bug shows up the same way every run) for uninitialized reads; ThreadSanitizer if threads are involved. - Misaligned access (the question mentions it) is checked by UBSan's alignment check; on some CPUs it traps, on others it is slow or silently fine.
- Pointer arithmetic UB includes forming a pointer more than one past the end of an array, not just dereferencing it; a sanitizer or
-Warray-boundsis how you find it.
Do not "fix" by lowering the optimization level or by adding -fno-strict-aliasing/-fwrapv as the end state. Those flags hide the symptom and keep the bug.
Proving the fix. Fix the bound (i < 8, or iterate with sizeof table / sizeof table[0]), then show: (a) the sanitizer build is clean on the test suite and on the original failing input, at -O0 and at -O2/-O3; (b) the outputs match between optimization levels; (c) the regression test for the failing input is committed and run in CI under both sanitizers; (d) a compiler other than GCC (clang) also builds clean, since each compiler exploits different UB.
This C code leaks memory in a long-running process: char *read_message(void){ char *buf = malloc(128); /* fill */ return buf; } void loop(void){ char *m = read_message(); process(m); }. Explain why it leaks, fix it, and describe a better ownership design for the interface.
Sample Answer
Direct answer
read_message allocates 128 bytes with malloc and returns the pointer, so ownership passes to the caller. loop stores it in m, uses it, and then m goes out of scope: the only pointer to that block is gone and nobody called free. The 128 bytes can never be reused by the process, and in a long-running loop every call leaks another 128, so memory grows without bound until the process is killed. Fix: call free(m) after process(m). The better design is to make ownership explicit in the interface: either the caller supplies the buffer (no allocation to forget) or the function returns an owned object with a documented matching release function.
Why it leaks (the mechanism)
malloc(128)takes a block from the heap and returns its address; the allocator now considers it in use.return bufcopies the address to the caller. Nothing frees it.- When
loopreturns, its localmdisappears. The allocator still believes the block is in use and has no way to know the program has lost the address. - The next call allocates a fresh block. Over 1,000,000 calls the process holds 128 x 1,000,000 = 128,000,000 bytes (about 122 MiB) of unreachable memory, plus the allocator's per-block overhead (the extra bookkeeping bytes the allocator keeps for every block it hands out; the allocator is the C library code behind
mallocandfree).
The original also never checks malloc for NULL; /* fill */ on a null pointer would crash.
Reproduce, fix, and verify
One file, three builds (original, fix 1, fix 2), compiled with gcc -g -O0 -Wall -Wextra -fsanitize=address leak.c plus an optional -DFIXED_CALLER or -DFIXED_API. AddressSanitizer includes LeakSanitizer, which reports blocks still reachable by no pointer at exit:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#ifndef FIXED_API
static char *read_message(void) {
char *buf = malloc(128);
if (!buf) return NULL;
strcpy(buf, "hello"); /* stands in for "fill" */
return buf;
}
#endif
static void process(const char *m) { printf("got %s\n", m); }
#if defined(FIXED_CALLER)
static void loop(void) {
char *m = read_message();
if (!m) return;
process(m);
free(m); /* the caller owns the buffer, so the caller frees it */
}
#elif defined(FIXED_API)
/* caller supplies the storage; returns the length the message needs (snprintf convention) */
static size_t read_message_into(char *buf, size_t cap) { return (size_t)snprintf(buf, cap, "hello"); }
static void loop(void) {
char m[128];
read_message_into(m, sizeof m);
process(m);
}
#else
static void loop(void) { /* original: m is never freed */
char *m = read_message();
process(m);
}
#endif
int main(void) { for (int i = 0; i < 3; i++) loop(); return 0; }
Original build (3 loop iterations), as printed by a real run in a GCC 14 container (the 0x... addresses and the ==13== process id change from run to run):
==13==ERROR: LeakSanitizer: detected memory leaks
Direct leak of 384 byte(s) in 3 object(s) allocated from:
#0 0xffff92ab3f08 in malloc (/usr/local/lib64/libasan.so.8+0xd3f08)
#1 0x4008d0 in read_message /w/leak.c:7
#2 0x40093c in loop /w/leak.c:32
#3 0x400968 in main /w/leak.c:37
#4 0xffff92842258 (/lib/aarch64-linux-gnu/libc.so.6+0x22258) (BuildId: 5f3926d9963ff284b8173bf3601bbb00c3abad16)
#5 0xffff92842338 in __libc_start_main (/lib/aarch64-linux-gnu/libc.so.6+0x22338) (BuildId: 5f3926d9963ff284b8173bf3601bbb00c3abad16)
#6 0x4007ec in _start (/w/leak+0x4007ec)
SUMMARY: AddressSanitizer: 384 byte(s) leaked in 3 allocation(s).
Read the stack from frame 1: the lost blocks were allocated by malloc inside read_message (leak.c:7), which loop called at leak.c:32, so loop is where the only pointer was dropped. Frames 4 to 6 are C-runtime start-up code. When stdout is a pipe, LeakSanitizer ends the process before the buffered got hello lines are flushed, so they may be missing from the capture; on a terminal they appear before the report.
The leak total is 3 calls x 128 bytes = 384 bytes, matching the mechanism. With -DFIXED_CALLER and with -DFIXED_API the run printed got hello three times and no leak report.
The snprintf convention in a call: snprintf writes at most cap - 1 characters plus a terminating NUL byte and returns the number of characters the full text would have needed, so the caller compares the return value with cap. This complete program, built with gcc -Wall -Wextra -fsanitize=address,undefined in a GCC 14 container, prints the lines below it:
#include <stdio.h>
#include <stddef.h>
/* caller supplies the storage; returns the length the message needs (snprintf convention) */
static size_t read_message_into(char *buf, size_t cap) { return (size_t)snprintf(buf, cap, "hello"); }
int main(void) {
char big[16], small[4];
size_t n1 = read_message_into(big, sizeof big); /* returns 5, text "hello" */
size_t n2 = read_message_into(small, sizeof small); /* returns 5, text "hel": truncated */
printf("big: returned %zu, cap %zu, text \"%s\", truncated: %s\n", n1, sizeof big, big, n1 >= sizeof big ? "yes" : "no");
printf("small: returned %zu, cap %zu, text \"%s\", truncated: %s\n", n2, sizeof small, small, n2 >= sizeof small ? "yes" : "no");
return 0;
}
big: returned 5, cap 16, text "hello", truncated: no
small: returned 5, cap 4, text "hel", truncated: yes
The test n >= cap is the truncation check: 5 >= 16 is false, 5 >= 4 is true.
Fix 1: free in the caller
free(m) after process(m), as in the FIXED_CALLER branch. This is correct but relies on every caller remembering, on every path. Add an early return or an error branch in loop and the leak is back. It is the minimal patch, not the interface fix.
Fix 2: a better ownership design
| Design | Interface | Who owns memory | Use when |
|---|---|---|---|
| Caller-supplied buffer | size_t read_message_into(char *buf, size_t cap) returns the length needed (the snprintf convention), caller can detect truncation | caller (here a stack array) | message has a known maximum size; hot loops; no heap needed |
| Owning return plus paired release | char *read_message(void) and void message_free(char *) | caller, but through the library's own release function | size is unbounded; library may change how it allocates; avoids freeing across different C runtimes (on Windows, a DLL is a separately built shared library, and each copy of the C runtime has its own heap manager, so a block allocated inside one must be freed by the same copy; Microsoft documents mismatches here as access violations or heap corruption) |
| Opaque handle (a pointer to a struct whose fields only the library can see, so callers can only pass it back to the library's functions) | message *m = msg_open(...) ... msg_close(m) | the handle's owner | message has internal state |
| Reuse one buffer | loop calls malloc once before the for, passes the same buffer to every iteration's read call, and calls free once after the last | the loop's owner | repeated calls with similar sizes |
Recommendation for this code: the caller-supplied buffer, because the loop reads fixed-size messages repeatedly: no allocation per iteration, no failure path from malloc, nothing to free, and the leak is impossible by construction. Switch to the owning-return-plus-message_free design if messages can be arbitrarily large, and document the ownership in the header comment ("returns a heap block; caller releases with message_free").
Edge cases
- Null return: check
mallocand make every caller check the result. - Truncation in the buffer design: compare the returned length with
cap, as in the run above. - Error paths: any early
returnbetween allocation andfreeleaks; in C the usual cure is a singlegoto out;cleanup label. - Ownership by name:
get_/borrow_for pointers the caller must not free,new_/create_/alloc_for ones it must; consistent naming makes the contract visible.
Complexity
Each call is O(1) allocation; the leak is O(n) memory over n calls. The caller-buffer design uses O(1) extra heap memory regardless of the number of calls.
You are designing a C event-loop API that stores a callback and a user-data pointer for later invocation. The callback may run long after registration, and callers can unregister at any time. What rules or safeguards would you put in place so the callback does not dereference freed memory?
Sample Answer
Direct answer
Give the callback registration a lifetime that is longer than any pending call, and make the owner of user_data explicit. Reference counting means keeping a small integer inside the object that says how many holders still use it; each holder increments it when it starts using the object and decrements it when done, and whoever brings it to zero frees the object. Practically: reference-count the registration object, let every queued or running invocation hold its own reference, make unregister only mark the entry inactive and drop the registration's reference (never free directly), check the active flag immediately before each call, and free the user data through a release callback exactly once when the count reaches zero. That removes the three failure modes: a callback that fires after unregister, a callback that unregisters itself or a neighbour while running, and a shared object freed while another callback still uses it.
Terms: an event loop is the central loop that waits for events (a timer, a socket becoming readable) and calls the callback registered for each; "re-enter" means a callback calling back into the loop (for example to unregister) while the loop is in the middle of dispatching; a deferred operation is one that records the request now and does the work later. The core mechanism is the reference count plus the active flag; the handle table and the deferred free queue under "Alternatives" are other ways to reach the same safety.
Rules to put in the API contract
- Registration returns a handle (the watcher). The caller never frees
user_datadirectly after registering; it hands ownership to the loop with arelease(user)function, or documents a different rule. unregisteris deferred, and it is called once per registration. It setsactive = 0and drops the registration reference. No further callbacks run for that watcher, including events already queued. The registration reference was the caller's only claim on the watcher, so when nothing is queued the watcher is freed right there: afterunregisterthe caller must drop its handle and must not callunregister(orpost) on it again. A second call would read freed memory (run exactly that way under AddressSanitizer, it reportsheap-use-after-free,READ of size 4). If you want a second call to be a harmless no-op, give the caller its own reference to the watcher, or use generation-numbered handles (below).- Every pending invocation owns a reference.
postincrements the count; the loop decrements it after the call or after skipping the call. This makes "the callback may run long after registration" safe, since the watcher outlives the queue entry. - Re-check
activeat call time, not at post time. - The running callback is protected too: the queue entry's reference is released only after the call returns, so a callback may unregister itself and still read its own context.
- One thread rule. The loop above is single-threaded (plain
unsignedcounts). Ifunregistercan be called from another thread, the count must be atomic and the active flag needs synchronization; say which in the API docs. - Shared objects (one callback frees what another uses): do not free by hand inside callbacks. Give shared state its own reference count (or have one owner and pass non-owning handles that are validated through the owner), so the last user frees it.
Worked example
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
typedef void (*cb_fn)(void *user, int event);
typedef void (*release_fn)(void *user);
typedef struct watcher {
cb_fn fn;
void *user;
release_fn release; /* frees user; called exactly once, when the last reference drops */
unsigned refs; /* 1 for the loop's registration + 1 per queued event (that entry's reference also covers the call while it runs) */
int active; /* cleared by unregister; checked before every call */
} watcher;
#define MAX_QUEUED 16
typedef struct { watcher *w; int event; } queued;
typedef struct { queued q[MAX_QUEUED]; size_t n; } loop;
static void w_unref(watcher *w)
{
if (--w->refs == 0) {
if (w->release) w->release(w->user);
free(w);
}
}
watcher *loop_register(cb_fn fn, void *user, release_fn release)
{
if (fn == NULL) return NULL; /* reject a NULL callback at registration, not at call time */
watcher *w = calloc(1, sizeof *w);
if (w == NULL) return NULL;
w->fn = fn; w->user = user; w->release = release; w->refs = 1; w->active = 1;
return w;
}
void loop_unregister(watcher *w)
{
if (w->active) {
w->active = 0; /* no further callbacks, even for events already queued */
w_unref(w); /* drop the registration reference; queued events keep w alive */
}
}
int loop_post(loop *lp, watcher *w, int event)
{
if (!w->active || lp->n == MAX_QUEUED) return -1;
w->refs++; /* the queue entry owns a reference until it is processed */
lp->q[lp->n++] = (queued){ w, event };
return 0;
}
void loop_run(loop *lp)
{
for (size_t i = 0; i < lp->n; i++) { /* callbacks may post more events: n can grow */
queued e = lp->q[i];
if (e.w->active)
e.w->fn(e.w->user, e.event);
w_unref(e.w); /* release the queue's reference AFTER the call */
}
lp->n = 0;
}
/* ---- demo ---- */
typedef struct { char name[16]; watcher *self; watcher *other; } session;
static void session_release(void *u) { session *s = u; printf(" free session %s\n", s->name); free(s); }
static session *new_session(const char *name)
{
session *s = calloc(1, sizeof *s);
strncpy(s->name, name, sizeof s->name - 1);
return s;
}
static void on_a(void *u, int ev)
{
session *s = u;
printf("A(%s) got event %d; unregistering B, then itself\n", s->name, ev);
loop_unregister(s->other); /* B's queued event still holds a reference, so B stays valid until that entry is skipped */
loop_unregister(s->self); /* A drops its own registration; the running call still holds a reference */
printf(" A still reads its own session after unregistering: %s\n", s->name);
}
static void on_b(void *u, int ev) { session *s = u; printf("B(%s) got event %d (must not print)\n", s->name, ev); }
int main(void)
{
loop lp = { .n = 0 };
session *sa = new_session("alpha"), *sb = new_session("beta");
watcher *wa = loop_register(on_a, sa, session_release);
watcher *wb = loop_register(on_b, sb, session_release);
sa->self = wa; sa->other = wb;
loop_post(&lp, wa, 1); /* A runs first and unregisters B */
loop_post(&lp, wb, 2); /* B is already queued when A unregisters it: its call must be skipped */
loop_run(&lp);
printf("done\n");
return 0;
}
Reading the code: loop_register creates a watcher with refs = 1 (the registration's own reference) and active = 1. loop_post adds a queue entry and bumps refs. loop_run takes each entry, calls the callback only if active is still set, and then drops the entry's reference. loop_unregister clears active and drops the registration reference. w_unref is the only place that frees, and it calls release first so the user data goes with it.
Compiled with gcc -O1 -g -Wall -Wextra -fsanitize=address,undefined loop.c (GCC 14.4.0, Linux arm64) and run, the program prints:
A(alpha) got event 1; unregistering B, then itself
A still reads its own session after unregistering: alpha
free session alpha
free session beta
done
Walk through the counts. Each watcher starts at 1 (registration); post raises each to 2. When A runs, it unregisters B (B: 2 to 1, inactive) and itself (A: 2 to 1). Reading s->name afterwards is still safe because the queue entry holds A's second reference. After the call returns, the loop drops it (A: 1 to 0) and the release function frees alpha. B's queued event is then skipped because B is inactive, and the loop drops its reference (B: 1 to 0), freeing beta. AddressSanitizer reported no errors and no leaks. ("Inactive means never called" is part of the contract: if the application unregistered A before the loop ran, A's callback would be skipped and would never get to unregister B, so cleanup must not depend on a callback running. LeakSanitizer checks that cleanup path too.)
Why not just NULL the pointer?
Setting a field to NULL on unregister only helps callers that look at that field. A copy of the pointer already in a queue still points at the freed block. Reference counting (or a handle table with generation numbers) is what makes stale references detectable or harmless. A handle table is an array of slots that callers refer to by index; a generation number is a counter stored in each slot and in the handle, bumped whenever the slot is reused, so a stale handle with an old generation is recognized as dead.
Alternatives and when to choose them
- Handle indirection: callers hold an integer ID and the loop looks it up at call time; a stale ID fails the lookup instead of dereferencing freed memory. Good when handles cross an API boundary or a thread.
- Deferred free queue: unregister marks the entry, and the loop frees marked entries only between dispatch rounds. Cheaper than counting, but you must guarantee no pending invocation can outlive a round.
- Choose reference counting when invocations can be queued across rounds, as here.
Trade-offs and pitfalls
- Reference cycles leak (a callback's context holding the watcher that holds the context): break them by unregistering explicitly.
- A release function that itself calls back into the loop can re-enter; forbid it in the contract.
- Do not decrement before the call returns; releasing early reproduces the bug.
- The caller's handle is dead once
unregisterreturns: callingunregistertwice on the same pointer is a use-after-free in this design (see rule 2).
Unlock Full Question Bank
Get access to all 44 Language-Level Memory Management (C/C++/Rust) interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.