Direct answer
Prove it on the binary that ships, not on the source. Rebuild with the identical compiler and flags, then disassemble the patched function in both builds and diff the instructions: the difference should be exactly the new check (a compare and a conditional branch to the error path, placed before the use it protects). Repeat the inspection on the final linked executable (inlining and link-time optimization can move or copy the function), and back the disassembly with a test that feeds the attack input to the shipped build. A check that is present in the source can still be missing or changed in the machine code, and the two examples below show how, for a bounds check and a constant-time comparison.
Step 1: diff the function before and after the patch
The listings in this answer are AArch64 assembly (Arm 64-bit; the destination operand comes first, as in cmp/ldrb/eor below). Under the Arm 64-bit calling convention (AAPCS64) the first eight integer or pointer arguments arrive in x0 to x7 (x86-64 System V passes the first six in rdi, rsi, rdx, rcx, r8, r9), so in copy_in, which has four, dst is x0, cap is x1, src is x2 and n is x3, and w0 is the 32-bit return value. A small function that copies into a buffer, first as before.c and then, with the fix, as after.c:
c
#include <string.h>
int copy_in(char *dst, size_t cap, const char *src, size_t n) {
(void)cap;
memcpy(dst, src, n);
return 0;
}
c
#include <string.h>
int copy_in(char *dst, size_t cap, const char *src, size_t n) {
if (n > cap)
return -1;
memcpy(dst, src, n);
return 0;
}
A helper that prints one function's instructions without addresses or encodings (fn_asm.sh). objdump -dr --no-show-raw-insn disassembles (-d) and shows relocations (-r: placeholders for addresses the linker fills in later, such as the call target) without the raw bytes; the awk program starts printing after the line <function>:, stops at the next blank line, and strips the leading address: from every line so only the instruction remains:
sh
#!/bin/sh
# usage: fn_asm.sh file.o function (prints the function's instructions without addresses or encodings)
objdump -dr --no-show-raw-insn "$1" | awk -v fn="$2" '
$0 ~ "<" fn ">:" { on = 1; next }
on && /^$/ { exit }
on { sub(/^ *[0-9a-f]+:[ \t]*/, ""); print }'
Built in a gcc:14 container (GCC 14.4.0, aarch64) with gcc -O2 -c and compared with diff <(./fn_asm.sh before.o copy_in) <(./fn_asm.sh after.o copy_in) (run with bash; the <(...) runs each helper and hands diff its output as if it were a file):
text
0a1,2
> cmp x3, x1
> b.hi 28 <copy_in+0x28> // b.pmore
6c8
< 10: R_AARCH64_CALL26 memcpy
---
> 18: R_AARCH64_CALL26 memcpy
8a11,12
> ret
> mov w0, #0xffffffff // #-1
Reading diff's notation: 0a1,2 means after line 0 of the first file, add lines 1 to 2 of the second (> marks lines only in the second file, < lines only in the first); 6c8 means line 6 changed into line 8; 8a11,12 means two lines were added after line 8. The added cmp x3, x1 (n against cap) and b.hi (branch if unsigned higher) come before the bl memcpy, and the branch goes to a block that returns -1. That is the check, in the right order relative to the use. The relocation line (R_AARCH64_CALL26 memcpy is a placeholder in a bl instruction that the linker resolves to the address of memcpy, possibly through a PLT stub, the procedure linkage table entry used for shared-library calls) changed only because its offset moved, from 0x10 to 0x18. Compiling after.c twice with the same command gave byte-identical object files (cmp reported no difference), so the same source and flags are reproducible and any real difference is meaningful. Compare with the same flags the release build uses, because changing -O level or -flto changes the instructions.
Step 2: look at the final binary
The check must be reachable in the final executable, not only in the object file. With inlining or link-time optimization the function may be merged into callers, so disassemble the linked binary (objdump -d prog), find every copy (search for the call target and for the compare constant), and confirm each one still has the compare-and-branch before the dangerous instruction. A stripped binary has no function names; locate it from the call to a known library function such as memcpy (visible through the PLT) or from the build ID's matching symbol file.
Bounds check example: present in source, bypassable in the binary
c
#include <stdio.h>
#include <limits.h>
/* Intended: reject if off + len overflows or runs past size. */
int in_bounds_wrapcheck(int off, int len, int size) {
if (off + len < off) /* "did it wrap?" */
return 0;
return off + len <= size;
}
/* Overflow-free version: the compiler is told what the check means. */
int in_bounds_builtin(int off, int len, int size) {
int end;
if (__builtin_add_overflow(off, len, &end))
return 0;
return end <= size;
}
int main(int argc, char **argv) {
(void)argv;
int off = INT_MAX - 2 + argc; /* INT_MAX - 1, so off + 100 overflows */
printf("wrapcheck: %d\n", in_bounds_wrapcheck(off, 100, 1000));
printf("builtin: %d\n", in_bounds_builtin(off, 100, 1000));
return 0;
}
Built with gcc -O2 -Wall bounds.c -o b && ./b (the same output appears at -O0):
text
wrapcheck: 1
builtin: 0
in_bounds_wrapcheck(INT_MAX - 1, 100, 1000) answers 1 (in bounds), which is wrong. The -O2 listing (gcc -O2 -S bounds.c; assembler directives and the .LFB/.LFE function-boundary labels removed, everything else as emitted) shows why:
asm
in_bounds_wrapcheck:
tbnz w1, #31, .L3
add w1, w1, w0
cmp w1, w2
cset w0, le
ret
.L3:
mov w0, 0
ret
Signed overflow is undefined behaviour (the C standard places no requirements on such a program, so the compiler may assume it never happens), so GCC rewrote off + len < off into "is len negative" (tbnz w1, #31, test bit 31 of len). The wrapped sum is never compared with off. in_bounds_builtin compiles to adds followed by bvs (branch if overflow), which is the real check. A build with -fsanitize=undefined printed a runtime error and wrapcheck: 0, which is a useful reminder that a sanitizer build is a different program: do the binary comparison on the release flags. Fix the source with __builtin_add_overflow, or by comparing as unsigned or as len <= size - off after checking 0 <= off <= size.
Constant-time comparison example
A constant-time comparison is one whose running time does not depend on the secret data being compared. It matters because an early-exit compare of, say, a secret token takes longer the more leading bytes match, and an attacker who can time it can recover the secret one byte at a time.
c
#include <stddef.h>
/* Early-exit comparison: running time depends on the first differing byte. */
int eq_early(const unsigned char *a, const unsigned char *b, size_t n) {
for (size_t i = 0; i < n; i++)
if (a[i] != b[i])
return 0;
return 1;
}
/* Accumulating comparison: touches every byte; the only branch tests the loop counter. */
int eq_accum(const unsigned char *a, const unsigned char *b, size_t n) {
unsigned char diff = 0;
for (size_t i = 0; i < n; i++)
diff |= a[i] ^ b[i];
return diff == 0;
}
c
#include <stdio.h>
#include <string.h>
int eq_early(const unsigned char *a, const unsigned char *b, size_t n);
int eq_accum(const unsigned char *a, const unsigned char *b, size_t n);
int main(void) {
unsigned char x[16], y[16];
int wrong = 0;
memset(x, 0xAB, sizeof x);
for (int pos = -1; pos < 16; pos++) { /* -1: identical; else differ at byte pos */
memcpy(y, x, sizeof y);
if (pos >= 0) y[pos] ^= 0x01;
int want = (pos < 0);
if (eq_early(x, y, 16) != want || eq_accum(x, y, 16) != want) wrong++;
}
printf("wrong results: %d\n", wrong);
return wrong != 0;
}
Built with gcc -O2 -Wall -Wextra ct.c ct_main.c -o ct_t (and run): wrong results: 0 (the test compares both functions against the expected answer for identical inputs and for a one-bit difference at each of the 16 positions). The -O2 listings (gcc -O2 -S ct.c; assembler directives and the .LFB/.LFE function-boundary labels removed, everything else as emitted):
asm
eq_early:
cbz x2, .L4
mov x3, 0
b .L3
.L8:
cmp x2, x3
beq .L4
.L3:
ldrb w5, [x0, x3]
ldrb w4, [x1, x3]
add x3, x3, 1
cmp w5, w4
beq .L8
mov w0, 0
ret
.L4:
mov w0, 1
ret
eq_accum:
cbz x2, .L12
mov x3, 0
mov w4, 0
.L11:
ldrb w5, [x0, x3]
ldrb w6, [x1, x3]
add x3, x3, 1
eor w5, w5, w6
orr w4, w4, w5
and w4, w4, 255
cmp x2, x3
bne .L11
cmp w4, 0
cset w0, eq
ret
.L12:
mov w0, 1
ret
What to check, and what the listings show:
- In
eq_early, the two ldrb instructions load one byte from each array, cmp w5, w4 compares them, and beq .L8 loops on equality while a mismatch falls through to mov w0, 0; ret. cbz x2 (compare and branch if zero) handles n == 0. The exit time depends on the first differing byte: this is not constant time. It is what the code asked for.
- In
eq_accum, the branches test only the length and counter (cbz x2, cmp x2, x3; bne); the data only flows through eor (exclusive or: nonzero where bytes differ) and orr (or: accumulates any difference) into w4. and w4, w4, 255 keeps only the low 8 bits, matching the unsigned char type of diff. w4 is tested once at the end with cmp w4, 0 and cset w0, eq, which sets w0 to 1 when equal and 0 otherwise, with no branch. That preserved the intended property at -O2.
- At
-O3 GCC 14.4 vectorized eq_accum (compiled to use SIMD instructions that handle 16 bytes at once: NEON loads, eor/orr on vectors, then a horizontal OR reduction that folds the lanes together), and every conditional branch in that listing is again a compare of a length or index register; the accumulated difference is tested only by cmp w3, 0 and cset. The main vector loop of that listing is .L12, shown complete here (the horizontal reduction after it and the scalar tail loops are left out):
asm
.L12:
ldr q16, [x0, x3]
ldr q7, [x1, x3]
add x3, x3, 16
eor v7.16b, v16.16b, v7.16b
orr v30.16b, v30.16b, v7.16b
cmp x3, x4
bne .L12
Each pass loads 16 bytes of each array (q registers), XORs them and ORs the result into v30; the only branch compares the byte offset x3 with the end offset x4.
So for this compiler and flags the property held, but it is a property of the generated code, not a guarantee of the C language: a different compiler version or flag set, or a different target, can produce something else, so the check is repeated on every shipped build configuration. Timing measurements on the target hardware (statistical tests such as dudect, a tool that times a function on two classes of input and tests whether the timing distributions differ) test the same property from the outside; the disassembly check cannot see data-dependent instruction latencies, only branches and addresses.
Checklist for the patch
- Same compiler, flags and source revision for the "before" and "after" builds, and a reproducibility check (build twice, diff).
- Per-function instruction diff, ignoring address noise, whose only change is the check, ahead of the protected use.
- The same inspection in the final linked, possibly LTO-optimized, binary, at every inlined copy.
- For each kind of check: bounds checks need a compare whose operands are the real quantities (no rewritten sign test, no removed branch); comparisons that must be constant-time need no data-dependent branch or data-dependent address.
- A negative test with the attack input against the shipped binary, which fails on the "before" build and passes on the "after" build.