Cryptographic Protocol Design and Analysis Questions
Designing and reasoning about cryptographic protocols and secure channels: how message flows, key-exchange handshakes, and end-to-end encryption systems are constructed so that composing individual primitives yields a provably or informally verified secure whole. Covers authentication and key-exchange protocol design (mutual authentication, forward secrecy, key confirmation, key-compromise-impersonation resistance), message-flow and state-machine security, formal and informal protocol verification (BAN logic, symbolic tools such as ProVerif and Tamarin, game-based reduction proofs), protocol-level vulnerability analysis (downgrade, replay, padding-oracle, algorithm-confusion attacks), TLS handshake and key-schedule internals, and end-to-end encryption system design (ratcheting, group key agreement, key transparency, post-compromise security). This is the design and analysis layer: why a protocol construction is secure, not which library call or key-management process to run in production. Distinct from selecting and operating cryptographic primitives day to day (certificate lifecycle management, TLS deployment monitoring and incident response, key rotation operations, algorithm and parameter selection for a given constraint set), which belongs to applied cryptography and key management; from core cryptographic vocabulary and primitive fundamentals; and from implementation-level bugs (side-channel leakage, memory-safety flaws, timing attacks in code), which belong to cryptographic implementation security.
Explain differences between MAC-then-encrypt, encrypt-then-MAC, and AEAD constructions used in TLS record protection. Provide examples of practical attacks (padding oracle, MAC oracle) that motivated the move to AEAD ciphers in modern TLS deployments.
Sample Answer
Direct answer
MAC-then-encrypt computes an integrity check (MAC, Message Authentication Code) over the plaintext first and then encrypts the plaintext-plus-MAC together; encrypt-then-MAC encrypts first and computes the integrity check over the resulting ciphertext; AEAD (Authenticated Encryption with Associated Data) does both in a single cipher operation instead of two separate steps. The reason modern TLS (Transport Layer Security) moved to AEAD-only is that MAC-then-encrypt with CBC (Cipher Block Chaining) mode forces the receiver to remove block padding before it can check the MAC, and whether that padding removal succeeds or fails is itself information an attacker can learn and exploit, the padding-oracle attack class behind Lucky 13 and POODLE.
Structured elaboration
MAC-then-encrypt (TLS 1.2 CBC suites): the sender computes MAC = HMAC(key, plaintext), appends it to the plaintext, pads the result to a block-size multiple, then encrypts with a block cipher in CBC mode. The receiver has no choice but to decrypt first and strip the padding before it can even see where the MAC bytes are to check them. This ordering is the root problem: an attacker who can send manipulated ciphertext to the receiver and observe whether it is rejected at the padding-removal step versus rejected later at the MAC-check step (or accepted) has a distinguishing oracle. Given many trial ciphertexts and that oracle, an attacker can recover plaintext byte by byte without ever knowing the encryption key, this is the padding-oracle attack.
Encrypt-then-MAC (RFC 7366, an option for TLS 1.2): the sender encrypts first, then computes MAC = HMAC(key, ciphertext) over the result. The receiver checks the MAC first, against the ciphertext exactly as received, and only proceeds to decrypt and remove padding if that check passes. A tampered ciphertext is rejected at the MAC-check step, uniformly, before any padding logic ever runs, so there is no second, distinguishable rejection path for an attacker to exploit. This ordering is generically secure for any combination of a secure cipher and a secure MAC, which is a real theoretical difference from MAC-then-encrypt, whose security depends on the specific pairing chosen.
AEAD (mandatory in TLS 1.3, for example AES-GCM or ChaCha20-Poly1305): confidentiality and integrity are provided by one cipher construction, with an authentication tag computed as part of the same operation, and no block-aligned padding step at all for stream/counter-mode AEAD ciphers. This removes the padding-oracle attack surface by construction rather than by getting the check-ordering right, there is no padding step for an attacker to distinguish behavior around in the first place.
Attacks that motivated the shift: POODLE (2014) exploited SSL 3.0's underspecified CBC padding to build a padding oracle; Lucky 13 (2013) showed that even TLS implementations which appeared to check padding and MAC correctly still leaked a measurable timing difference between the two rejection paths, letting the same class of attack work over a network purely from response timing, without needing an outright different error message.
Worked example
This demonstrates the structural mechanism (which rejection path fires) rather than a full timing attack, since a real timing side channel is environment-dependent and not something to claim reproducible numbers for. Both constructions are built with the same key and plaintext, then a byte in the second-to-last ciphertext block is XOR-perturbed with every possible mask, exactly the classic CBC padding-oracle setup, since that block's bytes get XORed into the final plaintext block (where the padding lives) during CBC decryption:
import hmac, hashlib
from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes
from cryptography.hazmat.primitives.padding import PKCS7
KEY, MAC_KEY, IV = b"\x11"*16, b"\x22"*32, b"\x33"*16
def aes_cbc_encrypt(pt):
padder = PKCS7(128).padder()
padded = padder.update(pt) + padder.finalize()
enc = Cipher(algorithms.AES(KEY), modes.CBC(IV)).encryptor()
return enc.update(padded) + enc.finalize()
def aes_cbc_decrypt_raw(ct):
dec = Cipher(algorithms.AES(KEY), modes.CBC(IV)).decryptor()
return dec.update(ct) + dec.finalize()
def mac_then_encrypt_receive(ct):
padded = aes_cbc_decrypt_raw(ct)
unpadder = PKCS7(128).unpadder()
try:
unpadded = unpadder.update(padded) + unpadder.finalize()
except ValueError:
return "REJECTED at padding-removal step"
msg, tag = unpadded[:-32], unpadded[-32:]
expected = hmac.new(MAC_KEY, msg, hashlib.sha256).digest()
if not hmac.compare_digest(tag, expected):
return "REJECTED at MAC-check step (padding parsed fine)"
return "ACCEPTED"
def encrypt_then_mac_receive(record):
ct, tag = record[:-32], record[-32:]
expected = hmac.new(MAC_KEY, ct, hashlib.sha256).digest()
if not hmac.compare_digest(tag, expected):
return "REJECTED at MAC-check step (ciphertext never touched)"
return "ACCEPTED (would decrypt now)"
plaintext = b"transfer=100"
mte_record = aes_cbc_encrypt(plaintext + hmac.new(MAC_KEY, plaintext, hashlib.sha256).digest())
ct_only = aes_cbc_encrypt(plaintext)
etm_record = ct_only + hmac.new(MAC_KEY, ct_only, hashlib.sha256).digest()
mte_outcomes, etm_outcomes = {}, {}
for pos in range(len(mte_record) - 32, len(mte_record) - 16):
for mask in range(1, 256):
t = bytearray(mte_record); t[pos] ^= mask
o = mac_then_encrypt_receive(bytes(t))
mte_outcomes[o] = mte_outcomes.get(o, 0) + 1
for pos in range(len(ct_only) - 16, len(ct_only)):
for mask in range(1, 256):
t = bytearray(ct_only); t[pos] ^= mask
o = encrypt_then_mac_receive(bytes(t) + etm_record[-32:])
etm_outcomes[o] = etm_outcomes.get(o, 0) + 1
print("MAC-then-encrypt outcomes:", mte_outcomes)
print("Encrypt-then-MAC outcomes:", etm_outcomes)
Output (actually executed, using the pinned key/IV/plaintext above and 4,080 tamper trials per construction):
MAC-then-encrypt outcomes: {'REJECTED at MAC-check step (padding parsed fine)': 3061, 'REJECTED at padding-removal step': 1019}
Encrypt-then-MAC outcomes: {'REJECTED at MAC-check step (ciphertext never touched)': 4080}
MAC-then-encrypt has two distinguishable rejection reasons, exactly the distinguisher a padding-oracle attack amplifies into full plaintext recovery given enough trial requests. Encrypt-then-MAC has exactly one rejection outcome across all 4,080 tamper trials, there is no distinguisher available to an attacker at all.
Trade-offs and pitfalls
Encrypt-then-MAC is a strict improvement over MAC-then-encrypt for any generic cipher/MAC pairing, but it was only ever an optional TLS 1.2 extension (RFC 7366), meaning a server or client not configured to negotiate it could still fall back to the weaker default; AEAD in TLS 1.3 removes that configuration risk by making the safe construction the only construction available. A separate, easy interview mistake is assuming AEAD is just "encrypt-then-MAC with extra steps," it is not, it is a single primitive with its own security proof, not a composition of two independently-chosen primitives, which matters because AEAD's security guarantee depends on the nonce never repeating under the same key, a failure mode with no analogue in the MAC-then-encrypt or encrypt-then-MAC discussion at all.
In a protocol with asynchronous delivery (for instance push notifications to mobile devices or intermittent IoT connectivity), how would you formally verify that the protocol remains secure under message reordering and duplicates? Propose modeling strategies, invariants to assert, and practical mitigations to handle network nondeterminism.
Sample Answer
Direct answer
Model the network itself as unreliable: instead of an ordered, exactly-once pipe, treat it as an unordered, replay-capable multiset (a "bag") of messages that an active adversary, or ordinary asynchronous delivery, can deliver in any order, drop, delay indefinitely, or hand over more than once, and then require every secrecy and authentication property to hold across every reachable interleaving of that bag, not just the one sequence a happy-path test would exercise. On top of that verification model, enforce two practical mitigations at the implementation layer: idempotent processing keyed by a stable message identifier, so a duplicate delivery produces no additional effect, and a sequence-number-plus-sliding-window check that bounds how far out of order a message can arrive before it is rejected outright.
Structured elaboration
Modeling the channel as a multiset, not a queue. The symbolic verification tools used for this kind of analysis (ProVerif and Tamarin, the two standard tools for this style of protocol verification) already default to something close to this: their attacker process owns the public channel, can intercept anything sent on it, and can resend a previously observed message as many times as it likes, in any order relative to other traffic. The risk is not that the tool lacks this capability, it is that the PROTOCOL model implicitly assumes an ordering the tool never actually enforces. A receiver-side process definition that pattern-matches "first this message, then that one" using a rigid sequential structure can silently encode an in-order assumption the real deployment does not honor, so the modeling discipline here is to write the receiver's logic so it accepts messages in ANY order the multiset can present them, and let the verifier explore the full space of orderings and duplicate counts rather than constraining it to a single expected sequence.
What changes about the invariants. In a model that assumes reliable, ordered, exactly-once delivery, an authentication property can informally lean on "the message the receiver just processed is the most recent one sent." That framing breaks the moment delivery is asynchronous: the invariant instead has to say "every accepted message corresponds to some message the claimed sender genuinely sent, independent of what order anything else arrived in," which is really the same non-injective agreement property used for synchronous protocols, just checked against a multiset of sent events instead of a sequence. Duplication adds a second, distinct invariant that a synchronous model does not need at all: processing the same message twice must produce the same observable state as processing it once (an idempotency invariant), rather than silently repeating whatever side effect the message triggers. It is worth being explicit that duplication here is not purely adversarial: a push-notification provider or an intermittently-connected device will legitimately re-send a message it never got an acknowledgment for, so the model has to prove the property holds under BOTH attacker-induced replay and the protocol's own normal, self-inflicted retransmission behavior, and a verification effort that only tests the adversarial case has not actually covered the deployment's real traffic pattern.
Concrete modeling strategies.
- Define the receiver's state-transition rules so that receiving a given message is always enabled whenever that message exists anywhere in the network multiset, with no additional precondition tying it to "the previous message was X." A rule that can only fire after some other specific rule has already fired is quietly reintroducing an ordering assumption.
- Verify the target properties as invariants that must hold in every reachable state of the resulting state-transition system, not as assertions checked only along one designated "expected" execution trace; this is exactly the difference between hand-written test cases (which check one ordering at a time) and a symbolic model (which explores the full reachable set).
- Explicitly include a self-replay action distinct from the attacker's replay action, representing the legitimate sender re-transmitting a message it believes was lost, so the model captures duplication that originates from the protocol's own retry logic rather than only from an adversary.
Practical mitigations, once the model tells you what to build.
- Idempotent processing keyed by a message identifier. Attach a stable identifier to every state-changing message (a content hash or a server-issued id, distinct from a raw sequence counter so it survives even if sequence numbers themselves get reset or reused). Before applying any effect, check a durable de-duplication record for that identifier; if it is already present, treat the message as successfully handled already and return the same result without reapplying the effect. This turns duplicate delivery into a no-op by construction, which is what the idempotency invariant above requires.
- Sequence-number-plus-sliding-window enforcement. A monotonic per-sender counter attached to each message lets the receiver bound how far out of order it will accept traffic: track the highest sequence number seen and accept anything within a bounded window behind it, rejecting anything older as stale and anything already marked seen inside the window as a duplicate. This exact mechanism, including how to size the window and how to cache keys for messages that arrive out of order in a ratcheting design, follows the same window-and-cache logic used for nonce and replay handling in general; the short version needed here is that the window bounds the reordering the receiver tolerates, while the identifier-keyed idempotency check above is what actually makes an accepted duplicate harmless rather than merely detected.
Worked example
Consider a tiny batch of three distinct messages, m1,m2,m3, sent to a device over a push-notification channel with at-least-once delivery. Even ignoring duplicates entirely, there are 3!=6 possible delivery orders the receiver might see. A hand-written test suite that only checks the intended order (m1,m2,m3 in sequence) exercises exactly one of those six orderings, meaning a property "proved" against that single test has not actually been checked against the other five reachable cases at all, which is precisely the gap a symbolic model closes by exploring the full multiset rather than one fixed list.
Now allow the channel to also duplicate delivery, a routine consequence of at-least-once push semantics when an acknowledgment is lost. Say the device goes briefly offline and reconnects to find m3 delivered first, followed by m1, then a duplicate re-delivery of m1 (the provider never received the device's earlier acknowledgment), then m2. Applying the two mitigations: each message carries an identifier and a sequence number (m1 = seq 10, m2 = seq 11, m3 = seq 12). The receiver processes m3 first, extending its window; processes m1, which is behind the window's current high-water mark but still within the tolerated span, and applies it, recording its identifier as seen; sees the duplicate m1 next, finds its identifier already recorded, and discards it as a no-op rather than reapplying the effect; then processes m2. The final state, message content applied exactly once each for m1, m2, and m3, is identical to what the "happy path" ordering would have produced, which is exactly the property the model was asked to prove holds across every reachable interleaving, not merely observed in this one illustrative trace.
Trade-offs and pitfalls
- Sizing the sliding window and the idempotency record's retention period both trade tolerance for state: too narrow, and a device that reconnects after a genuinely long gap (a week offline, not a few missed pushes) gets treated as sending stale traffic and rejected outright rather than being resynchronized through a separate recovery path; too wide, and a large device fleet forces the server to retain a correspondingly large amount of per-device de-duplication state indefinitely.
- A model that only verifies the property under reordering, or only under duplication, has not verified the actual threat model; both must be exercised together, since a protocol can pass a reordering-only check and still fail once duplicates are added (for example, if the idempotency key were derived from the sequence number alone rather than a full message identifier, a duplicate with a re-used sequence number could slip past a window check that only looks at "have I seen this sequence number as the newest one" without a proper seen-set).
- Internet of Things (IoT) devices in particular tend to combine long offline gaps with severely limited local storage, which pulls in opposite directions: the server-side window and de-duplication state want to be generous to tolerate the gaps, while the device's own retry and acknowledgment bookkeeping wants to be minimal; a design that only reasons about server-side state without accounting for what the constrained device itself can durably remember across a power cycle will look correct in the model and still fail in the field.
- It is tempting to treat "the properties still hold" as a one-time result; a protocol change that adds a new message type or a new state-changing side effect needs the multiset-delivery verification re-run against the updated model, since the earlier proof said nothing about a message type that did not exist when it was checked.
Given a protocol that uses certificates for client authentication, contrast the security implications when the attacker is (a) an active network MITM, (b) a rogue CA that issues certificates, and (c) a compromised client device. For each scenario, explain which properties (confidentiality, authentication, non-repudiation) are affected and recommend mitigations.
Sample Answer
Direct answer
"Certificates protect the connection" is not one fact, it is three separate claims that fail independently. An active network MITM (man-in-the-middle: an attacker positioned between the two parties who can read and alter traffic in transit) generally cannot break a correctly validated TLS (Transport Layer Security) connection's authentication or confidentiality. A rogue CA (certificate authority: an entity relying parties trust to vouch that a public key belongs to a claimed identity) breaks authentication for every relying party at once, because it can mint a certificate that says anything it wants. A compromised client device breaks confidentiality of traffic tied to that key and destroys non-repudiation, even though the protocol itself worked exactly as designed.
Structured elaboration
| Attacker | Confidentiality | Authentication | Non-repudiation | Primary mitigation |
|---|---|---|---|---|
| Active network MITM (no CA access) | Intact, if certificate validation (chain, hostname, revocation) is done correctly and no downgrade succeeds | Intact for the same reason: the attacker cannot produce a certificate any client will accept for an identity it doesn't hold | N/A here (this attacker never gets to sign anything as the client) | Correct chain/hostname validation, certificate pinning for high-value clients, downgrade-resistant negotiation |
| Rogue CA | At risk downstream: the server now grants the attacker's connection the trust level of a legitimate client, exposing whatever that client role can see | Broken: the whole point of a CA is that a valid certificate means something, and a rogue one can issue a certificate for any identity it likes | Broken: a signature under a fraudulently issued certificate proves nothing reliable about who really controls the corresponding key | Certificate Transparency (CT: public, append-only logs of every certificate a participating CA issues, so unexpected issuances are auditable), CA pinning/short trusted-CA lists, multi-CA issuance monitoring |
| Compromised client device | At risk for any traffic the key ever protected: past sessions are exposed if the handshake used static/non-forward-secret key transport, future sessions are exposed regardless | Intact at the protocol layer: the "right" key really did authenticate, because the attacker now holds it | Broken: any signature produced with the stolen key is indistinguishable from one the legitimate user produced | Hardware-backed (HSM/TPM: hardware security module / trusted platform module, physical hardware that stores a private key and performs signing or decryption with it internally, without ever exposing the key in plaintext) non-exportable keys, short-lived certificates with fast revocation, ephemeral (EC)DHE key exchange for forward secrecy, step-up authentication for sensitive operations |
Worked example
Take a concrete client-cert-authenticated banking API and walk each attacker through the same request, "transfer $500":
- MITM, no CA access: the attacker cannot present a certificate the server accepts as the real client, so the request is simply never authenticated as coming from the legitimate account holder. The attack fails at the handshake, before any application data is exposed.
- Rogue CA: the attacker obtains a certificate that names the victim's identity (the CA issued it without real verification). The server accepts the handshake as if the real client connected, and the fraudulent request is processed as authentic. Confidentiality of the underlying transport channel to the attacker is not "broken" by an eavesdropper's standard, because the attacker legitimately negotiated that channel, but the server's authentication decision was wrong, which is the actual damage.
- Compromised device: the attacker has the real client's key material directly (say, exfiltrated from an unlocked laptop). Every future request from that device looks completely legitimate to the server, and if the original handshake years ago used static RSA key transport rather than ephemeral (EC)DHE, an attacker who additionally captured old encrypted traffic can now decrypt it too.
Trade-offs and pitfalls
- TLS client certificates generally do not provide strong non-repudiation, regardless of which attacker scenario applies. The CertificateVerify signature in the handshake only proves possession of the private key during that specific handshake transcript; it says nothing about a specific business action, and the key usually lives in a general-purpose OS/browser keystore any process on the device can invoke. Treating "the connection was certificate-authenticated" as courtroom-grade proof of a user's intent is a common and incorrect leap.
- People frequently say "MITM" for both the plain network attacker and the rogue-CA case, but the trust-model implications are completely different: correct validation defeats the first and is powerless against the second, because a rogue CA attacks the trust anchor itself, not the network path.
- Revocation checking (OCSP/CRL: Online Certificate Status Protocol and Certificate Revocation List, the two standard mechanisms a relying party uses to check whether a certificate has been revoked before its stated expiry) is the mitigation everyone assumes is "on," and it is frequently soft-fail or effectively disabled in practice, which quietly removes the primary defense against both the rogue-CA and compromised-device scenarios.
Design a handshake and server-side anti-replay strategy that supports 0-RTT data while preserving replay protection and forward secrecy for non-0-RTT data. Specify client state to reuse, server-side caches or tokens, KDF binding of 0-RTT to context, and operational limitations to impose on 0-RTT traffic.
Sample Answer
Direct answer
Combine two independent mechanisms: a PSK (pre-shared key) binder that cryptographically ties an offered resumption ticket to this specific ClientHello, so a captured ticket can't be spliced onto a different one, and single-use redemption of that ticket for early data specifically, so a replayed ClientHello can still complete the connection as ordinary 1-RTT but can never get its early data applied twice. Non-0-RTT data keeps its forward secrecy by using psk_dhe_ke mode, PSK combined with a fresh (EC)DHE (elliptic-curve Diffie-Hellman Ephemeral) contribution, once the full handshake completes.
Structured elaboration
- Client state to reuse. The resumption ticket issued by the server after a prior full handshake, and the resumption PSK it names; the client presents the ticket identifier and computes a binder over the new ClientHello using a key derived from that PSK.
- Server-side caches or tokens. A small map from ticket identifier to (PSK, expiry); on top of that, a separate, bounded set of ticket identifiers already redeemed for 0-RTT, checked and updated atomically at connection time and evicted once the ticket's own validity window ends, so this state never grows past "recent tickets currently valid."
- KDF (key derivation function) binding of 0-RTT to context. The binder itself is
HMAC(binder_key, transcript_so_far), an HMAC being a keyed cryptographic checksum, wherebinder_keyis derived from the PSK via HKDF (HMAC-based key derivation function). Because it covers the ClientHello bytes up to that point, an attacker who captured a valid (ticket, binder) pair cannot reuse it with a different ClientHello, changing even one byte before the binder invalidates the match; this is a distinct property from anti-replay of the identical message, it stops splicing rather than resending. - Operational limitations on 0-RTT traffic. Restrict early data to requests the application has independently marked idempotent (the same discipline that keeps a mutating request safe under any replay, not just this transport's); bound ticket lifetime tightly enough that the anti-replay set never needs to hold more than one lifetime's worth of identifiers; and, per RFC 8446 (the TLS, Transport Layer Security, 1.3 specification), section 8.1 (Single-Use Tickets, the mechanism implemented below) together with the fallback rule in section 4.2.10, never abort the handshake purely because early data was rejected as a replay, fall back to completing the connection as 1-RTT so the mechanism can't be turned into a denial-of-service lever against legitimate retries.
Worked example
"""
0-RTT (zero round-trip time: the client sends encrypted application data
in its very first flight, before the handshake finishes) anti-replay design.
Two mechanisms shown together, matching real TLS 1.3 (RFC 8446):
(a) a PSK (pre-shared key) binder: an HMAC over the ClientHello-so-far that
cryptographically ties the offered resumption ticket to THIS specific
ClientHello, so a captured ticket can't be spliced into a different one.
(b) single-use ticket redemption for early data: replaying the exact same
ClientHello+ticket a second time must not let the *early data* (e.g. a
purchase) execute twice, even though the connection itself may still be
allowed to proceed as ordinary 1-RTT.
Stdlib only, pinned inputs.
"""
import hashlib
import hmac
HASH = hashlib.sha256
def hkdf_extract(salt: bytes, ikm: bytes) -> bytes:
return hmac.new(salt, ikm, HASH).digest()
def hkdf_expand(prk: bytes, info: bytes, length: int) -> bytes:
t, okm, counter = b"", b"", 1
while len(okm) < length:
t = hmac.new(prk, t + info + bytes([counter]), HASH).digest()
okm += t
counter += 1
return okm[:length]
class Server:
def __init__(self):
# ticket_id -> (resumption_psk, expiry_epoch)
self.tickets = {}
# ticket_ids already redeemed for 0-RTT early data (bounded by ticket lifetime)
self.zero_rtt_redeemed = set()
def issue_ticket(self, resumption_psk: bytes, now: int, lifetime: int = 3600) -> bytes:
ticket_id = bytes.fromhex("560b7ff61ec849cfa8cd145a46e66520")[:16] # pinned, synthetic
self.tickets[ticket_id] = (resumption_psk, now + lifetime)
return ticket_id
def handle_client_hello(self, ticket_id: bytes, binder: bytes, transcript_prefix: bytes,
early_data: bytes | None, now: int):
entry = self.tickets.get(ticket_id)
if entry is None:
return {"psk_accepted": False, "early_data_accepted": False, "reason": "unknown/expired ticket"}
psk, expiry = entry
if now > expiry:
return {"psk_accepted": False, "early_data_accepted": False, "reason": "ticket expired"}
# (a) PSK binder check: proves the client possesses the PSK AND that
# this exact ClientHello (not a different one) is what was bound.
binder_key = hkdf_expand(hkdf_extract(b"", psk), b"res binder", 32)
expected_binder = hmac.new(binder_key, transcript_prefix, HASH).digest()
if not hmac.compare_digest(binder, expected_binder):
return {"psk_accepted": False, "early_data_accepted": False, "reason": "binder mismatch"}
result = {"psk_accepted": True, "early_data_accepted": False, "reason": None}
if early_data is not None:
# (b) anti-replay: a ticket may fund early data AT MOST ONCE,
# the "Single-Use Tickets" mechanism of RFC 8446 section 8.1.
# Per section 4.2.10, a replayed early-data attempt does not
# abort the handshake; it just falls back to ordinary 1-RTT
# (early data is simply not applied) so we don't hand the
# attacker a free connection-establishment DoS lever either.
if ticket_id in self.zero_rtt_redeemed:
result["reason"] = "early data rejected: ticket already redeemed for 0-RTT (falling back to 1-RTT)"
else:
self.zero_rtt_redeemed.add(ticket_id)
result["early_data_accepted"] = True
return result
def client_hello(ticket_id: bytes, resumption_psk: bytes, session_nonce: bytes) -> tuple[bytes, bytes]:
transcript_prefix = b"ClientHello||" + ticket_id + b"||" + session_nonce
binder_key = hkdf_expand(hkdf_extract(b"", resumption_psk), b"res binder", 32)
binder = hmac.new(binder_key, transcript_prefix, HASH).digest()
return transcript_prefix, binder
if __name__ == "__main__":
server = Server()
resumption_psk = bytes.fromhex("deadbeef" * 8)
now = 100_000
ticket_id = server.issue_ticket(resumption_psk, now=now, lifetime=3600)
print("ticket issued:", ticket_id.hex())
# -- legitimate client sends 0-RTT early data (e.g. "PUT /cart/checkout") --
session_nonce = bytes.fromhex("a1a2a3a4a5a6a7a8") # pinned, synthetic
transcript_prefix, binder = client_hello(ticket_id, resumption_psk, session_nonce)
early_data = b'{"op":"checkout","cart":"c_501"}'
r1 = server.handle_client_hello(ticket_id, binder, transcript_prefix, early_data, now=now + 1)
print("live 0-RTT attempt: ", r1)
# -- network attacker captures and replays the exact same ClientHello+early data --
r2 = server.handle_client_hello(ticket_id, binder, transcript_prefix, early_data, now=now + 2)
print("replayed 0-RTT attempt: ", r2)
# -- connection can still proceed WITHOUT early data (1-RTT fallback) --
r3 = server.handle_client_hello(ticket_id, binder, transcript_prefix, early_data=None, now=now + 3)
print("1-RTT fallback (no early data claimed):", r3)
# -- attacker tries to splice the captured ticket_id+binder onto a DIFFERENT
# ClientHello transcript (different session_nonce) -> binder fails to verify --
forged_prefix = b"ClientHello||" + ticket_id + b"||" + bytes.fromhex("b1b2b3b4b5b6b7b8") # pinned, different nonce
r4 = server.handle_client_hello(ticket_id, binder, forged_prefix, early_data, now=now + 4)
print("binder spliced onto a different ClientHello:", r4)
Output:
ticket issued: 560b7ff61ec849cfa8cd145a46e66520
live 0-RTT attempt: {'psk_accepted': True, 'early_data_accepted': True, 'reason': None}
replayed 0-RTT attempt: {'psk_accepted': True, 'early_data_accepted': False, 'reason': 'early data rejected: ticket already redeemed for 0-RTT (falling back to 1-RTT)'}
1-RTT fallback (no early data claimed): {'psk_accepted': True, 'early_data_accepted': False, 'reason': None}
binder spliced onto a different ClientHello: {'psk_accepted': False, 'early_data_accepted': False, 'reason': 'binder mismatch'}
The live 0-RTT attempt is accepted. The identical replayed ClientHello is rejected for early data specifically but the PSK itself still validates, so a later attempt with no early-data claim still succeeds as a normal 1-RTT-equivalent connection. A binder computed for one ClientHello transcript, spliced onto a different one, fails immediately, independent of the anti-replay cache.
Trade-offs and pitfalls
- The binder and the single-use cache defend against different attacks and are not substitutes for each other: the binder stops transcript-splicing, the cache stops verbatim replay of an otherwise-identical message; a design with only one of the two is still exploitable by the attack the other one covers.
- Sharding the redemption cache across a load-balanced fleet (the same operational concern as a high-throughput mesh) trades a slightly weaker replay guarantee (a ticket might briefly be usable on two different shards before state converges) for avoiding a single shared-cache bottleneck; this is a deliberate, stated trade, not an oversight, as long as the window of inconsistency is bounded and short.
- It is tempting to reject the whole connection on a detected 0-RTT replay; doing so turns an attacker's captured, replayed ClientHello into a way to deny service to the legitimate client's own retries, which is why the correct response is to drop only the early data and continue the handshake normally.
A protocol you maintain allows third-party negotiated extensions at handshake time. A new extension that bypasses a key confirmation step caused a security regression. Design an extension-safety policy that allows safe extensibility without weakening core security guarantees. The policy should cover a specification language for extensions, static checks, dynamic runtime guards, and the vetting and deployment process.
Sample Answer
Direct answer
The policy needs four layers working together, because no single layer alone would have caught a real extension that "bypassed a key confirmation step": a machine-checkable specification language that forces every extension to declare its security-relevant behavior, static checks (including formal verification) run before an extension ships, runtime guards that enforce a monotonic security floor regardless of what an extension claims, and a staged vetting and deployment process with a kill switch. The specific regression described, an extension silently skipping key confirmation, is caught by the runtime guard layer even if the static-check layer misses it, which is exactly why the design needs both, not either.
Structured elaboration
- Specification language. A small, typed, signed manifest per extension declaring: capabilities requested, intended state changes, cryptographic primitives used, and explicit security properties it claims to preserve (for example, "does not bypass key confirmation"). Machine-checkable structure, not free-text documentation, is what makes the next two layers possible.
- Static checks (pre-deployment). Type and schema validation of the manifest; semantic checks that enforce protocol invariants (an extension cannot declare that it skips key confirmation and still be accepted); where feasible, translate the extension's effect on the handshake into a model for a symbolic verification tool (ProVerif or Tamarin) and prove no downgrade of authentication or confidentiality properties results.
- Runtime guards. Capability negotiation with least privilege: an extension is granted only the specific capabilities its manifest requested, nothing implicit. A monotonic security lattice: no operation an extension performs is allowed to lower the security level already established, specifically including a hard rule that key confirmation cannot be skipped by any extension, full stop, enforced by the core handshake state machine rather than by trusting the extension to behave. Unknown or invalid extensions fail closed (handshake aborts) rather than silently degrading.
- Vetting and deployment. Design review, formal verification, cryptographic peer review, fuzzing and interoperability testing, then a staged canary rollout with monitoring, in that order. Require signed manifests and reproducible builds so what was reviewed is provably what ships. Keep an operational kill switch that can disable a specific extension fleet-wide without a full deployment cycle.
Worked example
Applying this concretely to the described regression: the offending extension's manifest would have had to declare either "skips key confirmation" (a static check rejects this outright, since it directly contradicts the core-guarantee invariant) or nothing about key confirmation at all (a runtime guard still catches it, because the guard does not trust the manifest's silence; it independently verifies, at the protocol state-machine level, that a key-confirmation message was actually exchanged before the session is marked established, regardless of which extensions ran). This is the concrete difference between a policy that only checks what an extension says it does, versus one that also verifies what the core protocol actually did happen, and it is why the runtime-guard layer exists even when static checks look sufficient on paper.
Trade-offs & pitfalls
Formal verification (ProVerif/Tamarin, symbolic tools for proving properties of cryptographic protocols) is powerful but has real limits: it verifies the model you gave it, and a mismatch between that model and the actual implementation is a known failure mode, so pair it with implementation-level testing rather than trusting the proof alone. A fail-closed policy on unknown extensions is safer but has an availability cost: a legitimate new extension that has not yet been reviewed will simply not negotiate, which needs to be an accepted trade-off up front, not discovered as a surprise outage during a partner integration. Finally, a kill switch is only useful if it was exercised before the incident it is meant to handle; an untested kill switch is a false sense of security, so it belongs in the same staged-rollout testing the extension itself goes through.
Unlock Full Question Bank
Get access to all Cryptographic Protocol Design and Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.