Cryptography Fundamentals Questions
Core concepts and vocabulary of cryptography: confidentiality, integrity, authentication, and non-repudiation; the difference between symmetric and asymmetric primitives; and how standard algorithms, libraries, and protocols fit together. Covers threat models, common standards, and applying primitives and cryptographic libraries correctly to real-world security problems. The entry point for the cryptography track.
You're reviewing a colleague's pull request that adds encryption for a new feature. Walk through what you'd look for during review, and for each issue you flag, explain why it's a problem and how you'd fix it.
Sample Answer
Direct answer
A crypto-focused code review is really a checklist walk: where does the randomness come from,
where do keys live, which algorithm and mode was picked, how is the IV/nonce generated, is
there an integrity check on top of confidentiality, and does error handling leak anything
useful to an attacker. Most real-world crypto bugs live in these plumbing decisions, not in
the mathematics of the algorithm itself.
Structured elaboration
- Randomness source. Is the key, IV, or nonce generated from a CSPRNG (the platform's
cryptographic random source), or from a general-purposerandom()/Math.random()-style
function? The latter is predictable and unsafe for any of these values; flag it and replace
it with the platform's cryptographic RNG API. - Key handling. Is the key hardcoded in source, a config file, or a constant, versus
pulled from a secrets manager or KMS (key management service) at runtime? A hardcoded key
means anyone with source access (or a leaked repo) has the key forever, and rotating it
requires a code deploy. Fix: load it from a secrets store, never commit it. - Algorithm and mode choice. Is the cipher mode ECB (Electronic Codebook)? ECB encrypts
identical plaintext blocks to identical ciphertext blocks, visibly leaking structure (the
classic demonstration is an ECB-encrypted image where the outline is still recognizable).
Flag any ECB usage and replace it with an AEAD mode like AES-GCM. - IV/nonce handling. Is a fresh IV/nonce generated per encryption, or is one hardcoded or
reused across calls? A repeated IV under CBC leaks structure; a repeated nonce under GCM is
a full break of that key (see the nonce-reuse mechanics above). Fix: generate fresh
per-message values from a CSPRNG and store them alongside the ciphertext. - Integrity, not just confidentiality. Does the code use a plain cipher with no
authentication at all? Encryption without an integrity check lets an attacker tamper with
ciphertext undetected. Fix: use an AEAD cipher, or add a properly-ordered encrypt-then-MAC
if the library does not offer AEAD. - Library maturity and API usage. Is this a hand-rolled implementation of a
well-known primitive, or a vetted library used through its high-level, misuse-resistant API?
Flag any custom crypto code; recommend a maintained library instead. - Error handling. Do error messages or timing distinguish "wrong key" from "wrong padding"
from "tampered ciphertext"? Distinguishable errors can become an oracle an attacker probes
repeatedly to recover plaintext without ever learning the key. Fix: return one generic
failure for all decryption errors, and use constant-time comparison for any secret-derived
values.
Worked example
A pull request adds AES.new(key, AES.MODE_ECB) with key = b"my-static-key-1234567890123456"
hardcoded at the top of the file, and no separate authentication step. The review comment
would name three separate, stackable problems: (1) ECB leaks plaintext structure regardless of
key strength, switch to AES.MODE_GCM; (2) the key is a literal in source, move it to the
secrets manager the rest of the service already uses; (3) there is no authentication tag at
all, so tampered ciphertext decrypts silently, which GCM's tag would catch for free once
adopted.
Trade-offs & pitfalls
- Not every review will find all of these at once; the discipline is running the same
checklist every time rather than pattern-matching on whatever looks unusual first. - A review that only checks "is the algorithm modern" and stops there misses the majority of
real crypto bugs, which live in key handling, IV/nonce handling, and error behavior, not
algorithm choice.
Design an efficient cryptographic scheme for group messaging where a sender encrypts a message to multiple recipients. Minimize redundant encryption work while preserving recipient confidentiality and supporting efficient revocation and forward/backward secrecy. Discuss use of key-wrapping, group keys, tree-based key exchange (e.g., MLS-like), and trade-offs.
Sample Answer
Direct answer
Minimize redundant work by encrypting message content once under a shared, regularly rotated group secret rather than once per recipient, then solve the actual hard problem, keeping that group secret in sync efficiently as membership changes, with a tree-structured key-agreement scheme so a single add or remove costs work proportional to the tree's height, not to the group's size.
Why not encrypt to each recipient individually
Naive per-recipient public-key encryption of the same message costs one expensive asymmetric operation per recipient per message, an O(N) cost that doesn't scale and that every real design in this space specifically avoids.
Approach A: a single shared group key
Encrypt content once with a symmetric AEAD (Authenticated Encryption with Associated Data) key shared by the whole group, cheap per message. The weakness is on membership change: removing a member for confidentiality requires distributing a brand-new key to every remaining member, an O(N) operation every time someone leaves, expensive for large or frequently-changing groups.
Approach B: sender keys
Each member distributes their own per-sender ratcheting key to the rest of the group once, over pairwise-encrypted one-to-one channels (an O(N) cost, but only at join time, not per message). After that, a member's messages are AEAD-encrypted once with their own evolving key and broadcast, no further asymmetric work per message. The weak point is removal: a removed member who cached earlier key material can still derive future messages unless every remaining sender's key is explicitly re-issued, so clean removal requires proactive re-keying rather than coming for free.
Approach C: tree-based group key agreement (TreeKEM, as standardized in MLS, the Messaging Layer Security protocol, RFC 9420)
The group secret is derived from a binary tree whose leaves are members and whose interior nodes each hold a secret derivable from their two children.
graph TD
Root[Group secret at root]
A[Interior node A]
B[Interior node B]
L1[Alice leaf]
L2[Bob leaf]
L3[Carol leaf]
L4[Dave leaf]
Root --> A
Root --> B
A --> L1
A --> L2
B --> L3
B --> L4
Adding or removing a member only requires updating the path from that member's leaf up to the root, at most log2N nodes, versus O(N) for approach A. Because the interior secrets along that path are refreshed, the design achieves both forward secrecy (compromising today's group secret doesn't reveal past epochs) and post-compromise security (a fresh update heals the group even after a past compromise, since future secrets don't derive from anything an attacker previously learned). To remove a member, blank their leaf and update the path to the root exactly as with an add, cryptographically excluding them from every subsequent epoch without touching the other members' keys individually.
Worked example: removing Carol from the four-person tree above
Blank Carol's leaf. Recompute interior node B's secret from Dave's leaf secret plus a fresh random contribution. Recompute the root secret from node A's existing, unaffected secret and B's new secret. That's two update messages, to node B and to the root, instead of individually re-keying Alice, Bob, and Dave under approach A's flat shared key. For four people the gap is small, two updates instead of three, but it widens as group size grows, since it stays log2N against approach A's O(N).
Content encryption itself
Regardless of which approach maintains the group secret, message content is always encrypted with a fast symmetric AEAD keyed by the current epoch's secret, or a key derived from it via HKDF per sender or per message. The group-key-management layer and the message-encryption layer stay cleanly separable.
Trade-offs and pitfalls
TreeKEM's logarithmic scalability is the reason MLS exists, but it brings real protocol complexity: two members updating the tree at the same moment need a defined conflict-resolution mechanism (MLS uses an ordering "delivery service" to serialize commits), something a flat shared-key scheme never has to think about. Sender keys are far simpler to implement correctly and remain a reasonable choice for smaller, less frequently changing groups where O(N) rekey-on-remove is acceptable. The flat shared-group-key approach, despite being the simplest to reason about, degrades worst under exactly the workload, large, frequently-changing membership, that makes efficient revocation matter in the first place.
Evaluate the security implications of moving random number generation for containerized workloads from the OS CSPRNG to a remote entropy service. Discuss risks to confidentiality, integrity, and availability of entropy; possible attack vectors (man-in-the-middle, compromised service, replay), bootstrapping and DRBG design choices, and propose cryptographic safeguards (authenticated channels, signed entropy, local DRBG reseeding, HSM-sealed seeds) and operational mitigations.
Sample Answer
Direct answer
Moving random-number generation from the OS CSPRNG (cryptographically secure pseudorandom number generator) to a remote entropy service introduces two separate threats that must be handled separately: the network path to the service can be attacked (man-in-the-middle, replay), and the service itself can be malicious or compromised. The safe design treats remote entropy only as one reseed input to a local DRBG (deterministic random bit generator), never as key material directly, and never as the sole trusted source.
Confidentiality, integrity, and availability risks
- Confidentiality: if an attacker can observe the entropy delivered to a container, they can predict any key or nonce derived from it. This is the core risk, "confidentiality of the RNG's output" and "confidentiality of everything downstream of it" are the same thing.
- Integrity: a compromised or malicious service can feed biased, low-entropy, or attacker-known values. This is not a hypothetical: Dual_EC_DRBG, a NIST-standardized random-bit generator later shown to contain a backdoor letting whoever held a specific secret value predict its output, is documented history showing a standardized, trusted randomness source can be compromised at the design level, not just in transit.
- Availability: if the remote service is unreachable, does the system fail closed safely, or silently fall back to a weaker local source without anyone noticing? A silent fallback is often the more dangerous failure mode.
Attack vectors
Man-in-the-middle substitution of predictable values in transit; a compromised or malicious provider supplying biased output from the source itself, which no transport security can fix; and replay of previously observed "random" values, which is catastrophic for nonce-based modes like CTR or GCM if the same value gets reused across time or across machines.
Bootstrapping and DRBG design
The moment a remote entropy service is most tempting, and most dangerous to trust exclusively, is at boot: a freshly booted container or VM has accumulated little local entropy and may even share initial state with sibling instances. The safe pattern is to seed and periodically reseed a local DRBG (NIST SP 800-90A's CTR_DRBG or HMAC_DRBG are standard choices) by mixing the remote entropy with local sources (the OS CSPRNG, a hardware RNG if present) through a secure combiner, typically a hash-based extractor, so the design stays secure as long as at least one input source is honest and high-entropy, rather than requiring the remote service alone to be trustworthy.
Cryptographic safeguards and operational mitigations
- An authenticated, encrypted channel (mutual TLS) to the entropy service to block a network-level substitution attack.
- Signed entropy batches, verified before mixing in, so a compromised network path (though not necessarily a compromised provider) cannot substitute predictable values.
- Statistical health testing of delivered entropy blocks (NIST SP 800-90B-style minimum-entropy estimation) to catch a degraded or stuck feed before it's trusted.
- Fail-closed behavior: if the channel can't be authenticated or entropy fails a health check, don't silently fall back to a weaker path without operator visibility.
- An HSM (hardware security module)-sealed local seed for the very first boot, so the bootstrapping moment doesn't depend on the remote service being trustworthy at all.
Trade-offs and pitfalls
An HSM-sealed local seed combined with a well-designed local DRBG mostly removes the need for a remote entropy service in the first place for most workloads. A remote service really only solves the "freshly cloned container with too little local entropy" problem, and in most cloud environments that problem has a cheaper, lower-risk fix already available (a hardware RNG or the platform's virtio-rng passthrough). This design should be scrutinized against whether the dependency is needed at all, just as much as how to harden it if it stays.
Give a high-level walkthrough of a TLS handshake: what happens at each step, which cryptographic primitives are used where (certificates, asymmetric key exchange, symmetric session keys, MAC/AEAD), and what properties TLS is actually trying to guarantee. What's one common way this can fail in practice?
Sample Answer
Direct answer
A TLS (Transport Layer Security) handshake lets a client and server who have never met agree
on a shared symmetric key, authenticate the server (and optionally the client), and start
exchanging data, with confidentiality, integrity, and (with modern configuration) forward
secrecy. It works by combining asymmetric key exchange and certificate-based identity checks
up front, then switching to fast symmetric encryption for the actual traffic.
Structured elaboration
sequenceDiagram
participant C as Client
participant S as Server
C->>S: ClientHello (supported versions, cipher suites, key share, random)
S->>C: ServerHello (chosen cipher suite, key share, random)
S->>C: Certificate (server public key, signed by a CA)
S->>C: Finished (proves possession of the certificate's private key)
Note over C,S: Both derive the same shared secret via ECDHE, then a session key via a KDF
C->>S: Finished (handshake integrity check)
C->>S: Application data (encrypted with the negotiated AEAD cipher)
S->>C: Application data (encrypted with the negotiated AEAD cipher)
- ClientHello / ServerHello: negotiate protocol version, cipher suite, and exchange
ephemeral key-exchange material (an elliptic-curve public value for ECDHE) plus random
nonces that feed the key derivation. - Certificate: the server presents its certificate, a public key bound to an identity and
signed by a certificate authority (CA), so the client can authenticate who it is actually
talking to (see PKI/certificate chains for how that trust is validated). - Key derivation: both sides combine the ECDHE exchange output through a KDF (key
derivation function) to produce the actual symmetric session keys, never sending the
session key itself over the wire. - Application data: all subsequent traffic is encrypted and authenticated with an AEAD
cipher (typically AES-GCM or ChaCha20-Poly1305) under the derived session key; older TLS
versions could instead pair a plain block cipher with a separate MAC (message
authentication code) computed over each record, a split that AEAD replaces with one
combined operation. - Properties TLS is guaranteeing: confidentiality (only the endpoints can read the data),
integrity/authenticity (tampering is detected), server authentication (the certificate
chain proves identity), and, with an ephemeral key-exchange suite, forward secrecy.
Worked example / common failure
The most common real-world failure is not a cryptographic break, it is trust management:
an expired certificate. Everything about the handshake's math can be flawless and the
connection will still be rejected (or a warning shown) the moment the leaf certificate's
validity window has passed, because clients check expiration as part of chain validation
before they will trust the negotiated key exchange at all. This is a frequent cause of
sudden, otherwise-unexplained outages, and is why certificate expiry monitoring is treated as
an operational, not just a security, concern.
Trade-offs & pitfalls
- TLS 1.3 (RFC 8446) simplified this flow and removed insecure options entirely (no more
static-RSA key exchange, no more weak cipher suites), which is why upgrading from TLS 1.2
is itself a meaningful security improvement, not just a version bump. - A valid, unexpired certificate chain proves identity, it does not by itself prove the
identity is trustworthy; certificate validation and business-level trust are separate
concerns.
Implement PKCS#7 padding and unpadding functions in Python for a block cipher with a given block size. Provide: def pad(data: bytes, block_size: int) -> bytes and def unpad(data: bytes, block_size: int) -> bytes. The unpad function must validate padding and raise an exception for invalid padding, and you should explain how to validate padding without creating an easy padding oracle vulnerability during decryption flows.
Sample Answer
Direct answer
PKCS#7 padding (defined in RFC 2315, and reused unchanged in modern standards for block ciphers)
fills the last block of a message so its length becomes a multiple of the cipher's block size: the
value of every padding byte equals the count of padding bytes added, from 1 up to the block
size. Unpadding must check that consistently and reject anything else, but it must do that check
without leaking, through timing, which specific byte was wrong, since that leak is exactly what a
padding oracle attack exploits.
Structured elaboration
- Approach.
padcomputes how many bytes are missing to reach a multiple ofblock_size
(always between 1 andblock_size, even when the input is already a multiple, since padding
must always be present sounpadcan find it) and appends that many bytes, each holding that
count as its value. - Approach.
unpadfirst checks the length is a nonzero multiple ofblock_size, reads the
last byte as the claimed padding count, and rejects it outright if it is 0 or bigger than
block_size. Only then does it verify every one of the claimed padding bytes actually holds
that value. - Avoiding a padding oracle. The verification loop below never returns as soon as it finds a
mismatch; it OR's every byte's difference into an accumulator and checks it once at the end, so
the number of comparisons performed does not depend on where the first bad byte was. That
alone does not make the surrounding system safe: the real defense is to authenticate the
ciphertext (via an AEAD cipher, authenticated encryption with associated data, or a Message
Authentication Code, MAC, such as HMAC, computed over the ciphertext) before this function
is ever called during decryption, so a tampered ciphertext is rejected before its padding is
inspected at all.
Worked example
class PaddingError(ValueError):
pass
def pad(data: bytes, block_size: int) -> bytes:
if not (1 <= block_size <= 255):
raise ValueError("block_size must be between 1 and 255")
pad_len = block_size - (len(data) % block_size)
return data + bytes([pad_len]) * pad_len
def unpad(data: bytes, block_size: int) -> bytes:
if not data or len(data) % block_size != 0:
raise PaddingError("invalid padding length")
pad_len = data[-1]
if pad_len < 1 or pad_len > block_size:
raise PaddingError("invalid padding value")
# Constant-time check: touch every candidate padding byte instead of
# returning on the first mismatch, so timing does not reveal *which*
# byte was wrong.
diff = 0
for b in data[-pad_len:]:
diff |= b ^ pad_len
if diff != 0:
raise PaddingError("invalid padding bytes")
return data[:-pad_len]
block_size = 16
msg = b"attack at dawn!!" # 16 bytes, exactly one block
padded_short = pad(b"hi", block_size)
padded_exact = pad(msg, block_size)
print("pad(b'hi', 16) =", padded_short, "len:", len(padded_short))
print("pad(16-byte msg, 16) =", padded_exact, "len:", len(padded_exact))
print("unpad round-trips 'hi' :", unpad(padded_short, block_size) == b"hi")
print("unpad round-trips full :", unpad(padded_exact, block_size) == msg)
tampered = bytearray(padded_short)
tampered[-1] ^= 0xFF
try:
unpad(bytes(tampered), block_size)
print("tampered padding accepted (BUG)")
except PaddingError as e:
print("tampered padding rejected:", e)
Output (deterministic, this exact run reproduces it):
pad(b'hi', 16) = b'hi\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e\x0e' len: 16
pad(16-byte msg, 16) = b'attack at dawn!!\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10\x10' len: 32
unpad round-trips 'hi' : True
unpad round-trips full : True
tampered padding rejected: invalid padding value
Trade-offs & pitfalls
- Complexity: both functions are O(n) in the input length; the constant-time scan adds no
asymptotic cost since it is bounded byblock_size(at most 255 byte comparisons). - Edge cases: an already-block-aligned input still gets a full extra block of padding (a common
point of confusion, PKCS#7 never leaves zero padding bytes); an empty ciphertext; apad_lenof
0 or greater thanblock_size, both invalid and must be rejected before the byte-by-byte check
runs; a ciphertext whose length is not a multiple ofblock_sizeat all, which is a format error
distinct from a padding error but should still surface a single generic exception. - A constant-time comparison inside
unpadreduces one class of timing leak, but if this function
is called on attacker-controlled ciphertext before a MAC or AEAD tag has been checked, the
system is still exploitable: the attacker does not need a timing side-channel if the application
simply returns a different HTTP status or error string for "bad padding" versus "bad MAC".
Encrypt-then-MAC (or AEAD outright) removes the padding oracle at the protocol level; this
function alone only removes it at the instruction level.
Unlock Full Question Bank
Get access to all Cryptography Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.