Cryptography Fundamentals Questions
Core concepts and vocabulary of cryptography: confidentiality, integrity, authentication, and non-repudiation; the difference between symmetric and asymmetric primitives; and how standard algorithms, libraries, and protocols fit together. Covers threat models, common standards, and applying primitives and cryptographic libraries correctly to real-world security problems. The entry point for the cryptography track.
Compare an encrypt-then-MAC construction (e.g. AES-CBC + HMAC) against a dedicated AEAD cipher like AES-GCM for protecting an HTTP API payload. Cover performance, hardware acceleration, streaming support, IV/nonce requirements, and where each approach is more likely to be implemented incorrectly.
Sample Answer
Direct answer
Both encrypt-then-MAC (AES-CBC plus a separately computed HMAC) and a dedicated AEAD cipher
like AES-GCM protect confidentiality and integrity together, but AEAD bundles them into one
call with one failure mode to get right, while encrypt-then-MAC is two separate primitives you
must sequence correctly yourself. For a new HTTP API payload, AES-GCM is the default choice
unless there's a specific reason to prefer the older combination.
Structured elaboration
- Performance and hardware acceleration: AES-GCM benefits from dedicated CPU instructions
(AES-NI plus carryless multiplication for the authentication step), making it very fast on
essentially all modern server and mobile hardware. Encrypt-then-MAC with AES-CBC plus HMAC
requires two separate passes over the data (one for the cipher, one for the MAC), roughly
doubling the work, though both individual pieces can still be hardware-accelerated. - Streaming support: HMAC can be updated incrementally as data arrives, since it just
needs the full message to compute a final digest, which suits streaming reasonably well.
AES-GCM also supports incremental processing but has an internal limit on how much data one
nonce can safely encrypt before the underlying counter risks exhaustion; for very large or
long-lived streams under one key, that limit needs to be tracked. - IV/nonce requirements: CBC needs an unpredictable IV per message (repetition leaks
structure but is not immediately catastrophic). GCM needs a nonce that is never reused
under the same key (repetition is catastrophic: see the nonce-reuse mechanics above, up to
full key-material exposure). - Where each is more likely to be implemented incorrectly: encrypt-then-MAC has several
places to get wrong: MAC-then-encrypt or encrypt-and-MAC (rather than encrypt-then-MAC)
order can reopen padding-oracle-style attacks; the MAC comparison must be constant-time to
avoid a timing oracle; padding itself must be handled carefully. AES-GCM collapses most of
that into a single library call, but concentrates all the risk into one rule: never reuse a
nonce. In distributed services, where many stateless instances encrypt independently under
a shared key, coordinating nonce uniqueness (a monotonic counter plus a per-instance ID, or
purely random 96-bit nonces) becomes the operational version of the same problem, and it is
what drives the key-rotation volume worked out below. A cipher with a larger nonce space,
like XChaCha20-Poly1305's 192-bit nonce, removes that
practical ceiling by making random-nonce collisions negligible even at very high message
volumes.
Worked example
Take NIST's rule of thumb that a random 96-bit nonce under one key should be retired once
roughly 2^32 messages have been encrypted. That 2^32 is NIST's deliberately conservative
invocation limit, chosen to keep the IV-collision probability down around 2^-32 to 2^-33; it
sits far below the true birthday bound of 2^48 (the square root of the 96-bit space, the point
where the chance of some collision reaches about 50 percent). 2^32 = 4,294,967,296. A service encrypting 10 million API payloads per day reaches
that count in 4,294,967,296 / 10,000,000 ≈ 429.5 days, a little over a year. That is a
concrete, calendar-relevant reason to build key rotation into the design up front rather than
treating AES-GCM as "safe forever" once it is wired up.
Trade-offs & pitfalls
- If you must support encrypt-then-MAC (say, for interoperability with an existing system),
the order matters: authenticate the ciphertext, not the plaintext, and use a
constant-time comparison when checking the tag. - "AEAD is simpler" does not mean "AEAD is foolproof"; it moves the single point of failure to
nonce management, which still requires real design attention in a distributed system.
Give a high-level walkthrough of a TLS handshake: what happens at each step, which cryptographic primitives are used where (certificates, asymmetric key exchange, symmetric session keys, MAC/AEAD), and what properties TLS is actually trying to guarantee. What's one common way this can fail in practice?
Sample Answer
Direct answer
A TLS (Transport Layer Security) handshake lets a client and server who have never met agree
on a shared symmetric key, authenticate the server (and optionally the client), and start
exchanging data, with confidentiality, integrity, and (with modern configuration) forward
secrecy. It works by combining asymmetric key exchange and certificate-based identity checks
up front, then switching to fast symmetric encryption for the actual traffic.
Structured elaboration
sequenceDiagram
participant C as Client
participant S as Server
C->>S: ClientHello (supported versions, cipher suites, key share, random)
S->>C: ServerHello (chosen cipher suite, key share, random)
S->>C: Certificate (server public key, signed by a CA)
S->>C: Finished (proves possession of the certificate's private key)
Note over C,S: Both derive the same shared secret via ECDHE, then a session key via a KDF
C->>S: Finished (handshake integrity check)
C->>S: Application data (encrypted with the negotiated AEAD cipher)
S->>C: Application data (encrypted with the negotiated AEAD cipher)
- ClientHello / ServerHello: negotiate protocol version, cipher suite, and exchange
ephemeral key-exchange material (an elliptic-curve public value for ECDHE) plus random
nonces that feed the key derivation. - Certificate: the server presents its certificate, a public key bound to an identity and
signed by a certificate authority (CA), so the client can authenticate who it is actually
talking to (see PKI/certificate chains for how that trust is validated). - Key derivation: both sides combine the ECDHE exchange output through a KDF (key
derivation function) to produce the actual symmetric session keys, never sending the
session key itself over the wire. - Application data: all subsequent traffic is encrypted and authenticated with an AEAD
cipher (typically AES-GCM or ChaCha20-Poly1305) under the derived session key; older TLS
versions could instead pair a plain block cipher with a separate MAC (message
authentication code) computed over each record, a split that AEAD replaces with one
combined operation. - Properties TLS is guaranteeing: confidentiality (only the endpoints can read the data),
integrity/authenticity (tampering is detected), server authentication (the certificate
chain proves identity), and, with an ephemeral key-exchange suite, forward secrecy.
Worked example / common failure
The most common real-world failure is not a cryptographic break, it is trust management:
an expired certificate. Everything about the handshake's math can be flawless and the
connection will still be rejected (or a warning shown) the moment the leaf certificate's
validity window has passed, because clients check expiration as part of chain validation
before they will trust the negotiated key exchange at all. This is a frequent cause of
sudden, otherwise-unexplained outages, and is why certificate expiry monitoring is treated as
an operational, not just a security, concern.
Trade-offs & pitfalls
- TLS 1.3 (RFC 8446) simplified this flow and removed insecure options entirely (no more
static-RSA key exchange, no more weak cipher suites), which is why upgrading from TLS 1.2
is itself a meaningful security improvement, not just a version bump. - A valid, unexpired certificate chain proves identity, it does not by itself prove the
identity is trustworthy; certificate validation and business-level trust are separate
concerns.
Explain the roles of salting and key stretching in password-based key derivation and storage. Describe how salts should be generated and stored, why unique salts prevent precomputation/rainbow-table attacks, and how key stretching (iterative hashing, memory-hard functions) increases attacker work. Illustrate with concrete examples of attacks that salting and stretching mitigate and note any remaining risks that require additional controls.
Sample Answer
Direct answer
Salting adds a random, unique-per-user value to a password before hashing so that identical
passwords never produce identical stored hashes; key stretching deliberately makes each
hashing attempt slow (or memory-hungry) so an attacker who steals the hash dump still has to
pay a real per-guess cost. Together they turn "crack the whole database at once with a
precomputed table" into "crack each password individually, slowly."
Structured elaboration
- Generating and storing salts: generate a fresh, random salt (commonly 16 bytes) from a
CSPRNG per user, and store it alongside the hash. The salt is not secret, it can sit in
plaintext in the same database row; its job is uniqueness, not confidentiality. - Why unique salts defeat precomputation: a rainbow table (a precomputed map from common
passwords to their hashes) is only useful because it lets an attacker look up a hash instead
of computing one. A unique salt per user means the same password produces a different hash
for every user, so an attacker would need a separate precomputed table per salt value, which
is infeasible at any real user-base size. - How key stretching increases attacker work: an iterated hash (many rounds of a hash
function) or a memory-hard function (Argon2id, scrypt) makes each individual guess
expensive in CPU time, memory, or both. This does not slow down a legitimate login (one
verification per attempt is cheap in absolute terms) but multiplies the cost of trying
billions of candidate passwords offline by the same factor.
Worked example
Suppose 5,000 users in a leaked database chose the password "password123". With a plain,
unsalted hash, all 5,000 rows show the identical hash value, so cracking it once (via a
rainbow table lookup or a single guess) reveals all 5,000 accounts simultaneously. With unique
per-user salts, those same 5,000 users produce 5,000 different hash values, and an attacker
must run the guess-and-check process separately against each one; a rainbow table built for
the unsalted case is useless here since it was never built for any of these specific salts.
Trade-offs & pitfalls (remaining risks needing additional controls)
- Salting and stretching only make offline cracking (against a stolen hash) expensive; they
do nothing against online guessing at the login endpoint, which still needs rate limiting
and account lockout/backoff. - A short or common password remains crackable even when properly salted and stretched,
because the search space itself is small; this is why NIST SP 800-63B and similar guidance
push for password length and breach-list checking rather than relying on hashing alone. - Choosing stretching parameters too low (to save server CPU) narrows the gap between "safe"
and "fast enough to crack anyway"; parameters need periodic revisiting as hardware gets
cheaper.
Your service currently stores passwords with a weak or legacy hashing scheme (say, unsalted SHA-1). Design the migration to a modern KDF like Argon2id: how you pick parameters for production versus constrained clients, how you support a rolling upgrade so users aren't forced to reset immediately, how you detect and re-hash the remaining weak legacy entries over time, and what you'd monitor to catch offline cracking attempts against the old hashes.
Sample Answer
Direct answer
Migrate by upgrading each user's hash the moment you have their plaintext password in hand,
which is only at successful login, so nobody is forced to reset. Track which scheme each
record is on, tune Argon2id's memory/time cost to the device that will actually run it, sweep
up stragglers who never log back in, and watch for signs the old dump is being cracked
offline in the meantime.
Structured elaboration
1. Argon2id parameters, production vs constrained clients. Argon2id has three costs:
memory (KiB), iterations, and parallelism. OWASP's Password Storage Cheat Sheet lists a
sliding scale of roughly equal-strength options that trade memory for time, for example
around 46 MiB of memory with 1 iteration at parallelism 1 for a normal server, versus roughly
12 to 19 MiB with 2 to 3 iterations when memory is the constrained resource (a mobile client
doing local verification, or a server handling very high login concurrency where 46 MiB per
concurrent hash would exhaust RAM). Pick the highest memory cost your slowest realistic
target (server fleet under peak login load, or the weakest client device) can sustain in
under roughly half a second; more memory is what actually makes GPU/ASIC cracking expensive,
since those devices have comparatively little fast memory per core.
2. Rolling upgrade with no forced reset. Store the hash in a self-describing format that
names its own algorithm (Argon2id's standard encoded string already does this;
$2b$... for bcrypt or a raw hex digest identifies legacy SHA-1). On login, verify against
whatever scheme the stored value says it is. If verification succeeds and the scheme is not
the current one, you now hold the plaintext password for one request: immediately hash it
with Argon2id and overwrite the stored value before returning the response. Every active user
silently upgrades on their next successful login; nobody sees a reset prompt.
3. Detecting and re-hashing stragglers. Track an algo/version column per credential
row and report the shrinking count of legacy-SHA1 rows over time as a rollout metric. Users
who never log in cannot be upgraded this way, because you never see their plaintext again.
For those, set a deadline (say, after a defined inactivity window) past which you expire the
legacy hash and require a password reset via email on next login attempt, since silently
leaving weak hashes in place indefinitely defeats the migration.
4. Monitoring for offline cracking against the old hashes. You cannot watch cracking
happen inside an attacker's own hardware, so the signal has to come from what a successful
crack produces: seed a handful of unused canary accounts with known passwords hashed under
the same legacy SHA-1 scheme, and alert if those exact credentials are ever used to log in;
watch for a spike in successful logins from unfamiliar IPs/devices concentrated on
still-legacy accounts, which suggests a cracked batch is being replayed; and correlate
against breach-notification feeds (has any of these emails' passwords shown up in a public
breach corpus). None of this proves nothing is being cracked, but it turns "we hope not" into
a concrete tripwire.
Worked example
User alice has a legacy row: algo=sha1, hash=<sha1 hex digest>. She logs in with her real
password. The server sees algo=sha1, verifies the SHA-1 digest against the submitted
password, and it matches. Because verification succeeded, the server now holds alice's
plaintext password for the remainder of this one request. It immediately computes
argon2id(password, salt, m=19456, t=2, p=1), overwrites the row to
algo=argon2id, hash=<new encoded hash>, and returns the normal login response. Alice never
saw a reset prompt; her next login will verify against Argon2id instead. A dashboard query
counting WHERE algo = 'sha1' trending toward zero over weeks is the migration's own progress
metric; whatever remains after the inactivity deadline gets force-reset instead of waiting
indefinitely.
Trade-offs & pitfalls
- Skipping the self-describing hash format is the single most common mistake: without a way
to tell schemes apart per row, you cannot dispatch verification correctly during the
transition. - Setting Argon2id memory too low "to be safe on old hardware" quietly reduces it to
something closer to a fast hash, defeating the point of the migration. - Forcing every user to reset immediately is simpler to build but throws away the accounts of
anyone who does not see the email in time; the rolling upgrade exists specifically to avoid
that support and churn cost.
Explain the core security properties a cryptographic hash function must have (preimage resistance, second-preimage resistance, collision resistance), name a couple of common algorithms, and give one example each of where a hash is the right tool (integrity checks, content-addressing) and where it is not (storing passwords without salting and stretching).
Sample Answer
Direct answer
A cryptographic hash function must be preimage resistant (given an output, you cannot find
an input that produces it), second-preimage resistant (given one input, you cannot find a
different input with the same output), and collision resistant (you cannot find any two
distinct inputs that share an output at all). Common algorithms today are SHA-256 and SHA-3;
MD5 and SHA-1 are broken for collision resistance and should not be used where that property
matters.
Structured elaboration
- Preimage resistance: given
H(m) = h, it should be computationally infeasible to find
anymthat producesh. This is what makes a hash "one-way." - Second-preimage resistance: given a specific
m1, it should be infeasible to find a
differentm2withH(m1) = H(m2). - Collision resistance: it should be infeasible to find any pair
m1 != m2with
H(m1) = H(m2), without fixing either input in advance. This is the property MD5 and SHA-1
have both practically broken (real collision attacks have been demonstrated for both). - Where a hash is the right tool: integrity checks (comparing a downloaded file's hash
against a published one to detect corruption or tampering), content-addressing (Git commit
hashes, package registries keying content by its digest so identical content always maps to
the same address), and as the digest step inside a digital signature: signing algorithms
sign a fixed-size hash of a message rather than the message itself, since asymmetric
operations are expensive on large inputs. - Where it is not the right tool: storing passwords with a bare, unsalted hash. Hash
functions are deliberately fast, which is exactly wrong for password storage: an attacker
with a stolen hash dump can try billions of guesses per second on commodity GPUs. Password
storage needs a slow, memory-hard key derivation function (Argon2id, scrypt, bcrypt), not a
general-purpose hash used directly.
Worked example
To see why collision resistance is the strongest property, compare a real hash to a toy,
clearly-broken one: H(x) = x mod 97. Take m1 = 5 and m2 = 102. 102 mod 97 = 5, so
H(5) = H(102) = 5: a collision found by inspection, with no search at all. That is exactly
what "not collision resistant" looks like. A real hash like SHA-256 produces a 256-bit output,
so the best known way to find any colliding pair is a generic birthday-bound search of roughly
2^128 attempts (the square root of 2^256 possible outputs), which is computationally out
of reach with current or foreseeable hardware. The mod-97 toy fails instantly because its
output space is only 97 values; a cryptographic hash needs both a large output space and no
structural shortcut into it.
Trade-offs & pitfalls
- Collision resistance and preimage resistance are different properties; a function can lose
one without immediately losing the other, which is why "MD5 has collisions" does not
automatically mean "MD5 leaks passwords," but it does mean you should not trust MD5 for
anything where an attacker choosing two colliding inputs would matter (like forging a signed
document). - "Hashing" and "encryption" are often confused: hashing is one-way and not meant to be
reversed at all, encryption is two-way and meant to be reversed with the right key. A hash
cannot be "decrypted" back to the original input, by design.
Unlock Full Question Bank
Get access to all 17 Cryptography Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.