Cryptography Fundamentals Questions
Core concepts and vocabulary of cryptography: confidentiality, integrity, authentication, and non-repudiation; the difference between symmetric and asymmetric primitives; and how standard algorithms, libraries, and protocols fit together. Covers threat models, common standards, and applying primitives and cryptographic libraries correctly to real-world security problems. The entry point for the cryptography track.
Compare block cipher modes of operation such as CBC and GCM. In a system that requires encryption for disks (at-rest logs) and for transport (TLS-like traffic), which mode(s) would you choose for each use case and why? Discuss IV/nonce requirements and integrity guarantees.
Sample Answer
Direct answer
For transport-like traffic (a live connection between two parties), use Galois/Counter Mode (GCM): it is authenticated, single-pass, and parallelizable, which matches a stream of packets where you want low latency and can generate a fresh nonce per message easily. For data at rest (logs sitting on disk with no live connection to negotiate per-message nonces), Cipher Block Chaining (CBC) is the classic textbook answer but is the wrong choice today for the same reason it always was: it provides no built-in integrity check. The honest, current recommendation for both cases is an authenticated mode, GCM (or an alternative authenticated construction) for both, with the at-rest case leaning on deliberate nonce management rather than live negotiation.
Comparing the modes
- CBC encrypts each plaintext block XORed with the previous ciphertext block, using an initialization vector (IV) to seed the first block. It needs the IV to be unpredictable (not necessarily secret, but not attacker-guessable in advance) and requires padding, which historically enabled padding-oracle attacks when an application leaked whether decryption padding was valid. Critically, CBC on its own gives you confidentiality but no integrity: an attacker can flip bits in ciphertext and produce predictable, attacker-influenced changes in the decrypted plaintext with no error raised.
- GCM combines counter-mode encryption with a polynomial-based authentication tag (GHASH) in one pass, so it gives you both confidentiality and integrity, and it detects any tampering by failing to verify rather than silently decrypting corrupted data. Its hard requirement is a nonce (a 96-bit value in the standard construction) that must never repeat under the same key, reuse breaks both the confidentiality guarantee (via keystream reuse) and the authentication guarantee (via leaking GCM's internal authentication key), so nonce uniqueness is non-negotiable.
Worked recommendation
- At-rest logs: use GCM, with nonces generated from a per-file or per-segment counter that is durably persisted alongside the encrypted data (so a process restart cannot repeat a value), or a per-record random 96-bit nonce if record volume per key stays well under the birthday-bound guidance for random nonces. Store the nonce and authentication tag next to the ciphertext; verification on read means corrupted or tampered log segments are detected immediately instead of silently trusted.
- Transport-like traffic: use GCM (or ChaCha20-Poly1305 where hardware AES acceleration isn't available) with a fresh, connection-scoped nonce per record, which is exactly the pattern Transport Layer Security (TLS) 1.2 and 1.3 use internally, a per-record explicit or implicit nonce derived from a sequence number.
Trade-offs and pitfalls
The single hardest constraint in both cases is the same one: nonce uniqueness under a fixed key. Choosing GCM for both use cases only pays off if the nonce-generation strategy for each is actually enforced (durable counters for files that survive restarts, connection state for transport), because a mode swap from a broken CBC deployment to a broken GCM deployment (repeating nonces) is not actually a fix, it trades one class of catastrophic failure for another.
Your service currently stores passwords with a weak or legacy hashing scheme (say, unsalted SHA-1). Design the migration to a modern KDF like Argon2id: how you pick parameters for production versus constrained clients, how you support a rolling upgrade so users aren't forced to reset immediately, how you detect and re-hash the remaining weak legacy entries over time, and what you'd monitor to catch offline cracking attempts against the old hashes.
Sample Answer
Direct answer
Migrate by upgrading each user's hash the moment you have their plaintext password in hand,
which is only at successful login, so nobody is forced to reset. Track which scheme each
record is on, tune Argon2id's memory/time cost to the device that will actually run it, sweep
up stragglers who never log back in, and watch for signs the old dump is being cracked
offline in the meantime.
Structured elaboration
1. Argon2id parameters, production vs constrained clients. Argon2id has three costs:
memory (KiB), iterations, and parallelism. OWASP's Password Storage Cheat Sheet lists a
sliding scale of roughly equal-strength options that trade memory for time, for example
around 46 MiB of memory with 1 iteration at parallelism 1 for a normal server, versus roughly
12 to 19 MiB with 2 to 3 iterations when memory is the constrained resource (a mobile client
doing local verification, or a server handling very high login concurrency where 46 MiB per
concurrent hash would exhaust RAM). Pick the highest memory cost your slowest realistic
target (server fleet under peak login load, or the weakest client device) can sustain in
under roughly half a second; more memory is what actually makes GPU/ASIC cracking expensive,
since those devices have comparatively little fast memory per core.
2. Rolling upgrade with no forced reset. Store the hash in a self-describing format that
names its own algorithm (Argon2id's standard encoded string already does this;
$2b$... for bcrypt or a raw hex digest identifies legacy SHA-1). On login, verify against
whatever scheme the stored value says it is. If verification succeeds and the scheme is not
the current one, you now hold the plaintext password for one request: immediately hash it
with Argon2id and overwrite the stored value before returning the response. Every active user
silently upgrades on their next successful login; nobody sees a reset prompt.
3. Detecting and re-hashing stragglers. Track an algo/version column per credential
row and report the shrinking count of legacy-SHA1 rows over time as a rollout metric. Users
who never log in cannot be upgraded this way, because you never see their plaintext again.
For those, set a deadline (say, after a defined inactivity window) past which you expire the
legacy hash and require a password reset via email on next login attempt, since silently
leaving weak hashes in place indefinitely defeats the migration.
4. Monitoring for offline cracking against the old hashes. You cannot watch cracking
happen inside an attacker's own hardware, so the signal has to come from what a successful
crack produces: seed a handful of unused canary accounts with known passwords hashed under
the same legacy SHA-1 scheme, and alert if those exact credentials are ever used to log in;
watch for a spike in successful logins from unfamiliar IPs/devices concentrated on
still-legacy accounts, which suggests a cracked batch is being replayed; and correlate
against breach-notification feeds (has any of these emails' passwords shown up in a public
breach corpus). None of this proves nothing is being cracked, but it turns "we hope not" into
a concrete tripwire.
Worked example
User alice has a legacy row: algo=sha1, hash=<sha1 hex digest>. She logs in with her real
password. The server sees algo=sha1, verifies the SHA-1 digest against the submitted
password, and it matches. Because verification succeeded, the server now holds alice's
plaintext password for the remainder of this one request. It immediately computes
argon2id(password, salt, m=19456, t=2, p=1), overwrites the row to
algo=argon2id, hash=<new encoded hash>, and returns the normal login response. Alice never
saw a reset prompt; her next login will verify against Argon2id instead. A dashboard query
counting WHERE algo = 'sha1' trending toward zero over weeks is the migration's own progress
metric; whatever remains after the inactivity deadline gets force-reset instead of waiting
indefinitely.
Trade-offs & pitfalls
- Skipping the self-describing hash format is the single most common mistake: without a way
to tell schemes apart per row, you cannot dispatch verification correctly during the
transition. - Setting Argon2id memory too low "to be safe on old hardware" quietly reduces it to
something closer to a fast hash, defeating the point of the migration. - Forcing every user to reset immediately is simpler to build but throws away the accounts of
anyone who does not see the email in time; the rolling upgrade exists specifically to avoid
that support and churn cost.
Give a high-level walkthrough of a TLS handshake: what happens at each step, which cryptographic primitives are used where (certificates, asymmetric key exchange, symmetric session keys, MAC/AEAD), and what properties TLS is actually trying to guarantee. What's one common way this can fail in practice?
Sample Answer
Direct answer
A TLS (Transport Layer Security) handshake lets a client and server who have never met agree
on a shared symmetric key, authenticate the server (and optionally the client), and start
exchanging data, with confidentiality, integrity, and (with modern configuration) forward
secrecy. It works by combining asymmetric key exchange and certificate-based identity checks
up front, then switching to fast symmetric encryption for the actual traffic.
Structured elaboration
sequenceDiagram
participant C as Client
participant S as Server
C->>S: ClientHello (supported versions, cipher suites, key share, random)
S->>C: ServerHello (chosen cipher suite, key share, random)
S->>C: Certificate (server public key, signed by a CA)
S->>C: Finished (proves possession of the certificate's private key)
Note over C,S: Both derive the same shared secret via ECDHE, then a session key via a KDF
C->>S: Finished (handshake integrity check)
C->>S: Application data (encrypted with the negotiated AEAD cipher)
S->>C: Application data (encrypted with the negotiated AEAD cipher)
- ClientHello / ServerHello: negotiate protocol version, cipher suite, and exchange
ephemeral key-exchange material (an elliptic-curve public value for ECDHE) plus random
nonces that feed the key derivation. - Certificate: the server presents its certificate, a public key bound to an identity and
signed by a certificate authority (CA), so the client can authenticate who it is actually
talking to (see PKI/certificate chains for how that trust is validated). - Key derivation: both sides combine the ECDHE exchange output through a KDF (key
derivation function) to produce the actual symmetric session keys, never sending the
session key itself over the wire. - Application data: all subsequent traffic is encrypted and authenticated with an AEAD
cipher (typically AES-GCM or ChaCha20-Poly1305) under the derived session key; older TLS
versions could instead pair a plain block cipher with a separate MAC (message
authentication code) computed over each record, a split that AEAD replaces with one
combined operation. - Properties TLS is guaranteeing: confidentiality (only the endpoints can read the data),
integrity/authenticity (tampering is detected), server authentication (the certificate
chain proves identity), and, with an ephemeral key-exchange suite, forward secrecy.
Worked example / common failure
The most common real-world failure is not a cryptographic break, it is trust management:
an expired certificate. Everything about the handshake's math can be flawless and the
connection will still be rejected (or a warning shown) the moment the leaf certificate's
validity window has passed, because clients check expiration as part of chain validation
before they will trust the negotiated key exchange at all. This is a frequent cause of
sudden, otherwise-unexplained outages, and is why certificate expiry monitoring is treated as
an operational, not just a security, concern.
Trade-offs & pitfalls
- TLS 1.3 (RFC 8446) simplified this flow and removed insecure options entirely (no more
static-RSA key exchange, no more weak cipher suites), which is why upgrading from TLS 1.2
is itself a meaningful security improvement, not just a version bump. - A valid, unexpired certificate chain proves identity, it does not by itself prove the
identity is trustworthy; certificate validation and business-level trust are separate
concerns.
Suppose a memory disclosure vulnerability similar to Heartbleed was discovered in an SSL/TLS library used by your fleet. Describe how you would detect whether exploitation occurred, immediate mitigations, patch rollout strategy, and post-incident tasks such as key and certificate rotation, forensic evidence collection, and customer notification.
Sample Answer
Framing
Heartbleed (CVE-2014-0160, disclosed in 2014) is the canonical example of a memory disclosure vulnerability in a Transport Layer Security (TLS) library: a bug in OpenSSL's implementation of the TLS "heartbeat" extension let an unauthenticated remote attacker read chunks of the server process's memory, with no login required and, critically, often no trace in normal logs. The response to a Heartbleed-class bug in your own fleet has to assume the worst rather than wait for proof, because the nature of the bug makes proof unusually hard to get.
Why detection is unusually hard here
The underlying flaw: a heartbeat request includes a payload and a claimed payload length, but the vulnerable code trusted the claimed length without checking it against the payload actually sent, and echoed back that many bytes starting from the payload's location in memory, silently including whatever adjacent memory happened to be there. Since the length field is a 16-bit value, the maximum claimed length is 2^16 - 1 = 65,535 bytes, roughly 64 kilobytes of arbitrary process memory per request, and an attacker can repeat the request as many times as they like, each time potentially recovering a different memory window. Because a heartbeat exchange is a normal, expected part of TLS and the request itself doesn't look malformed at the protocol level, it produces no error and, under most default logging configurations, no log entry at all. Detection cannot rely on finding evidence of exploitation, its absence is not reassuring.
Detecting whether exploitation occurred
- Check the installed OpenSSL version against the known-vulnerable range (OpenSSL 1.0.1 through 1.0.1f were affected; 1.0.1g and later fixed it) across every server, load balancer, and embedded device in the fleet, including things people forget are running TLS termination (internal load balancers, management interfaces, third-party appliances).
- Search any available packet captures or intrusion-detection signatures for heartbeat requests with an anomalous claimed payload length relative to the actual payload sent, this is the one concrete artifact that can prove exploitation if you happen to have captured it, but its absence proves nothing given how rarely this traffic gets captured in practice.
- Treat any server that was running a vulnerable version and was reachable from the internet during the vulnerability window as a presumed compromise for planning purposes, even with zero direct evidence.
Immediate mitigations
- Patch OpenSSL to a fixed version on every affected system, prioritized by internet-facing exposure, and where an immediate patch isn't possible on a given system, disable the heartbeat extension as an interim compensating control if your TLS stack supports doing so.
- Take internet-facing vulnerable endpoints out of service or behind a patched proxy rather than leaving them exposed while a fleet-wide patch rolls out.
Patch rollout strategy
Patch by exposure, not by convenience: internet-facing systems first, then internal systems that terminate TLS for anything sensitive, then everything else. A patch alone does not undo prior exposure, so rollout speed matters less than what happens next.
Post-incident tasks
- Key and certificate rotation. Because the bug could leak a server's private key directly from memory, and there is no reliable way to prove it didn't, every private key that was ever loaded into a vulnerable process's memory should be treated as compromised: revoke the old certificates and reissue new ones with fresh keys, don't just reissue a certificate for the same key.
- Credential rotation. Session tokens, passwords, and any other secrets that could plausibly have passed through a vulnerable process's memory (anything handled during requests to that server in the exposure window) should be treated the same way, force credential resets rather than assuming they were untouched.
- Forensic evidence collection. Preserve whatever logs, packet captures, and system snapshots exist from the exposure window before they age out or get overwritten, even incomplete evidence is useful for later analysis and for any compliance or legal review of the incident.
- Customer notification. Because exploitation cannot be conclusively ruled out or confirmed from typical logging, notification and disclosure decisions should be made on the assumption of potential compromise, honestly communicating that certainty isn't available, rather than waiting for definitive proof that will likely never come.
Trade-offs and pitfalls
The single biggest mistake in a Heartbleed-class incident is treating "we found no evidence of exploitation" as equivalent to "we confirmed no exploitation occurred." Given how the bug works, those are not the same statement, and a post-mortem that conflates them tends to under-rotate keys and credentials, leaving a genuinely compromised key in service simply because nobody could prove it was used.
Compare PBKDF2, bcrypt, scrypt, and Argon2 at a high level. For each describe its primary design goals, whether it is CPU-bound or memory-hard, how resistant it is to GPU/ASIC acceleration, and any known side-channel concerns. For a greenfield web service today, state which you would choose by default and justify that choice in terms of security and deployability.
Sample Answer
Direct answer
All four are key derivation functions (KDFs) built specifically to make password guessing slow, unlike a plain fast hash like SHA-256, which an attacker with commodity graphics-processing-unit (GPU) hardware can evaluate billions of times a second. They differ in what resource they force an attacker to spend: pure computation time (PBKDF2), or computation time plus memory (bcrypt, scrypt, Argon2). For a new, general-purpose web service today, Argon2id is the right default.
Comparing the four
| Function | Primary resource cost | Memory-hard? | GPU/ASIC resistance | Notable side-channel concern |
|---|---|---|---|---|
| PBKDF2 (Password-Based Key Derivation Function 2) | CPU time only (many rounds of an underlying hash-based message authentication code) | No | Weak: its small, fixed memory footprint makes it comparatively cheap to parallelize on GPUs and custom application-specific integrated circuit (ASIC) hardware | None specific to the algorithm itself |
| bcrypt | CPU time, with a small, fixed memory requirement from its Blowfish-based design | Modest, not true memory-hardness by modern standards | Better than PBKDF2, still meaningfully weaker against GPU/ASIC attackers than scrypt or Argon2 | Its internal state fits comfortably in a GPU's cache, reducing the memory-bandwidth advantage a defender would want |
| scrypt | CPU time and a large, tunable memory requirement | Yes, this was its original selling point | Strong, since large memory requirements are expensive to replicate at scale on GPUs and especially on ASICs | Its data-dependent memory access pattern can leak information via cache-timing side channels |
| Argon2 (Argon2id variant) | CPU time and tunable memory, explicitly designed as the winner of the 2015 Password Hashing Competition | Yes, and tunable independently of CPU cost | Strong, comparable to or better than scrypt, with more flexible tuning of the memory/time/parallelism trade-off | Argon2id specifically mitigates the side-channel weakness of Argon2d by using a data-independent memory access pattern for its first pass |
Worked recommendation
For a greenfield web service today, default to Argon2id with parameters sized to your actual server hardware (a common practical target is tuning memory and iteration count so a single hash takes somewhere in the range of a few hundred milliseconds on your production hardware, adjusted for how much concurrent login load you need to serve). Argon2id is chosen specifically because it is both memory-hard (expensive to parallelize on GPUs and application-specific hardware) and resistant to the cache-timing side-channel weakness of the purely data-dependent Argon2d variant, while remaining widely supported in current cryptography libraries and formalized in a published standard (RFC 9106). Where a platform or compliance requirement mandates a National Institute of Standards and Technology (NIST) approved algorithm and Argon2 isn't an option, PBKDF2 with a high iteration count is the fallback, but it should not be the first choice when Argon2id is available, since it offers no meaningful memory-hardness at all.
Trade-offs and pitfalls
Picking a KDF is necessary but not sufficient: tuning it too weakly (low iteration count, low memory) to keep login latency low defeats the purpose, while tuning it too aggressively can create a denial-of-service risk on a login endpoint under load. Every one of these functions must also be used with a unique, randomly generated salt per user, reusing a salt (or omitting one) lets an attacker precompute or share cracking work across every account that used the same value, regardless of which of these four algorithms is chosen.
Unlock Full Question Bank
Get access to all 26 Cryptography Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.