Applied Cryptography and Key Management Questions
Selecting and applying cryptographic primitives correctly: symmetric and asymmetric encryption, hashing, digital signatures, key derivation, secure random number generation, and public key infrastructure. Covers key lifecycle management, key exchange and distribution, choosing appropriate algorithms for a given constraint set including resource-constrained environments, and the forward-looking side of algorithm lifecycle: cryptographic agility and algorithm-migration strategy, forward secrecy, and the post-quantum cryptography transition and planning upgrades without breaking existing data or interoperability. The applied-crypto engineering layer, distinct from compliance-driven crypto standards.
Why can't you just hash a password with SHA-256 and call it done? Walk through what a key derivation function actually is (how it differs from a plain hash or a PRF) and compare PBKDF2, bcrypt, scrypt, and Argon2 as password-storage choices: the core mechanism each relies on (iteration count vs memory hardness), typical parameter knobs, and strengths/weaknesses. Then explain how the calculus changes when you're deriving a session key from an already-random shared secret instead of a low-entropy password, and where HKDF fits that case.
Sample Answer
Direct answer
A plain hash like SHA-256 is deliberately fast and has no built-in salt, so two users with the same password get identical hashes, and an attacker holding stolen hashes can try billions of guesses per second on off-the-shelf hardware. A key derivation function (KDF) exists to do the opposite: turn a secret (a password, or an existing high-entropy secret) into cryptographic key material through a process that is either deliberately expensive, for password-based KDFs, to slow down guessing, or that safely spreads entropy into independent-looking keys, for a secret that is already random. PBKDF2, bcrypt, scrypt, and Argon2 are password-hashing KDFs; HKDF is built for the second case.
KDF vs hash vs pseudorandom function (PRF)
- A cryptographic hash function (SHA-256) maps arbitrary input to a fixed-size output, deterministically and quickly, with no secret key. Good for integrity checks; bad for password storage precisely because it is fast.
- A pseudorandom function (PRF), commonly HMAC-SHA256, is a keyed function whose output is computationally indistinguishable from a truly random function's output to anyone without the key. PRFs are building blocks, not complete answers: they say nothing about work factor or memory cost.
- A KDF composes a hash or PRF with an explicit strategy (iteration, memory hardness, or entropy extraction/expansion) to solve "slow this down for guessing resistance" or "expand this safely."
Comparing the four password-hashing KDFs
| Algorithm | Core mechanism | Typical knobs | Strength | Weakness |
|---|---|---|---|---|
| PBKDF2 | Iterates a PRF (usually HMAC-SHA256) c times | Iteration count only | Simple, standardized in NIST SP 800-132, widely available in FIPS-validated libraries | Needs almost no memory, so custom hardware (ASICs, GPUs) can run millions of cheap parallel guesses |
| bcrypt | Repeated Blowfish key schedule | Cost factor (rounds = 2^cost) | Decades of scrutiny, resists naive GPU ports better than PBKDF2 | Fixed, small (roughly 4 KB) memory footprint; silently truncates passwords past 72 bytes |
| scrypt | Sequential memory-hard mixing | Memory/CPU cost N, block size r, parallelism p | True memory hardness (RFC 7914), first widely deployed | More complex to tune correctly; mostly superseded by Argon2 for new work |
| Argon2 | Memory-hard, tunable in three dimensions | Memory m, iterations t, parallelism p | Winner of the 2015 Password Hashing Competition, documented in RFC 9106, three variants for different threat models | Needs careful parameter tuning per target hardware |
The line that matters most: PBKDF2 and bcrypt scale cost mainly by adding sequential compute, which cheap, highly parallel hardware absorbs well. scrypt and Argon2 scale cost by requiring a large memory footprint per guess, and memory is expensive to duplicate across thousands of parallel attempts, which is why they resist GPU/ASIC cracking far better per unit of legitimate server CPU time spent.
When the input is already random: HKDF
Once you have a high-entropy shared secret, for example the output of an ECDHE key exchange, you have no guessing problem to defend against; you have an "I need several independent-looking keys from one secret" problem (a separate encryption key and a MAC key, or per-direction keys). Deliberately slowing that down the way PBKDF2/Argon2 do would only add latency with no security benefit, since there is no offline dictionary attack to resist. HMAC-based key derivation function (HKDF, RFC 5869) is built for exactly this: an "extract" step concentrates the input's entropy into a fixed-length pseudorandom key using HMAC, then an "expand" step derives as much labeled output keying material as needed, bound to a context string so an encryption key and a MAC key derived from the same secret cannot be confused with each other.
Worked example
Deriving 64 bytes of key material (a 32-byte AES key plus a 32-byte HMAC key) from a 32-byte ECDHE shared secret with HKDF-SHA256 is one Expand call producing two concatenated 32-byte outputs. HKDF-SHA256 has a hard ceiling from its own specification: the expand step can produce at most 255 times the hash length, which for SHA-256 (32-byte output) is
255×32=8160 bytes
Deriving a handful of session keys never comes close to that ceiling, which underlines the point: HKDF is built to expand cheaply, not to resist guessing.
Trade-offs and pitfalls
- Argon2id is the current default recommendation over PBKDF2 or bcrypt precisely because of the memory-hardness gap above.
- Using a password KDF where you actually have a random shared secret wastes CPU and adds latency for no security gain; using HKDF where you actually have a low-entropy password gives an attacker no additional cost at all, since HKDF has no configurable work factor.
- Never build a KDF out of raw hash calls yourself; use the vetted PBKDF2/bcrypt/scrypt/Argon2/HKDF implementation in your language's standard cryptography library.
Describe practical techniques for protecting key material in the memory of a running system: zeroization, mlock/VirtualLock to prevent paging to disk, guard pages, dedicated secure-memory APIs, and hardware secure enclaves. What are the limitations of these techniques on modern operating systems and managed runtimes like the JVM or CPython?
Sample Answer
Direct answer: Key material sitting in ordinary process memory can be paged to disk (swap), captured in a crash dump, or read by another process or a privileged attacker while still in RAM. Protecting it means marking those pages non-swappable, actively overwriting ("zeroizing") them the instant the key is no longer needed rather than waiting for garbage collection, and, where available, keeping the key inside a hardware-isolated boundary so it never exists in general process memory at all. On managed runtimes like the JVM or CPython, none of these are fully achievable, because the runtime itself controls memory layout and timing in ways application code cannot override.
Zeroization. Explicitly overwrite the memory holding a key as soon as it is no longer needed. The practical difficulty is that compilers and runtimes can legally eliminate a "dead store," a write to memory that is never read afterward looks, from the optimizer's point of view, like it can simply be removed. Real implementations use a function specifically documented to survive optimization (explicit_bzero on Linux/BSD, memset_s from the C11 Annex K extension on other platforms) rather than a plain memset call.
mlock/VirtualLock. mlock() on POSIX systems or VirtualLock() on Windows pins a memory region so the operating system never writes it to swap or the page file. This matters because swapped-out key material can persist on disk indefinitely, surviving process exit or even a reboot, somewhere a "wipe it from RAM" strategy provides zero protection. The limitation: these calls typically require an elevated privilege or a raised resource limit, and they only protect against swapping, not against a privileged attacker reading live RAM or a core dump.
Guard pages. Allocate an inaccessible memory page immediately before or after the sensitive region, so an out-of-bounds read or write elsewhere in the process hits the guard page and crashes immediately and loudly, rather than silently reading adjacent key material. This converts a silent information leak into a visible crash, a trade of availability for confidentiality that is usually the right one for key material specifically.
Dedicated secure-memory APIs. Libraries like libsodium bundle guard pages, memory locking, and guaranteed zeroing on free behind a single call (sodium_malloc). Prefer these over hand-rolling the individual primitives, since getting all the interacting details right (page alignment, the dead-store problem) is easy to get subtly wrong.
Hardware secure enclaves. The strongest option: the key material never exists in the general process's addressable memory at all, only inside a hardware-isolated boundary that even a compromised operating system kernel cannot read directly. The trade-off is a much narrower programming interface (typically you can only ask the enclave to perform specific operations, not general-purpose computation on the key), and real limitations exist even here, several hardware enclave technologies have had published side-channel attacks over the years, so this is a strong control, not an absolute guarantee.
Limitations on managed runtimes (JVM, CPython). Garbage-collected runtimes routinely copy objects during collection, for example a compacting collector moves live objects to consolidate free space, so even if application code zeroizes the byte array it was handed, one or more copies of the original bytes may already be scattered elsewhere in memory from a prior garbage collection cycle, copies the code has no handle to and cannot zeroize. This is a structural limitation of the runtime, not a bug that can be coded around within the language. Both the JVM and CPython also make memory pinning awkward, since neither exposes direct control over which native pages back a given object. Practical mitigations in both ecosystems push sensitive key handling toward native, non-moving memory: the JVM's Foreign Function and Memory API can allocate raw, off-heap memory outside the garbage collector's control, and Python code doing serious key handling typically delegates the sensitive cryptographic operations to a native extension (such as the OpenSSL bindings in the cryptography package) so the plaintext key mostly lives in C-managed memory rather than a Python bytes object the interpreter may have copied or interned.
Worked example (a real, compiled demonstration of the dead-store problem). Compiling this C function, which loads a "secret" into a stack buffer, uses it, and then tries to wipe it with a plain memset:
void process_secret(void) {
char key[32];
for (int i = 0; i < 32; i++) key[i] = 'A' + (i % 26);
int sum = 0;
for (int i = 0; i < 32; i++) sum += key[i];
printf("checksum=%d\n", sum);
memset(key, 0, sizeof(key)); /* meant to wipe the secret before returning */
}
At -O0 (no optimization), running the compiled binary prints:
checksum=2420
and the compiled assembly shows the zeroing store still present right after the printf call, as expected.
At -O2, clang determines the entire computation is a compile-time constant (nothing here is treated as genuinely secret, unpredictable input, so the optimizer is free to fold it), and the compiled assembly for the whole function collapses to just loading the constant 2420 and calling printf. The array, the loops, and the memset are all eliminated, not because the compiler is targeting the zeroization specifically, but because it cannot tell the difference between "this is a secret that must be materialized and then wiped" and "this is dead code with no observable effect." Recompiling the same logic with memset_s instead of memset keeps an explicit call memset_s in the -O2 output, because the C11 Annex K specification for that function mandates the call is never optimized away, exactly the guarantee memset does not provide.
Trade-offs and pitfalls. The example above is the pitfall, in miniature: a plain memset-based zeroization routine can be silently and completely removed by the optimizer with no warning, and the bug is invisible from reading the source code alone. Always use a zeroization primitive the platform explicitly documents as immune to this optimization, and treat any custom zeroization helper you write yourself as unverified until you have actually inspected the compiled output.
Summarize NIST's post-quantum cryptography standardization effort: which algorithm families and specific algorithms has NIST selected for KEMs and signatures, and how should an organization use that activity when planning a migration? Explain why production deployments favor hybrid classical-plus-PQ patterns (parallel key encapsulation, dual signatures) rather than switching outright, and the main pitfalls to avoid when implementing a hybrid.
Sample Answer
Direct answer: NIST finalized three post-quantum cryptography (PQC) standards in August 2024: FIPS 203 (ML-KEM, a key encapsulation mechanism, meaning an algorithm two parties use to agree on a shared secret), FIPS 204 (ML-DSA, the primary general-purpose signature scheme), and FIPS 205 (SLH-DSA, a hash-based signature kept as a conservative backup built on different mathematical assumptions). A fourth, FIPS 206 (FN-DSA, offering smaller signatures), is still in draft as of 2026, not yet finalized, so organizations should plan around the first three now and treat FN-DSA as directional rather than a firm dependency. Production deployments favor hybrid classical-plus-PQ constructions rather than switching outright, because the new lattice-based math has roughly a decade of serious cryptanalysis behind it versus elliptic-curve cryptography's three-plus decades, and a properly combined hybrid is only as weak as its strongest component.
What this means for migration planning. The "we're waiting for NIST to decide" excuse is over for key exchange and general-purpose signing. Cryptographic inventory work (knowing exactly where every algorithm is actually used) and algorithm-agility groundwork, versioned ciphertext and certificate formats, a documented decision record for each choice, should start now, targeting these three finalized families. Do not architect a hard dependency on FN-DSA yet since its parameters and even its finalization timeline could still shift.
Why hybrid, not a straight switch
- Defense against implementation immaturity of the new math. If a properly designed hybrid combiner is used, breaking the hybrid requires breaking both the classical and the post-quantum component. If lattice cryptography turns out to have an unforeseen flaw, the classical component alone still protects you; the reverse is also true against a future quantum computer.
- Compliance transition. Many existing compliance regimes still mandate specific classical, FIPS-approved algorithms; a hybrid can satisfy that mandate while also adding post-quantum protection, whereas a pure PQ-only deployment might not yet satisfy an unmodified compliance requirement.
- Interoperability during rollout. Not every peer supports the new algorithms yet; a hybrid degrades gracefully where a hard switch would simply fail to connect.
Main pitfalls implementing a hybrid
- Getting the combiner wrong. Naively concatenating two shared secrets is not automatically secure on its own; the combiner needs to be a proven construction. For TLS 1.3, RFC 9954 specifies exactly how the classical and post-quantum shared secrets are combined via the existing key schedule's HKDF step; do not invent your own combining function.
- Bandwidth and fragmentation. ML-KEM-768 alone adds roughly 2.2 KB of key-exchange material to a handshake (see the worked arithmetic below); combined with a hybrid signature too, a TLS ClientHello or ServerHello can exceed a single network packet's typical size and trigger fragmentation issues on middleboxes that mishandle unusually large TLS handshake records, a real, documented rollout pain point.
- Downgrade risk. If negotiation is not done carefully, an active attacker could try to strip the post-quantum component and force a classical-only connection. TLS 1.3's transcript binding already protects against this when the hybrid group is negotiated through the standard mechanism; do not build a separate, unprotected negotiation path.
- Assuming "hybrid" automatically doubles security. A coding bug, such as reusing the same randomness source for both key-generation operations, can quietly defeat the whole point.
Worked example. FIPS 204's ML-DSA-65 signature is roughly 3,309 bytes; a classical ECDSA P-256 signature is roughly 70 bytes (DER-encoded). Combining them for a hybrid signature:
3,309 + 70 = 3,379 bytes combined signature
3,379 / 70 = ~48.3x the size of the classical signature alone
That is not a rounding error, it is nearly a fiftyfold increase, which is exactly why hybrid signatures (as opposed to hybrid key exchange, which adds a comparatively modest amount of overhead) are the harder piece to fit into size-constrained protocols like DNSSEC or small IoT payloads.
Trade-offs and pitfalls, restated. The single most important framing for planning purposes: symmetric algorithms (AES, HMAC) do not need to change for post-quantum readiness, only key exchange and signatures do, since Shor's algorithm targets the number-theoretic structure those two rely on, not symmetric primitives.
Design a migration strategy to move a user database from PBKDF2 to Argon2id without forcing a mass password reset. Cover the schema changes needed to version hashes, the authentication-flow change that detects an old hash and re-hashes on successful login, options for migrating accounts that never log in again, and the metrics you'd watch to confirm the migration is succeeding.
Sample Answer
Direct answer
Store the algorithm and its parameters alongside every hash so old (PBKDF2) and new (Argon2id) records are distinguishable, verify a login against whichever algorithm the stored hash says it used, and silently re-hash with Argon2id the moment a login succeeds, since a successful login is the one moment the plaintext password is legitimately available in memory to spend the cost of the new, stronger hash on.
Schema changes to version hashes
Store a self-describing hash string rather than a bare hash value: the algorithm identifier, its parameters, the salt, and the derived hash, all together, for example the widely used PHC string format ($argon2id$v=19$m=65536,t=3,p=4$<salt>$<hash>) versus a PBKDF2 equivalent naming its iteration count. This avoids a separate "algorithm version" column that could drift out of sync with the actual stored value; parsing the stored string alone tells you everything needed to verify it.
Authentication-flow change: detect old hash, re-hash on success
- On login, parse the stored hash string to determine which algorithm and parameters it used.
- Verify the submitted password against that algorithm (PBKDF2 for old records, Argon2id for already-migrated ones).
- If verification succeeds and the stored hash is not already the current target algorithm/parameters, immediately derive a new Argon2id hash from the plaintext password, available only during this one request, and overwrite the stored record with the new hash string.
- If verification fails, leave the stored record untouched; a failed login gives no legitimate opportunity to touch it, and doing so would be actively wrong.
Accounts that never log in again
Some accounts never trigger the natural re-hash-on-login path because they simply stop logging in. Options, usually combined: leave them on the old algorithm indefinitely, accepting the residual risk since these tend to also be the least active, lowest-value targets, with a monitoring dashboard tracking what fraction of the user base remains un-migrated over time; force a password reset for accounts that have not logged in, and therefore not migrated, after a defined long window, one of the rare cases where forcing a reset for a subset of users, rather than the whole base, is reasonable, trading a small, bounded friction against an indefinite tail of weakly hashed accounts; or, for genuinely dormant accounts past a retention policy, consider whether they should simply be deactivated rather than migrated at all.
Metrics to confirm the migration is succeeding
- Percentage of the user base still on the old algorithm, trending down over time as logins naturally trigger re-hashing; a stalled or flat trend signals the natural mechanism alone will not finish in a reasonable timeframe and the forced-reset path for dormant accounts needs to trigger sooner or more aggressively.
- Error rate on the re-hash step itself, distinct from the login verification step, since a bug here could silently fail to upgrade accounts while users still log in successfully, hiding the problem from ordinary login-success monitoring.
- Latency impact on logins that include a re-hash, since Argon2id's memory-hard computation adds real cost to that one request; this should show up as a small, bounded latency bump specifically on migrating logins, not a general regression across all logins.
Trade-offs and pitfalls
Never derive the new hash from anything except the plaintext password already available during a successful verification; there is no way to re-hash an existing PBKDF2 output into an Argon2id one without the original plaintext, since these are one-way functions by design. If the login endpoint is heavily rate-limited or load-shed under peak traffic, remember the re-hash step adds real, non-trivial compute cost, that is the point of Argon2id, so capacity-plan for a transition period where a meaningful fraction of logins carry that extra cost, not just the eventual steady state once migration is complete.
Compare the security assurances of TPMs, dedicated HSMs, and cloud-provider KMS hardware-backed stores: tamper resistance, certification levels (FIPS/Common Criteria), remote-attestation strength, supply-chain concerns, and firmware-update posture. How do these differences actually drive key-lifecycle and policy decisions in a highly regulated environment?
Sample Answer
Direct answer
A Trusted Platform Module (TPM), a dedicated Hardware Security Module (HSM), and a cloud provider's
hardware-backed Key Management Service (KMS) sit on a spectrum from cheap-and-general to
expensive-and-specialized, and the right choice for a given key depends less on which is "most
secure" in the abstract and more on which specific property (tamper response, certified physical
custody, control over firmware-patch timing) a regulation or threat model actually requires.
Structured elaboration
| Property | TPM | Dedicated HSM | Cloud KMS (hardware-backed) |
|---|---|---|---|
| Purpose | Platform/boot integrity, modest key sealing | High-throughput crypto operations for high-value keys | Elastic, managed crypto operations at cloud scale |
| Tamper resistance | Basic: resists casual extraction, not a well-funded lab | Strong: many models are tamper-responsive, zeroizing keys on detected intrusion | Strong hardware, but you never see or audit it directly, you trust the provider's operation of it |
| Certification | Commonly FIPS 140-2 or 140-3 validated at a modest level | Commonly FIPS 140-3 Level 3, sometimes Level 4, or Common Criteria evaluated | Commonly FIPS 140-3 validated hardware underneath, inherited by the managed service |
| Remote attestation | Its signature strength: signed "quotes" over Platform Configuration Registers (PCRs) proving what firmware/bootloader ran | Vendor-specific attestation and audit mechanisms, less standardized than TPM quotes | Largely "trust the provider's compliance reports" (SOC 2, ISO 27001) rather than a quote you personally verify |
| Supply chain | Fixed part of the motherboard from a small set of vendors, generally traceable | Vendor-locked appliance, a narrower but auditable vendor relationship | Risk shifts to trusting the cloud provider's hardware supply chain and staff entirely |
| Firmware updates | Rare, vendor-controlled, you have little say in timing | You control update timing, which matters for signing ceremonies and change freezes | Provider-controlled, patched quickly, but you cannot delay a patch to align with your own freeze window |
How these differences drive key-lifecycle and policy decisions.
- If a regulation requires provable physical custody or single-tenant hardware, a shared multi-tenant
cloud KMS may not satisfy it on its own; many cloud providers offer a distinct "dedicated HSM" tier
(single-tenant hardware, still cloud-hosted) specifically to close that gap. - Root Certificate Authority (CA) keys and other long-lived signing keys usually justify a dedicated
HSM: you want tamper-response, a higher validation level, and control over exactly when firmware
changes so an unexpected forced patch cannot interrupt a signing ceremony. - TPMs are the right tool for platform and boot-integrity attestation (measured boot, sealing a
disk-encryption key to a known-good boot state), not for general organizational secret management;
treating a TPM as a substitute for a proper key-management infrastructure is a common design
mistake. - A cloud KMS's fast, provider-controlled patch cadence reduces your exposure window to a disclosed
firmware vulnerability, but you lose the ability to hold a patch back to align with your own change
freeze, which is a real, explicit trade-off, not a strictly better or worse choice.
Worked example
A payment company must keep card-data-encryption keys inside hardware that satisfies its Payment
Card Industry (PCI) obligations. Using only a shared, multi-tenant cloud KMS, an auditor would
typically expect evidence of dedicated, non-shared hardware and controlled physical access, which a
standard multi-tenant offering cannot provide by design. The company instead chooses a cloud
provider's dedicated HSM tier: single-tenant, FIPS 140-3 validated hardware, with customer-controlled
key material and access policy, satisfying the "not shared with other tenants" requirement without
the company operating its own physical data center.
Trade-offs & pitfalls
- Do not treat "FIPS 140-3 validated" as a single blanket property of a whole product. It describes a
specific certified module and boundary at a specific validation level (1 through 4); always state
the module and level rather than a bare "certified" claim. - Common Criteria evaluation levels describe how rigorously a product was tested and documented, not
raw security strength; a higher evaluation level means "more thoroughly evaluated," not
automatically "harder to break." - Remote attestation on a TPM proves what code ran at boot, not that the code is free of
vulnerabilities; do not conflate integrity attestation with a general security guarantee.
Unlock Full Question Bank
Get access to all Applied Cryptography and Key Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.