Cryptography Fundamentals Questions
Core concepts and vocabulary of cryptography: confidentiality, integrity, authentication, and non-repudiation; the difference between symmetric and asymmetric primitives; and how standard algorithms, libraries, and protocols fit together. Covers threat models, common standards, and applying primitives and cryptographic libraries correctly to real-world security problems. The entry point for the cryptography track.
Describe Public Key Infrastructure (PKI) fundamentals: what a certificate is, what a Certificate Authority (CA) does, certificate chains and trust anchors, and a typical use of certificates in TLS. Keep the explanation high-level but include how trust is established and how certificate expiration affects secure channels.
Sample Answer
Direct answer
Public Key Infrastructure (PKI) is the system of certificates, certificate authorities, and
trust chains that lets a stranger's public key be trusted without meeting them first. A
certificate authority (CA) signs a certificate binding a public key to an identity; anyone who
already trusts that CA can then trust the certificate, and by extension the key inside it.
Structured elaboration
flowchart TD
Root["Root CA (trust anchor, pre-installed in OS/browser)"]
Inter["Intermediate CA (signed by Root)"]
Leaf["Server certificate (signed by Intermediate)"]
Root -->|signs| Inter
Inter -->|signs| Leaf
Leaf -->|presented during TLS handshake| Client["Client validates the whole chain"]
- Certificate: a data structure containing a public key, an identity (like a domain
name), a validity window, and a signature from whoever issued it. - Certificate authority (CA): a party trusted to verify identity claims before signing a
certificate. Trust in a CA is not proven cryptographically from scratch each time, it comes
from the CA's root certificate being pre-installed as a trust anchor in operating systems
and browsers. - Certificate chains: in practice, a root CA rarely signs end-entity (leaf) certificates
directly; it signs intermediate CAs, which sign the leaf certificates servers actually
present. Validating a certificate means walking this chain, checking each signature, until
reaching a root the client already trusts. - Use in TLS: during the handshake, the server presents its certificate (and usually the
intermediate chain); the client verifies the signature chain up to a trusted root, checks
the domain name matches, and checks the validity window, before it will trust the public
key for the key exchange.
Worked example
A browser connecting to example.com receives a leaf certificate for example.com signed by
"Intermediate CA X," which is itself signed by "Root CA R." The browser already has Root CA
R's certificate baked into its trust store. It verifies Intermediate CA X's signature using
Root CA R's public key, then verifies the leaf certificate's signature using Intermediate CA
X's public key. Only after both signatures check out, the domain name matches, and the
current date falls inside the leaf certificate's validity window does the browser proceed to
trust that public key for the handshake.
Trade-offs & pitfalls
- Certificate expiration: every certificate in the chain has a validity window; if the
leaf certificate's window ends and it is not renewed, clients will refuse the connection (or
warn loudly) even though nothing about the underlying keys has actually changed. This is a
frequent, entirely avoidable cause of production outages, which is why certificate expiry is
tracked as an operational alert, not just a one-time setup step. - Self-signed certificates skip the CA entirely (the certificate signs itself), which is fine
for internal testing but provides no independent identity verification for anyone who
hasn't been told in advance to trust that specific certificate.
Explain the fundamental differences between symmetric and asymmetric encryption. Name real algorithms in each category, describe typical enterprise use cases for each (data at rest, key exchange, code signing, TLS), and give one concrete advantage and one limitation of each approach.
Sample Answer
Direct answer
Symmetric encryption uses one shared secret key for both encrypting and decrypting; it is
fast but requires that key to somehow reach every party who needs it. Asymmetric (public-key)
encryption uses a mathematically linked key pair, a public key anyone can have and a private
key only the owner holds, which solves the distribution problem but at a real performance
cost.
Structured elaboration
- Symmetric algorithms: AES (Advanced Encryption Standard), ChaCha20. Typical enterprise
use: encrypting data at rest (a database, a disk volume) and encrypting bulk traffic once a
session key exists.- Advantage: very fast, cheap enough to encrypt gigabytes of data with negligible overhead.
- Limitation: both sides need the same key beforehand; there is no built-in mechanism for
two parties who have never met to agree on one safely over an open channel.
- Asymmetric algorithms: RSA, ECDSA/ECDH (elliptic curve variants). Typical enterprise
use: key exchange (establishing a shared secret over an untrusted network), and code/document
signing (proving who published something, and that it was not altered).- Advantage: solves key distribution; the public key can be published openly and only the
matching private key can decrypt or sign. - Limitation: orders of magnitude slower than symmetric ciphers and produces much larger
ciphertext relative to the data size, which makes it impractical for encrypting bulk data
directly.
- Advantage: solves key distribution; the public key can be published openly and only the
Worked example
TLS (the protocol behind HTTPS) is the everyday case: a browser and a server have never
shared a secret in advance, so they use asymmetric key exchange (ECDHE, backed by the
server's certificate) to agree on a fresh symmetric session key, then switch to a symmetric
AEAD cipher like AES-GCM for the actual page data. Neither algorithm family alone does the
whole job well: symmetric alone cannot solve "how do two strangers agree on a key," and
asymmetric alone is too slow to encrypt a multi-megabyte response.
Trade-offs & pitfalls
- Almost no production system uses pure asymmetric encryption for bulk data; if you see it
proposed, that is usually a design smell. - Symmetric key management (rotating and distributing shared secrets safely) is its own hard
problem and does not disappear just because the algorithm is fast; it is why systems like
KMS (key management service) and PKI (public key infrastructure) exist.
Explain the core security properties a cryptographic hash function must have (preimage resistance, second-preimage resistance, collision resistance), name a couple of common algorithms, and give one example each of where a hash is the right tool (integrity checks, content-addressing) and where it is not (storing passwords without salting and stretching).
Sample Answer
Direct answer
A cryptographic hash function must be preimage resistant (given an output, you cannot find
an input that produces it), second-preimage resistant (given one input, you cannot find a
different input with the same output), and collision resistant (you cannot find any two
distinct inputs that share an output at all). Common algorithms today are SHA-256 and SHA-3;
MD5 and SHA-1 are broken for collision resistance and should not be used where that property
matters.
Structured elaboration
- Preimage resistance: given
H(m) = h, it should be computationally infeasible to find
anymthat producesh. This is what makes a hash "one-way." - Second-preimage resistance: given a specific
m1, it should be infeasible to find a
differentm2withH(m1) = H(m2). - Collision resistance: it should be infeasible to find any pair
m1 != m2with
H(m1) = H(m2), without fixing either input in advance. This is the property MD5 and SHA-1
have both practically broken (real collision attacks have been demonstrated for both). - Where a hash is the right tool: integrity checks (comparing a downloaded file's hash
against a published one to detect corruption or tampering), content-addressing (Git commit
hashes, package registries keying content by its digest so identical content always maps to
the same address), and as the digest step inside a digital signature: signing algorithms
sign a fixed-size hash of a message rather than the message itself, since asymmetric
operations are expensive on large inputs. - Where it is not the right tool: storing passwords with a bare, unsalted hash. Hash
functions are deliberately fast, which is exactly wrong for password storage: an attacker
with a stolen hash dump can try billions of guesses per second on commodity GPUs. Password
storage needs a slow, memory-hard key derivation function (Argon2id, scrypt, bcrypt), not a
general-purpose hash used directly.
Worked example
To see why collision resistance is the strongest property, compare a real hash to a toy,
clearly-broken one: H(x) = x mod 97. Take m1 = 5 and m2 = 102. 102 mod 97 = 5, so
H(5) = H(102) = 5: a collision found by inspection, with no search at all. That is exactly
what "not collision resistant" looks like. A real hash like SHA-256 produces a 256-bit output,
so the best known way to find any colliding pair is a generic birthday-bound search of roughly
2^128 attempts (the square root of 2^256 possible outputs), which is computationally out
of reach with current or foreseeable hardware. The mod-97 toy fails instantly because its
output space is only 97 values; a cryptographic hash needs both a large output space and no
structural shortcut into it.
Trade-offs & pitfalls
- Collision resistance and preimage resistance are different properties; a function can lose
one without immediately losing the other, which is why "MD5 has collisions" does not
automatically mean "MD5 leaks passwords," but it does mean you should not trust MD5 for
anything where an attacker choosing two colliding inputs would matter (like forging a signed
document). - "Hashing" and "encryption" are often confused: hashing is one-way and not meant to be
reversed at all, encryption is two-way and meant to be reversed with the right key. A hash
cannot be "decrypted" back to the original input, by design.
Explain forward secrecy (and perfect forward secrecy): what does it protect against, and what doesn't it protect against (e.g. server long-term key compromise vs a single session key's exposure)? Describe how ephemeral Diffie-Hellman (ECDHE) achieves it in TLS, and what a server operator must configure correctly for it to actually apply.
Sample Answer
Direct answer
Forward secrecy (also called perfect forward secrecy, PFS) means that if an attacker later
steals a server's long-term private key, they still cannot decrypt traffic they recorded in
the past, because each past session used its own key that was never derivable from the
long-term key in the first place. It does not protect a session whose own key is compromised
while that session is still relevant, or protect anything on either endpoint after
decryption.
Structured elaboration
- What it protects against: mass retroactive decryption. An attacker who passively
records encrypted traffic for years, then later compromises the server's long-term private
key (via a breach, subpoena, or key-extraction bug), gains nothing from the recording: the
long-term key was only ever used to authenticate, not to directly encrypt. - What it does not protect against: compromise of a specific session's own ephemeral key
or the derived session key while that session's data is still sensitive; endpoint
compromise (malware reading plaintext directly on a client or server); and metadata (who
talked to whom, when, how much data), which forward secrecy says nothing about. - How ECDHE achieves it in TLS: ECDHE (elliptic-curve Diffie-Hellman, ephemeral) means
both sides generate a brand-new elliptic-curve key pair for every handshake, exchange the
public halves, and derive a shared secret via elliptic-curve Diffie-Hellman. The server's
long-term certificate key is only used to sign that ephemeral exchange (proving "this really
is me"), never to encrypt anything directly. Once the handshake finishes, both sides discard
the ephemeral private keys, so nothing recoverable from the long-term key remains. - What an operator must actually configure: enable and prefer cipher suites containing
ECDHE (or classic DHE), for exampleTLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256, and disable
static-RSA key-exchange suites (TLS_RSA_WITH_...), where the session key is directly
RSA-encrypted with the server's long-term public key and is therefore fully recoverable if
that private key ever leaks. TLS 1.3 (RFC 8446) removes static-RSA key exchange entirely and
makes ephemeral exchange mandatory, so upgrading to it is itself a forward-secrecy fix.
Session resumption is the other place this quietly breaks: if session tickets are encrypted
under a long-lived ticket key, compromising that ticket key can undo forward secrecy for
every resumed session, so ticket keys need their own rotation schedule.
Worked example
An attacker passively records two years of encrypted traffic between a client and
bank.example.com. In year three, the bank suffers a breach and its long-term certificate
private key leaks. If the bank's TLS configuration used ECDHE, the attacker's two-year
recording is still worthless: every one of those sessions used its own ephemeral key pair,
generated and discarded at handshake time, and the leaked long-term key was never used to
encrypt anything, only to sign the ephemeral exchange. If the bank had instead used a static-RSA
key-exchange suite, the same leaked key directly decrypts every recorded session key from the
last two years, because in that mode the session key was RSA-encrypted straight to the server's
long-term public key and never expires on its own.
Trade-offs & pitfalls
- A server can support ECDHE in principle and still fail to get forward secrecy in practice
if a legacy static-RSA suite is still enabled and a client or load balancer negotiates down
to it. - Forward secrecy adds a small amount of per-handshake computation (a fresh key exchange
every time versus a single static RSA operation), which is why it was historically seen as
a cost/security trade rather than a default; TLS 1.3 settled that trade by removing the
alternative.
Explain the roles of salting and key stretching in password-based key derivation and storage. Describe how salts should be generated and stored, why unique salts prevent precomputation/rainbow-table attacks, and how key stretching (iterative hashing, memory-hard functions) increases attacker work. Illustrate with concrete examples of attacks that salting and stretching mitigate and note any remaining risks that require additional controls.
Sample Answer
Direct answer
Salting adds a random, unique-per-user value to a password before hashing so that identical
passwords never produce identical stored hashes; key stretching deliberately makes each
hashing attempt slow (or memory-hungry) so an attacker who steals the hash dump still has to
pay a real per-guess cost. Together they turn "crack the whole database at once with a
precomputed table" into "crack each password individually, slowly."
Structured elaboration
- Generating and storing salts: generate a fresh, random salt (commonly 16 bytes) from a
CSPRNG per user, and store it alongside the hash. The salt is not secret, it can sit in
plaintext in the same database row; its job is uniqueness, not confidentiality. - Why unique salts defeat precomputation: a rainbow table (a precomputed map from common
passwords to their hashes) is only useful because it lets an attacker look up a hash instead
of computing one. A unique salt per user means the same password produces a different hash
for every user, so an attacker would need a separate precomputed table per salt value, which
is infeasible at any real user-base size. - How key stretching increases attacker work: an iterated hash (many rounds of a hash
function) or a memory-hard function (Argon2id, scrypt) makes each individual guess
expensive in CPU time, memory, or both. This does not slow down a legitimate login (one
verification per attempt is cheap in absolute terms) but multiplies the cost of trying
billions of candidate passwords offline by the same factor.
Worked example
Suppose 5,000 users in a leaked database chose the password "password123". With a plain,
unsalted hash, all 5,000 rows show the identical hash value, so cracking it once (via a
rainbow table lookup or a single guess) reveals all 5,000 accounts simultaneously. With unique
per-user salts, those same 5,000 users produce 5,000 different hash values, and an attacker
must run the guess-and-check process separately against each one; a rainbow table built for
the unsalted case is useless here since it was never built for any of these specific salts.
Trade-offs & pitfalls (remaining risks needing additional controls)
- Salting and stretching only make offline cracking (against a stolen hash) expensive; they
do nothing against online guessing at the login endpoint, which still needs rate limiting
and account lockout/backoff. - A short or common password remains crackable even when properly salted and stretched,
because the search space itself is small; this is why NIST SP 800-63B and similar guidance
push for password length and breach-list checking rather than relying on hashing alone. - Choosing stretching parameters too low (to save server CPU) narrows the gap between "safe"
and "fast enough to crack anyway"; parameters need periodic revisiting as hardware gets
cheaper.
Unlock Full Question Bank
Get access to all 11 Cryptography Fundamentals interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.