Applied Cryptography and Key Management Questions
Selecting and applying cryptographic primitives correctly: symmetric and asymmetric encryption, hashing, digital signatures, key derivation, secure random number generation, and public key infrastructure. Covers key lifecycle management, key exchange and distribution, choosing appropriate algorithms for a given constraint set including resource-constrained environments, and the forward-looking side of algorithm lifecycle: cryptographic agility and algorithm-migration strategy, forward secrecy, and the post-quantum cryptography transition and planning upgrades without breaking existing data or interoperability. The applied-crypto engineering layer, distinct from compliance-driven crypto standards.
Design a scalable PKI and key-management solution for 10 million IoT devices, many with intermittent connectivity and limited TPM/HSM capability. Address secure provisioning, key storage choices, rotation, revocation strategies where CRL/OCSP don't fit well, OTA updates, and how you'd bootstrap a root of trust.
Sample Answer
Direct answer
At 10 million devices with intermittent connectivity, favor short-lived certificates that are
renewed opportunistically on next contact over classic revocation checking (Certificate Revocation
Lists, CRLs, or the Online Certificate Status Protocol, OCSP), because a device that cannot reliably
reach a live revocation server cannot be trusted to check one anyway. Bootstrap trust from a small,
offline root Certificate Authority (CA) that only ever signs intermediate CAs, and generate each
device's identity keypair inside its Trusted Platform Module (TPM) or secure element when one is
available, falling back to a documented, weaker posture when it is not.
Structured elaboration
- Secure provisioning. At manufacturing time, generate the device's identity keypair inside its
secure element or TPM if it has one, so the private key never exists outside hardware. For cheaper
devices with no secure hardware, inject a per-device secret at a controlled factory line using a
Hardware Security Module (HSM), and treat that factory HSM itself as a critical root-of-trust asset. - Key storage choices. Prefer Elliptic Curve Cryptography (ECC, for example the P-256 curve) over
RSA for constrained devices: smaller keys and faster operations mean less flash storage and battery
drain. Where there is no secure element, the private key is encrypted at rest using a device-unique
wrapping key derived from the factory-provisioned secret, an explicit compromise, not equivalent to
hardware-backed storage, and worth stating plainly to reviewers rather than glossing over. - Bootstrapping the root of trust. Keep the root CA offline (air-gapped, used only in occasional,
audited signing ceremonies) and have it sign intermediate CAs, which are what actually issue device
certificates day to day. A compromised intermediate can then be revoked and replaced without
invalidating every device's trust in the root itself. - Revocation where CRL/OCSP do not fit. A CRL that lists revoked certificates for 10 million
devices becomes too large for constrained devices to fetch or parse, and OCSP needs an always
reachable responder many IoT deployments cannot guarantee. Short-lived certificates sidestep this:
if a device is compromised, you simply refuse to renew its certificate on its next check-in, and it
naturally stops being trusted once its current certificate expires. For urgent revocation before
natural expiry, push a small, compact deny-list (a Bloom filter or delta CRL) opportunistically to
gateways and provisioning servers via over-the-air (OTA) updates, rather than requiring every device
to check it directly. - OTA firmware updates. Sign firmware images with a separate code-signing key, distinct from each
device's identity key, and verify the signature plus a rollback-protection counter before flashing.
Rotate the signing key on a much slower cadence than device identity certificates, using a small
hierarchy of trusted signing keys so devices are not locked to a single one forever. - Handling long offline periods. Allow a short grace window where a recently expired certificate
is still accepted for the sole purpose of renewal, so a device left unplugged in a warehouse for a
few weeks does not get permanently bricked the moment it reconnects.
Worked example
Suppose devices check in roughly every 7 days on average, and you issue 30-day certificates with
automatic renewal starting at day 20. A device that misses up to two consecutive weekly check-ins (14
days late) is still comfortably inside its 30-day validity window and renews cleanly the next time it
connects. A device offline for 45 days has exceeded that window and needs the grace-window or
re-bootstrap path instead. This is a direct consequence of the numbers you choose (30-day validity,
20-day renewal start, roughly 7-day check-in cadence), and it is exactly the kind of arithmetic worth
running explicitly before committing to certificate lifetimes at this scale.
Trade-offs & pitfalls
- Reaching for a classic CRL at this device count is a common wrong first instinct; the list itself
becomes the bottleneck long before any cryptographic weakness would. - OCSP assumes reliable outbound connectivity that many locked-down IoT network policies deliberately
restrict, so building a revocation architecture around it for this fleet often fails in the field
even if it works in a lab. - Keeping the root CA offline and having it sign only intermediates is what lets you recover from a
compromised intermediate without a full 10-million-device re-provisioning event; skipping that
layer is a design mistake that only becomes visible during an actual incident.
Design a remote-attestation and secure firmware-update pipeline for a fleet of HSM-like devices or secure elements. Include boot-chain verification, signed firmware images, an attestation protocol to verify device state, rollback protection, an emergency-revocation path, and a safe staged rollout with recovery if a firmware push goes wrong.
Sample Answer
Treat the device fleet as a chain of custody from silicon to running firmware: every boot stage verifies the next stage's signature before running it, the device can prove its current state to a remote verifier through a signed attestation report, and firmware rollout happens in reversible, monitored stages with a fast kill switch if something goes wrong.
Framework
- Boot-chain verification: an immutable first-stage bootloader, burned into read-only memory or write-protected at manufacturing (the hardware root of trust), verifies the signature of the second-stage bootloader before executing it; that stage verifies the next, and so on up to the application firmware. Each link is checked against a public key that itself only changes through a signed, audited update to the trust anchor. This is the standard secure-boot chain pattern.
- Signed firmware images: every firmware image is signed by the vendor's release key, ideally HSM-backed, and carries a monotonically increasing version number inside the signed metadata, not just in an unsigned filename or header.
- Attestation protocol: the device measures each boot-chain component, meaning it computes a cryptographic hash of each stage, into a protected register (conceptually similar to a Trusted Platform Module's platform configuration registers). On request, it signs a report of those measurements plus a fresh, server-supplied random value (a nonce, which prevents an attacker from replaying an old, valid-looking report) using a device-unique attestation key. A remote verifier checks the signature and compares the measurements against known-good values to confirm the device is running genuine, unmodified firmware.
- Rollback protection: store the minimum acceptable firmware version in protected, monotonic storage inside the secure element, meaning it can only increase, so an attacker cannot present an older, validly-signed-but-vulnerable firmware image to downgrade the device. The boot chain refuses to run any signed image whose version is below that stored minimum.
- Emergency revocation: maintain a mechanism, often a simple monotonic "minimum acceptable version" bump rather than a full certificate-style revocation list, that the fleet checks on next contact, so a newly discovered vulnerability in a specific firmware version can be blocked fleet-wide, even for devices that already have that validly signed image installed.
- Staged rollout with recovery: push a new firmware version to a small canary population first, monitor attestation reports and health telemetry, and widen the rollout ring by ring. Keep the previous known-good image available in a separate flash partition (A/B partitioning) so a device that fails to boot the new image, or fails a post-update attestation check, automatically falls back to the previous verified-good image rather than becoming unreachable.
Worked example
A fleet of secure elements runs firmware v12. A canary ring of 1% of devices receives v13 first; each canary device attests after updating, and the backend confirms the signed measurements match the expected v13 hashes and that error rates stay flat before widening the rollout to the next ring. If a device fails to boot v13, its bootloader automatically falls back to the v12 image in the alternate partition and reports the failed boot in its next attestation, which halts the fleet-wide rollout. A critical vulnerability later found in v11 triggers bumping the fleet-wide minimum acceptable version, so any remaining device still on v11 is blocked from booting it and forced onto an update path, even though that v11 image remains validly signed.
Trade-offs and pitfalls
Rollback protection and emergency revocation both depend on a genuinely tamper-resistant, monotonic counter; a counter that can be reset defeats the whole purpose, since an attacker could simply reset it and replay an old image. A common mistake is checking signature validity without also checking the version against the rollback floor, which lets an attacker replay an old, validly-signed, vulnerable firmware image. The A/B fallback partition needs its own attestation check too, otherwise "recovery" can itself become a downgrade path if that partition is allowed to hold an arbitrarily old image instead of being kept in step with the current rollback-protection floor.
Design a key-management service to provision, authenticate, and rotate cryptographic keys for a fleet of intermittently-connected, low-power IoT devices (from tens of thousands to hundreds of millions of units). Cover manufacturing provisioning (pre-seeded identities), initial bootstrap, use of TPMs or secure elements where available, over-the-air key updates for devices offline for months, and how you handle revocation or compromise at that scale.
Sample Answer
Direct answer
Design the Key Management Service (KMS) around a per-device identity seeded at manufacturing time, an
initial bootstrap that authenticates each device to the provisioning backend on its first network
contact using that identity, and an idempotent "sync to the latest desired key state" model for
over-the-air (OTA) updates, since a device offline for months should not need to replay every
intermediate update in sequence. At hundreds of millions of devices, revocation has to work at the
level of manufacturing batches and cohorts, not one device at a time, because you often will not know
which individual devices are affected before you know which production run was compromised.
Structured elaboration
- Manufacturing provisioning. Pre-seed each device with a unique identity at the factory line,
generated inside a Trusted Platform Module (TPM) or secure element when the device has one, so the
private identity material never leaves hardware, or injected via a factory-line Hardware Security
Module (HSM) for devices with no secure element, an explicit, documented weaker posture rather than
something left implicit. - Initial bootstrap. On first network contact, the device authenticates to the provisioning backend
using its pre-seeded identity (for example proving possession of its factory-issued key), and only
after that authentication succeeds does the backend issue whatever session or operational keys the
device needs going forward. This separates "prove who you are" (the identity key, rarely touched
again) from "what you can do right now" (operational keys, rotated normally). - TPM or secure element usage where available. The identity key stays sealed inside hardware for
devices that have it; for devices that do not, the identity secret is encrypted at rest under a
device-unique wrapping key, and this fallback should be tracked as a known, weaker posture in your
device inventory, not treated as equivalent to hardware-backed storage. - OTA key updates for devices offline for months. Rather than a sequential update log a device must
replay in order (which breaks badly for a device that missed many updates), have each device fetch
its full CURRENT desired key state on next contact and reconcile directly to it. This is idempotent:
a device offline for a week and a device offline for eight months both converge to the same correct
state in one round trip, rather than needing eight months of intermediate steps replayed correctly. - Revocation or compromise handling at hundreds-of-millions scale. A single global list of revoked
devices, whether a classic Certificate Revocation List or a simple deny-list, grows too large for
constrained devices to fetch or for the backend to distribute efficiently at this scale. Instead,
organize revocation by manufacturing batch or cohort: if a specific factory production run's
provisioning HSM is later found to have been compromised, you can revoke and re-provision that entire
cohort directly, without first needing to identify which individual devices within it were actually
affected. This is a distinctly larger-scale, batch-oriented operational concern than revoking
individual devices one at a time, and it depends entirely on having tracked which cohort or
manufacturing batch every device belongs to from the moment it was provisioned.
Worked example
Suppose a fleet of 200 million devices is organized into manufacturing cohorts of roughly 500,000
devices each, one per factory production run, and the provisioning system logs each device's cohort ID
at manufacturing time. If cohort 137's factory-line HSM is later found to have been compromised during
a specific production window, the KMS can immediately flag all roughly 500,000 devices in cohort 137
for forced re-provisioning on next contact, a targeted, bounded action, rather than either revoking
all 200 million devices (unnecessary and operationally catastrophic) or trying to first identify which
individual devices among 200 million were actually affected (infeasible without cohort tracking). The
cohort boundary is what turns an intractable, fleet-wide problem into a bounded, roughly 0.25 percent
of the fleet action.
Trade-offs & pitfalls
- A sequential OTA update log that a device must replay in strict order is a common early design choice
that breaks badly for exactly the long-offline devices this question is about; idempotent
state-reconciliation avoids that failure mode entirely. - Skipping cohort or batch tracking at provisioning time is invisible until the first factory-level
compromise, at which point you discover you cannot scope the incident without it, and the honest fallback
is treating the entire fleet as suspect. - Devices without a secure element carry a real, ongoing risk that should be tracked explicitly in your
device inventory (which devices have hardware-backed identity versus software-wrapped identity),
rather than treated as a one-time caveat and then forgotten in day-to-day operations.
Design a secure, post-quantum-ready code-signing and firmware-update architecture for highly constrained embedded devices (think 32KB RAM, a slow CPU, limited flash). Walk through your signature-scheme selection, verifier performance in the bootloader, how you handle storage and transmission of larger PQ signatures, and your strategy for backward compatibility with legacy devices in the field.
Sample Answer
At 32KB of RAM you cannot naively swap in a general-purpose post-quantum signature scheme, since several standardized options are too large or memory-hungry for a bootloader verifier. The practical path is a hash-based signature scheme built for small, fast verification, run alongside the existing classical signature during the transition, with the larger post-quantum signature streamed from flash rather than held fully in memory.
Framework
- Signature-scheme selection: of the schemes the U.S. National Institute of Standards and Technology (NIST) has standardized, ML-DSA (Dilithium, specified in FIPS 204) and SLH-DSA (SPHINCS+, specified in FIPS 205) are the general-purpose post-quantum signature options. SLH-DSA is stateless and relies only on hash-function security rather than newer lattice-hardness assumptions (a newer family of hard-math problems that lattice-based schemes rely on, whereas SLH-DSA rests only on the security of hash functions), but has larger signatures and slower signing; its verification, which is the only operation this constrained device ever performs, is comparatively cheap. A stateful hash-based scheme like LMS or XMSS (specified in RFC 8554 and RFC 8391) is a common embedded choice specifically for firmware signing, with very small, fast verification code and a long track record, at the cost of needing careful state management on the signer's side, not the device, to avoid catastrophic key reuse.
- Verifier performance in the bootloader: the constrained device only ever verifies, it never signs, so pick the scheme whose verification is cheapest in code size and processing time, even if signing (which happens once, on well-resourced build infrastructure) is slower or requires careful state tracking. LMS, XMSS, and SLH-DSA all have comparatively lightweight verification relative to their signing cost, which is the right trade-off to make for a pure verifier.
- Storage and transmission of larger signatures: post-quantum signatures and public keys are meaningfully larger than a classical elliptic-curve signature's tens of bytes; SLH-DSA signatures commonly range from roughly 8 kilobytes up to 30-plus kilobytes depending on the parameter set chosen, and LMS signatures typically run to a few kilobytes depending on the tree height used. At 32KB total RAM you already cannot buffer a full firmware image in memory, so the update flow should already stream the image and verify incrementally against a running hash; the fix is to also stream the signature and any authentication-path data (the set of sibling hashes a verifier needs to walk the signature's hash tree back to the trusted public key) from flash rather than holding them fully in RAM, budgeting flash space (which is comparatively cheap) rather than RAM (which is not) for the larger signature.
- Backward compatibility: run in hybrid mode during the transition, signing firmware with both the existing classical scheme and the new post-quantum scheme, and verifying with whichever the specific device's bootloader supports. Legacy devices already in the field keep working unmodified against the classical signature, while newer bootloader revisions also check the post-quantum signature, and can later be configured to reject updates lacking it once the fleet has migrated.
Worked example
A firmware update bundle carries the image, an unchanged elliptic-curve signature (about 64 bytes) for legacy verifiers, and an LMS signature (a few kilobytes) for updated verifiers. The bootloader streams the image from the update partition through a hash function block by block, never holding the whole image in RAM, then streams the LMS signature and its authentication path from flash to verify the final hash against the embedded public key, all within a working set comfortably under 32KB. A device still running the old bootloader simply ignores the LMS signature and verifies only the classical one, so the same update artifact serves both fleets during a multi-year rollout.
Trade-offs and pitfalls
Choosing a stateful scheme like LMS or XMSS pushes real operational risk onto the signer: reusing a one-time signing state catastrophically breaks the scheme's security, so the build and release infrastructure needs hardened, HSM-backed state tracking even though the constrained device itself stays stateless. Shipping both a classical and a post-quantum signature roughly doubles signature size and verification code, a real cost against limited flash that has to be budgeted for, not assumed away. A common mistake is selecting the post-quantum scheme with the best signing performance and only later discovering its verification routine doesn't fit the bootloader's code-size budget; for a device that only ever verifies, verification cost should be the first selection criterion, not signing cost.
Design a secure key-derivation and device-authentication scheme for resource-constrained IoT devices that have low-quality entropy at manufacturing time and may be provisioned offline. The scheme needs to prevent device cloning, provide a durable device identity, and support firmware updates while minimizing long-term secret disclosure. Describe your provisioning, storage, and update flows and the trade-offs involved.
Sample Answer
The core fix for low-quality entropy at manufacturing time is to never rely on the device's own random-number generator for its root identity. Instead, inject or derive a strong per-device secret during a controlled manufacturing step, bind it to a hardware root of trust so cloning requires physically extracting protected hardware secrets, and keep the identity key separate from any key that needs to change, like a firmware-verification key, so the device's cryptographic identity survives updates.
Framework
- Provisioning despite low entropy: instead of trusting the device to generate its own randomness at first boot, either inject a unique secret from a trusted external entropy source during manufacturing (a factory-side HSM generates and writes a strong per-device key into protected, one-time-programmable storage), or use a physically unclonable function (PUF), a technique that derives a device-unique secret from tiny, unrepeatable manufacturing variations in the silicon itself. A PUF-derived secret is available fully offline and doesn't depend on the device's runtime random-number generator at all.
- Durable device identity: derive the device's long-term identity key pair from that injected or PUF-derived root secret using a key derivation function (KDF), and record the public identity key, or a certificate binding it to the device's serial number, in a manufacturing database, so it can be verified later without needing network access to the device itself.
- Anti-cloning: because the root secret is bound to physical hardware, copying the device's flash memory to a clone does not give the clone the identity key, since the secret either never existed as extractable data (PUF) or is locked in fused, one-time-programmable storage. Any authentication protocol should require a proof-of-possession challenge-response against the identity key, not just presenting a certificate, so a cloned flash image without the physical secure element fails authentication.
- Offline provisioning: since devices may be provisioned with no network connection, generate identity certificates at manufacturing time and load them into the backend's device registry out of band, so the device's first network contact can be verified against a pre-loaded record rather than requiring a live certificate authority interaction.
- Firmware updates without long-term secret disclosure: use a separate firmware-signing key, held by the vendor's release infrastructure and ideally HSM-backed, to sign firmware images. The device verifies signatures with a public key baked in at manufacturing, so firmware updates never touch or re-provision the device's identity secret, and a firmware-signing key can be rotated independently without re-manufacturing devices.
Worked example
A sensor is manufactured with an embedded secure element containing a PUF. At the factory, a provisioning station reads the PUF-derived secret via a challenge-response exchange, derives a device identity key pair from it plus the device's serial number, and records the public key against that serial in a manufacturing database before the device ever touches a network. In the field, the device authenticates by producing a proof-of-possession signature over a server-issued nonce, so a cloned firmware image without the physical secure element fails authentication. A later firmware update ships signed by the vendor's release key and is verified using a public key baked into the boot ROM, without touching the device's identity key at all.
Trade-offs and pitfalls
Physically unclonable functions add manufacturing cost and complexity, while cheaper alternatives (fused-in secrets from a factory HSM) require trusting the manufacturing supply chain not to leak the provisioning database. A common mistake is deriving both device identity and firmware trust from the same root secret; that couples them unnecessarily, so a firmware-key rotation now also has to touch device identity and vice versa. Keeping the two cryptographically separate is exactly what gives you the property that identity survives updates.
That is every published Applied Cryptography and Key Management question for Embedded Developer so far. Browse the other topics in this category, or practice this one interactively.