Identity, Authentication, and Access Management Questions
Designing and operating identity and access control systems. Covers authentication protocols and standards (OAuth, SAML, OIDC, MFA), authorization models (RBAC, ABAC), identity lifecycle and privilege management, IAM architecture and automation, and access control across cloud and on-premises environments. The 'who can do what' control plane, distinct from cryptographic key management.
Explain the structure of a JSON Web Token (JWT): header, payload (claims), and signature/encryption. Describe the difference between signed JWTs (JWS) and encrypted JWTs (JWE). When is confidentiality required (use JWE) versus only integrity/authenticity (JWS)? Describe risks of placing sensitive data in token claims and best practices for claim design.
Sample Answer
Direct answer
A JSON Web Token (JWT) is a compact, URL-safe way to represent a set of claims, cryptographically protected so a recipient can trust them. Structurally it is base64url-encoded segments joined by dots: a header describing how the token is protected, a payload carrying the claims, and a final segment (or segments) that either signs or encrypts the rest. A JSON Web Signature (JWS) is the signed variant, and it's what most JWTs in practice actually are. A JSON Web Encryption (JWE) is the encrypted variant, used when the payload itself must be hidden from everyone except its intended recipient, not merely protected from tampering.
Structured elaboration
- Header: a JSON object with at minimum
alg(the signing or encryption algorithm) andtyp(typically"JWT"), often alsokid(a key id identifying which key, among several, was used). - Payload (claims): a JSON object with registered claims such as
iss(issuer),sub(subject),aud(audience),exp(expiration),nbf(not-before),iat(issued-at), andjti(a unique token id), plus whatever public or private claims the issuer chooses to add, such as roles or a tenant id. - JWS versus JWE: a JWS token is three segments,
base64url(header).base64url(payload).base64url(signature). Base64url is an encoding, not encryption, so anyone holding a JWS token can read the payload in full; the signature only proves the payload hasn't been tampered with and really came from whoever holds the signing key. A JWE token has five segments instead of three (an encrypted content key, an initialization vector, ciphertext, and an authentication tag, alongside the header), and the payload is genuinely unreadable without the decryption key. - When to use which: reach for JWS, by far the more common choice, when the concern is only integrity and authenticity, that is, "did this really come from the issuer, unmodified," and the claims are not a problem to expose to anyone who happens to be holding the token, such as a user id, a role, or an expiration. Reach for JWE when the payload itself carries data that would be a genuine confidentiality problem if seen by an intermediary that merely relays the token without being its intended audience. In practice most access and ID tokens are JWS-only; JWE, sometimes nested inside a JWS ("sign then encrypt"), shows up specifically when a token has to pass through an untrusted intermediary that shouldn't be able to read its contents.
Worked example
Compare two claim designs for the same login:
Risky: {"sub": "user-42", "email": "alice@example.com", "ssn": "123-45-6789", "full_address": "...", "role": "admin"}, signed as a plain JWS. Because JWS provides no confidentiality at all, every one of those fields is visible in plaintext to anything that ever sees the token: the browser holding it, any proxy or load balancer that logs request headers, any error-tracking tool capturing request context, and every downstream service the token passes through, regardless of whether that service actually needs the social security number.
Better: {"sub": "user-42", "roles": ["admin"], "tenant_id": "acme-corp", "exp": 1735689600}. Only what's needed to make an authorization decision travels inside the token; anything genuinely sensitive, such as a social security number or a home address, stays server-side and is looked up by sub only when actually needed, rather than riding along inside a bearer credential that ends up in logs by default.
If an integration genuinely needs to carry sensitive data inside the token itself, for instance a short-lived token handed to a third-party partner across an intermediary message queue that shouldn't be able to read one specific field, that is the real JWE case: encrypt just that field, or wrap the whole JWS in a JWE layer, rather than defaulting every token to encryption. Encrypting everything adds real overhead across every relying party, which now needs a decryption key before it can read even routine claims like exp.
Where the resulting token gets stored on the client afterward, a cookie versus localStorage, is a related but separate decision with its own distinct attack surface; that trade-off belongs to the session-storage question, and this answer's concern is strictly what goes inside the token, not where it lives once issued.
Trade-offs and pitfalls
- The most common mistake is treating base64url encoding as a form of protection. It is trivially reversible by anyone, paste any JWT into a decoder and read it, so never put anything in a JWS payload you wouldn't be comfortable writing directly into a URL query parameter.
- Reaching for JWE by default "just to be safe" has a real cost: every relying party now needs the decryption key, or access to a decryption service, before it can read even routine claims, which complicates key distribution and rotation for little benefit if the payload was never sensitive to begin with.
- Claim design interacts with lifetime and revocation more than people expect: a token with no
jtito reference individually and a longexpis hard to revoke or reason about during an incident. Keep both the claim set and the lifetime as small as the use case actually requires.
Design a high-performance Attribute-Based Access Control (ABAC) policy evaluation engine capable of handling 1,000,000 authorization checks per second with complex policies and dynamic attributes. Include your choice of policy language, attribute retrieval and caching strategies, policy compilation or pre-evaluation techniques, consistency vs freshness trade-offs, horizontal scaling, and how you'd test correctness and performance under load.
Sample Answer
Direct answer
To hit 1,000,000 authorization checks per second, compile policies into an efficient in-process representation, for example Rego compiled to WebAssembly (Wasm), so each check never leaves the process, cache both raw attributes and full decisions with explicit freshness bounds, and scale horizontally by embedding the evaluator as a sidecar next to each calling service rather than routing every check through one centralized decision point. This design deliberately accepts bounded staleness (attributes and cached decisions can lag their source of truth by a tunable window) in exchange for the throughput a synchronous 1M-checks-per-second workload requires; correctness is then validated with a differential test harness that compares the fast path against a reference, always-correct evaluator.
Structured elaboration
Policy language
- Open Policy Agent's Rego is the strongest default: it separates policy authoring from the evaluator's implementation language, has a mature WebAssembly compilation target that turns a policy into a portable, sandboxed bytecode module embeddable directly in the evaluating process, and has first-class support for partial evaluation (below).
- Alternative: a narrower, purpose-built domain-specific language compiled directly to a native decision tree, trading Rego's generality for a smaller, more predictable per-check cost. Worth it only once profiling shows the evaluator itself, not attribute I/O, is the bottleneck.
Attribute retrieval and caching strategies
- Attributes split into two latency classes: request-local (already present on the incoming call, free) and remote (user/resource/environment attributes living in another system).
- Cache remote attributes locally on each evaluator node with a short time-to-live (TTL), plus push-invalidation for attributes where staleness has real consequences (a just-revoked project membership).
- Cache full decisions too, keyed by a canonical hash of the request's relevant attributes, since a real workload repeats the same (subject, resource-class, action) tuple often; this is the highest-leverage cache, since it skips policy evaluation entirely, not just remote attribute lookup.
Policy compilation and pre-evaluation techniques
- Compile the policy bundle ahead of time (Rego to Wasm, or a custom DSL to bytecode) rather than interpreting source text per request; a policy change is a build-and-deploy-a-new-bundle event, not a live text reload.
- Partial evaluation: where some attributes are known in advance (an endpoint where
resource.typeis always"invoice"), pre-specialize the compiled policy against those known values so the runtime evaluator has fewer branches per check.
Consistency vs freshness trade-offs
- Checking every attribute against a strongly consistent source of truth on every request is not achievable within a synchronous latency budget at this scale; the design deliberately trades strict consistency for bounded staleness, a decision may be computed against attributes up to the cache's TTL old.
- Make the staleness window asymmetric by attribute sensitivity: fast-changing, high-consequence attributes get push-invalidation and a near-zero TTL fallback; slow-changing, low-consequence attributes get a longer TTL, since staleness there is a minor annoyance, not a security incident.
Horizontal scaling
- Deploy the compiled evaluator as a sidecar (or in-process library) next to each calling service, not as one centralized policy decision point (PDP) cluster every request round-trips to; this removes the network hop from the hot path entirely and lets throughput scale linearly with the calling service's own replica count, a resource already scaled to meet its own traffic.
- A small central control plane distributes compiled bundles and cache-invalidation events to every sidecar, but stays off the per-request path.
flowchart LR
Req[Authorization Check]
Cache[(Attribute + Decision Cache)]
Compiler[Policy Compiler]
Bundle[(Compiled Policy Bundle)]
PIP[Attribute Sources: PIP]
Eval[In-Process Evaluator]
Req --> Cache
Cache -->|hit| Req
Cache -->|miss| Eval
Eval --> Bundle
Eval --> PIP
Compiler --> Bundle
PIP -->|push update| Cache
Testing correctness and performance under load
- Correctness: maintain a straightforward, unoptimized reference evaluator (no caching, no partial evaluation) that is trivially correct by construction. Run the same corpus of (policy, request) pairs through both the reference evaluator and the compiled fast path and assert identical decisions; any divergence is a caching or compilation bug caught before production.
- Performance: load-test the compiled evaluator in isolation (no network) to find its true per-core ceiling, then load-test end-to-end with a realistic, non-100% attribute-cache hit rate, since a 100% hit rate in a test would hide the cost of a real cache-miss rate.
Worked example
Capacity arithmetic for reaching the 1,000,000 checks/sec target. The per-node throughput and cache-miss rate below are stated design assumptions (labeled as such), not measurements; only the arithmetic that follows is the claim:
target_checks_per_sec = 1_000_000
per_node_checks_per_sec = 50_000 # assumption: one evaluator instance, warm in-process cache
cache_miss_rate = 0.01 # assumption: 1% of checks need the shared attribute store
nodes_needed_no_headroom = target_checks_per_sec / per_node_checks_per_sec
nodes_with_headroom = nodes_needed_no_headroom + 2 # N+2 for rolling deploys / node loss
fallback_checks_per_sec = target_checks_per_sec * cache_miss_rate
print(f"target: {target_checks_per_sec:,} checks/sec")
print(f"nodes needed at {per_node_checks_per_sec:,}/sec/node, no headroom: {nodes_needed_no_headroom:.1f}")
print(f"nodes with +2 headroom: {nodes_with_headroom:.1f} -> round up to {int(nodes_with_headroom) + 1}")
print(f"at {cache_miss_rate:.0%} local-cache-miss rate, the shared attribute store must sustain: {fallback_checks_per_sec:,.0f} lookups/sec")
Output (actually run):
target: 1,000,000 checks/sec
nodes needed at 50,000/sec/node, no headroom: 20.0
nodes with +2 headroom: 22.0 -> round up to 23
at 1% local-cache-miss rate, the shared attribute store must sustain: 10,000 lookups/sec
Twenty evaluator nodes cover the raw target with no margin; twenty-three gives standard rolling-deploy headroom. The shared attribute store, the one piece of this design that is NOT embarrassingly parallel per calling service, only needs to sustain 10,000 lookups/sec at a 1% local-cache-miss rate, two orders of magnitude below the raw target, which is exactly why the local-cache-hit path, not the shared store, has to carry the bulk of the throughput.
Trade-offs and pitfalls
- The core trade-off, bounded staleness, is also the biggest risk: a revoked permission that has not propagated yet is a live security gap for the length of the TTL or propagation lag. Push-invalidation for sensitive attributes reduces, never eliminates, this window, so it must be sized and monitored, not assumed away.
- Compiling policy to Wasm or bytecode adds a build/deploy pipeline for policy changes that a live-interpreted system didn't need; a policy fix becomes "compile, differentially test, deploy a new bundle" rather than "edit a file", slower by design, but exactly what makes the hot path fast and safe.
- A decision cache keyed on a hash of "relevant" attributes is only as good as that definition: omit an attribute the policy actually depends on and the cache serves wrong decisions that look like correct hits. Differential testing must include cache-key coverage, not just evaluator-logic coverage.
- Sidecar-per-service scaling avoids a central bottleneck but multiplies operational surface: every sidecar needs bundle updates and cache warm-up, and a bad bundle rollout becomes a fleet-wide problem instead of a single service's problem.
Design a system for issuing, rotating, and revoking API keys used by services and external partners. The system must support automated rotation without downtime, bind keys to service identities and scopes, provide audit trails, integrate with CI/CD and secret managers, and support emergency revocation. Describe issuance flows, rotation strategies (rolling, short-lived), migration for consumers, and how to deprecate long-lived keys in favor of ephemeral tokens.
Sample Answer
Direct answer
An API key lifecycle system treats a key the way this domain treats any other credential: bound to a specific service identity and an explicit scope at the moment it is issued, rotated on a schedule using an overlap window so rotation itself never causes downtime, delivered through a secrets manager and the continuous integration/continuous delivery (CI/CD) pipeline so it is never hand-copied into a config file, and backed by a separate, faster emergency-revocation path for the rare case where a confirmed compromise cannot wait for the next scheduled rotation.
Structured elaboration
Issuance flows. Every key is bound at issuance to a specific service identity, which team or service owns this key, and an explicit scope, the specific resources and actions it may access, never a default of full account access. Both must be stated before the key is minted, not attached afterward. Keys are generated with sufficient entropy (a cryptographically random value, never a guessable or sequential identifier), and the issuance event itself is the first audit-trail entry: who requested it, for which service, with what scope, and who approved it if the scope requested crosses an approval threshold.
Rotation strategies: rolling and short-lived. Rolling rotation is dual-key validity: generate a new key while the old one stays fully valid, distribute the new key to consumers, and only retire the old one once usage telemetry confirms nothing is still calling with it, never retiring the old key before consumers have had a chance to adopt the new one, since that ordering is exactly what turns a routine rotation into an outage. Short-lived keys take a different approach: instead of periodically rotating a long-lived key, issue keys with a short built-in expiration (hours to days) and have consumers fetch a fresh one automatically before the old one expires, moving rotation from a scheduled event into a continuous property of how the key is obtained at all. This bounds a leaked key's usable life automatically, without depending on anyone noticing the leak. Rolling rotation is simpler to retrofit onto systems built around a stable, occasionally-changing credential; short-lived keys need consumer-side support for automatic re-fetching, more upfront engineering, particularly for legacy consumers, but bound exposure far tighter without depending on a rotation schedule actually running on time.
Binding keys to service identities and scopes, and preserving that binding through rotation. Rotating a key must never silently change what it is allowed to do: the new key generated during rotation should carry the exact same identity and scope binding as the one it replaces, and any scope change should be its own deliberate, separately audited action, not a side effect of a routine rotation someone bundled it into.
Providing audit trails. Log every lifecycle event, issued, rotated, aggregated usage, revoked, expired, with the key identifier, the service identity it belongs to, who or what automation initiated the event, and a timestamp, so a later security review can answer when a specific key was created, whether it has actually rotated on schedule, and when it was last used, without archaeology.
Integrating with CI/CD and secret managers. A key should never be pasted into a config file, set by hand as an environment variable, or committed to a repository, a leading real-world source of leaked credentials. Instead, the CI/CD pipeline and running services fetch the current key directly from a secrets manager, a dedicated service built to store and serve credentials with its own access control and audit trail, at deploy or runtime. Rotation then becomes updating one value in the secrets manager and letting consumers pick it up on their next fetch or restart, not hunting down every place the key was ever manually copied.
Supporting emergency revocation. This path is deliberately separate from, and faster than, scheduled rotation: a confirmed-compromised key needs to be invalidated immediately, accepting the disruption risk that scheduled rotation exists specifically to avoid, because a compromised key's continued validity is a worse outcome than a brief consumer failure. Emergency revocation should be one fast, well-tested action, not a hurried variant of the normal rotation process improvised under pressure.
Migration for consumers. Whenever the key-delivery mechanism itself changes, for example moving from a key baked into a deploy-time config file to a key fetched live from a secrets manager, the migration should follow the same dual-validity principle as key rotation itself: support both delivery mechanisms simultaneously for a transition window, track which consumers have actually moved based on which path their traffic uses, and retire the old mechanism only once telemetry confirms every consumer is off it, the identical generate-distribute-confirm-retire pattern, applied one level up, to the delivery mechanism rather than the key value.
Deprecating long-lived keys in favor of ephemeral tokens. The end state many organizations are moving toward eliminates standing, long-lived API keys for service-to-service calls entirely, replacing them with workload-identity-issued ephemeral tokens: a service authenticates using its own platform-attested identity and receives a short-lived token, rather than presenting a static secret at all. This is a different security posture, not merely a faster rotation cadence on the same one: a long-lived key, however well-rotated, is still a bearer secret usable by anyone who has a copy of it, while a workload-identity token is bound to, and only obtainable by, the legitimate workload itself, as verified by the platform. Migrating existing consumers to this model is a genuine, longer-term project, since it requires workload-identity infrastructure to exist and every consumer to support it, so disciplined rotation of long-lived keys is the realistic interim posture for systems that cannot make that jump yet, not a permanently acceptable end state.
sequenceDiagram
participant Rotator as Rotation automation
participant Vault as Secrets manager
participant Svc as billing-service
participant Proc as Payment processor
Rotator->>Proc: Generate key_v13 (same identity + scope as key_v12)
Rotator->>Vault: Write key_v13 to new version slot (key_v12 untouched)
Svc->>Vault: Fetch current key on next deploy/refresh
Vault-->>Svc: key_v13
Svc->>Proc: Calls now use key_v13
Rotator->>Proc: Check usage telemetry for key_v12
Proc-->>Rotator: key_v12 usage at zero for grace period
Rotator->>Proc: Retire key_v12
Rotator->>Vault: Log rotation event to audit trail
Worked example
Acme's billing-service calls a third-party payment processor's API using an API key, on a 30-day rotation policy. key_v12 is currently active, bound to service identity billing-service with scope payments:charge, payments:refund.
Five days before scheduled rotation, automation generates key_v13 with the identical identity and scope binding as key_v12, and writes it to the secrets manager under a new version slot without touching key_v12. Both keys are valid simultaneously at the processor. Over the following days, billing-service's deployment pipeline picks up key_v13 on its next scheduled redeploy, and usage telemetry at the processor shows calls on key_v12 dropping toward zero as instances roll over. On the scheduled rotation day, automation checks that telemetry and confirms key_v12 has shown zero usage for the defined grace period before retiring it, and the audit trail records the full sequence: key_v13 issued, key_v12 confirmed at zero usage, key_v12 retired.
Separately, if a security scan discovers on day 27 that key_v12 was accidentally committed to a public repository months earlier, that is not handled by waiting for the scheduled retirement. Emergency revocation invalidates key_v12 immediately, accepting that any consumer instance not yet migrated to key_v13 will see failed calls until it redeploys, which is the correct trade-off given a confirmed leak.
Longer-term, Acme's platform team is migrating internal callers of its own internal services away from static API keys entirely, toward workload-identity-issued short-lived tokens for service-to-service calls. billing-service's call to the external payment processor still uses a rotated API key, since Acme does not control the processor's own authentication model; the ephemeral-token end state applies most readily to systems an organization controls end-to-end, while integrations built on an external provider's own issuance model realistically stay on the disciplined-rotation posture for longer.
Trade-offs and pitfalls
- Retiring an old key before telemetry confirms every consumer has migrated is the single most common way "automated rotation" causes the outage it was built to prevent. The grace-period confirmation step is not optional overhead; it is what makes the automation safe rather than merely automatic.
- A rotation that also quietly changes scope conflates two different kinds of change, and makes it much harder to know later which one caused a problem; scope changes deserve their own deliberate, separately-audited action rather than riding along with a routine rotation.
- Committing a key, even briefly and even to a private repository, remains a common real leak vector no matter how good the rotation policy is. Secrets-manager integration exists specifically to remove the need to ever put a key value somewhere that gets committed, version-controlled, or manually copied between people.
- Assuming short-lived keys are strictly better than rolling rotation ignores the real engineering cost on the consumer side. A legacy consumer that cannot support automatic re-fetching gets no benefit from a short-lived-key policy; it will simply hardcode whatever key it was given until it breaks, so the migration cost has to be budgeted, not assumed away by choosing the theoretically better strategy.
- Rotating a long-lived key frequently is not equivalent to eliminating long-lived keys. However tight the rotation window, the key remains a bearer secret usable for its full remaining validity if stolen; workload-identity-issued ephemeral tokens are a categorically different posture, not a faster version of the same one, and treating them as interchangeable understates how much protection the shift actually buys.
Compare JWT (self-contained) tokens and opaque tokens. Explain verification differences, revocation strategies, performance characteristics, security tradeoffs, and when you might choose one format over the other for access tokens and refresh tokens.
Sample Answer
Direct answer
A JWT (JSON Web Token) is self-contained: its claims (who the token belongs to, what it grants, when it expires) travel inside the token itself, base64url-encoded and signed, so any service holding the right public key or shared secret can verify it alone, with no network call. An opaque token is a random string with no embedded meaning; it only becomes useful after a lookup (a database query or an introspection call to the authorization server) resolves it to the claims and validity state stored server-side. That one structural difference (state carried in the token versus state carried in a server-side store) drives every trade-off below.
Structured elaboration
| Dimension | JWT (self-contained) | Opaque token |
|---|---|---|
| Verification | Local signature check (HMAC secret or RSA/ECDSA public key); no network hop, no shared-store dependency | Requires a lookup: a call to the authorization server's introspection endpoint, or a query against a shared token store |
| Revocation | Hard: the token is valid until its own expiry unless the system adds an out-of-band revocation mechanism (a deny-list, short lifetimes, or a Bloom-filter pre-check) | Trivial: mark or delete the server-side row, and the very next introspection call reflects it immediately |
| Performance at scale | Verification cost does not grow with request volume across many resource servers, since each one verifies independently; the cost is the (cheap) signature check itself | Every verification adds a real network or store round trip; this can be reduced with caching, but caching then re-introduces a revocation-delay window proportional to the cache time-to-live |
| Security exposure | The payload is encoded, not encrypted, so anyone who intercepts the token can read its claims (never put secrets or sensitive personal data in a JWT payload unless it is also encrypted as a JWE, JSON Web Encryption, token); a stolen JWT is immediately usable until expiry | The token itself reveals nothing (it is meaningless outside the issuing system); a stolen opaque token is still usable until revoked, but revocation is instantaneous once detected |
Choosing per token type. Access tokens are commonly issued as JWTs, because they are verified on nearly every request across potentially many independent microservices, and paying a database or introspection round trip on every single one of those checks does not scale the way a local signature check does; the downside (hard to revoke early) is contained by keeping access-token lifetimes short. Refresh tokens are commonly issued as opaque tokens instead: they are validated in exactly one place (the authorization server itself, never fanned out to resource servers), used far less often, and the whole point of a refresh token is to be the thing you can shut off immediately (on logout, on a detected compromise, on password change). That means the opaque token's strong revocability matters far more for a refresh token than the JWT's fan-out verification speed does, since there is no fan-out to speed up in the first place.
Worked example
Consider an API gateway in front of 12 microservices, each independently verifying the caller's access token on every request, at a combined rate of 5,000 requests/second across the fleet.
- With JWT access tokens: each service verifies the signature locally. No service makes an extra network call to authenticate a request; the 5,000 req/sec figure adds zero load to any shared authentication store, because there is not one.
- With opaque access tokens and no caching: every one of those 5,000 requests/second becomes a lookup against the shared token store or an introspection call, meaning the authorization server (or shared cache) now has to sustain 5,000 lookups/second just for token checks, on top of its own normal load, and every microservice's request latency now includes that round trip.
This is exactly why the fan-out access-token case favors JWT and the single-point refresh-token case does not need it: a refresh token is checked once per refresh cycle at one server, not 5,000 times/second across 12 independent callers.
Trade-offs and pitfalls
The most common mistake is treating "self-contained" as "safe to put anything in," and shipping personal data or authorization decisions in a JWT payload that then leaks through logs or browser storage; encode only what each verifier needs. The second is picking JWTs for a system that actually needs fast, precise revocation (a banking session, an admin token) without also building a revocation mechanism, which reintroduces the exact problem opaque tokens solve for free. The reverse mistake is picking opaque tokens for a large microservice fleet without a caching strategy, which silently turns the authorization server or token store into a single point of contention that every request depends on.
In Node.js using Express, outline middleware that verifies a Bearer JWT access token signed with RS256 by using a JWKS endpoint. Provide the high-level pseudocode, list error cases to handle (expired, malformed, unknown kid), and explain how to cache JWKS keys safely to avoid performance problems.
Sample Answer
Direct answer
An Express middleware that verifies a Bearer RS256 JWT (JSON Web Token) needs to: extract the token from the Authorization header, resolve the signing key by the token's kid from a JWKS (JSON Web Key Set) endpoint through a caching client rather than fetching on every request, verify the signature with the algorithm pinned to RS256, and map each distinct failure (missing header, malformed token, expired token, unknown kid) to a 401 response with its own machine-readable error code, so logs and clients can tell the failure modes apart instead of seeing one undifferentiated "unauthorized."
Structured elaboration (approach)
- Extraction: read
Authorization: Bearer <token>; a missing or malformed header is rejected before any verification work happens at all. - Key resolution with caching: use a JWKS client (
jwks-rsabelow) configured withcache: true, acacheMaxAge, andrateLimit: truewith a requests-per-minute ceiling. This matters specifically because a burst of requests can present several differentkidvalues in quick succession right after the identity provider rotates keys; without a cache, that burst becomes a burst of outbound calls to the JWKS endpoint, turning someone else's infrastructure into a dependency of your hot path. - Verification:
algorithms: ['RS256']is passed explicitly, the same defense used against algorithm-confusion attacks as in a Python implementation of the same idea, plusaudienceandissuerchecks and aclockTolerancefor clock-skew. - Error mapping:
TokenExpiredErrorbecomestoken_expired;JsonWebTokenError(which covers a malformed token and a bad signature) becomesinvalid_token; anything else, in practice the JWKS client's own "no signing key found for this kid" error, becomesunauthorized.
Worked example (executed)
"use strict";
const http = require("http");
const crypto = require("crypto");
const jwt = require("jsonwebtoken");
const jwksClient = require("jwks-rsa");
// npm install jsonwebtoken jwks-rsa
const EXPECTED_ISSUER = "https://idp.example.com";
const EXPECTED_AUDIENCE = "orders-api";
function requireBearerJwt(client) {
function getKey(header, callback) {
client.getSigningKey(header.kid, (err, key) => {
if (err) return callback(err);
callback(null, key.getPublicKey());
});
}
return function (req, res, next) {
const authHeader = req.headers["authorization"] || "";
const match = /^Bearer (.+)$/.exec(authHeader);
if (!match) {
return res.status(401).json({ error: "missing_bearer_token" });
}
const token = match[1];
jwt.verify(
token,
getKey,
{
algorithms: ["RS256"],
audience: EXPECTED_AUDIENCE,
issuer: EXPECTED_ISSUER,
clockTolerance: 30,
},
(err, decoded) => {
if (err) {
if (err.name === "TokenExpiredError") {
return res.status(401).json({ error: "token_expired" });
}
if (err.name === "JsonWebTokenError") {
return res.status(401).json({ error: "invalid_token", detail: err.message });
}
return res.status(401).json({ error: "unauthorized", detail: err.message });
}
req.user = decoded;
next();
}
);
};
}
// --- test harness: real RSA keypair, real local JWKS HTTP server, 5 cases ---
function buildKeypairAndJwks() {
const { publicKey, privateKey } = crypto.generateKeyPairSync("rsa", { modulusLength: 2048 });
const jwk = publicKey.export({ format: "jwk" });
jwk.kid = "test-key-1";
jwk.alg = "RS256";
jwk.use = "sig";
return { privatePem: privateKey.export({ format: "pem", type: "pkcs8" }), jwksDoc: { keys: [jwk] } };
}
function startJwksServer(jwksDoc) {
const body = JSON.stringify(jwksDoc);
const server = http.createServer((req, res) => {
res.writeHead(200, { "Content-Type": "application/json" });
res.end(body);
});
return new Promise((resolve) => {
server.listen(0, "127.0.0.1", () => {
const { port } = server.address();
resolve({ server, jwksUrl: `http://127.0.0.1:${port}/jwks.json` });
});
});
}
async function runCase(middleware, authHeaderValue) {
const req = { headers: { authorization: authHeaderValue } };
let nextCalled = false;
const settled = new Promise((resolve) => {
const res = {
_status: 200, _body: null,
status(code) { this._status = code; return this; },
json(body) { this._body = body; resolve(); return this; },
};
req._res = res;
middleware(req, res, () => { nextCalled = true; resolve(); });
});
await settled;
const res = req._res;
return { nextCalled, status: res._status, body: res._body, user: req.user };
}
async function main() {
const { privatePem, jwksDoc } = buildKeypairAndJwks();
const { server, jwksUrl } = await startJwksServer(jwksDoc);
const client = jwksClient({
jwksUri: jwksUrl, cache: true, cacheMaxEntries: 5,
cacheMaxAge: 10 * 60 * 1000, rateLimit: true, jwksRequestsPerMinute: 10,
});
const middleware = requireBearerJwt(client);
const now = Math.floor(Date.now() / 1000);
const basePayload = { sub: "user-42", iss: EXPECTED_ISSUER, aud: EXPECTED_AUDIENCE };
const validToken = jwt.sign(basePayload, privatePem, { algorithm: "RS256", keyid: "test-key-1", expiresIn: 300 });
const expiredToken = jwt.sign({ ...basePayload, iat: now - 1000, exp: now - 700 }, privatePem,
{ algorithm: "RS256", keyid: "test-key-1" });
const unknownKidToken = jwt.sign(basePayload, privatePem, { algorithm: "RS256", keyid: "does-not-exist", expiresIn: 300 });
const malformedToken = "not-a-jwt-at-all";
const results = [];
const r1 = await runCase(middleware, `Bearer ${validToken}`);
results.push(["valid token calls next() with req.user", r1.nextCalled && r1.user && r1.user.sub === "user-42"]);
const r2 = await runCase(middleware, `Bearer ${expiredToken}`);
results.push(["expired token -> 401 token_expired", !r2.nextCalled && r2.status === 401 && r2.body.error === "token_expired"]);
const r3 = await runCase(middleware, `Bearer ${malformedToken}`);
results.push(["malformed token -> 401", !r3.nextCalled && r3.status === 401]);
const r4 = await runCase(middleware, `Bearer ${unknownKidToken}`);
results.push(["unknown kid -> 401", !r4.nextCalled && r4.status === 401]);
const r5 = await runCase(middleware, "");
results.push(["missing bearer header -> 401 missing_bearer_token", !r5.nextCalled && r5.status === 401 && r5.body.error === "missing_bearer_token"]);
let allPass = true;
for (const [name, passed] of results) {
console.log(`[${passed ? "PASS" : "FAIL"}] ${name}`);
if (!passed) allPass = false;
}
console.log(`\nALL_PASS=${allPass}`);
server.close();
}
main();
Output, from an actual run (node verify_jwt_middleware.js):
[PASS] valid token calls next() with req.user
[PASS] expired token -> 401 token_expired
[PASS] malformed token -> 401
[PASS] unknown kid -> 401
[PASS] missing bearer header -> 401 missing_bearer_token
ALL_PASS=true
Complexity and edge cases
Excluding a JWKS cache miss (one HTTP round trip), each request costs one RSA signature verification and a handful of claim comparisons, O(1) in the number of requests. The three named error cases plus a missing-header case are exercised directly above: expired token, malformed token, and unknown kid all resolve to a 401 with a distinct error code rather than an unhandled exception or a generic 500, and each is tested against a real signature so the failure is genuine, not simulated by mocking the verification step itself.
Trade-offs and pitfalls
- Don't collapse "the credential is missing" and "the credential is invalid" into identical log output; both end up as a
401, but a client integration bug (forgetting the header) looks nothing like a compromised or expired token in an incident review, and conflating them in logs erases that distinction. - The JWKS client's cache has a maximum entry count; under a fleet-wide key rotation with many keys active simultaneously, a low limit can evict entries that are still legitimately in use, causing avoidable cache misses. Size it against how many keys the identity provider actually keeps active concurrently, not against a copied default.
getKeybeing asynchronous means a slow, uncached key lookup adds real latency to that one request, not to the whole server, but "not blocking everyone else" is a much lower bar than "fast." The cache configuration is what actually keeps the common case fast, not the mere fact that the lookup is async.
Unlock Full Question Bank
Get access to all Identity, Authentication, and Access Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.