Identity, Authentication, and Access Management Questions
Designing and operating identity and access control systems. Covers authentication protocols and standards (OAuth, SAML, OIDC, MFA), authorization models (RBAC, ABAC), identity lifecycle and privilege management, IAM architecture and automation, and access control across cloud and on-premises environments. The 'who can do what' control plane, distinct from cryptographic key management.
You must onboard external partners with SAML or OIDC federation. Draft a federation onboarding checklist covering metadata exchange, certificate validation, required attributes, scopes/claims, test cases, operational contacts, and trust lifecycle management including periodic validation and revocation procedures.
Sample Answer
Direct answer
A federation onboarding checklist turns "add a new SAML or OpenID Connect (OIDC) partner" from an ad hoc integration exercise into a repeatable process with an explicit go-live gate. Exchange and verify both sides' metadata over a trusted channel, validate the actual signing certificate rather than take it on faith, agree the exact attributes and scopes or claims that will flow before the first real login, prove the integration against a written set of test cases rather than a single successful try, record specific people to call on both sides when something breaks, and treat the trust relationship itself as something with an ongoing lifecycle, periodically re-checked and revocable, rather than a one-time setup nobody looks at again once it works.
Structured elaboration
| Checklist item | What it covers |
|---|---|
| Metadata exchange | Exchange each side's federation metadata (entity ID, endpoints, supported bindings, signing certificate) over an authenticated, verified channel, never an unverified email attachment, and confirm both sides are configured against the current metadata rather than a stale copy from an earlier draft of the integration |
| Certificate validation | Verify the signing certificate's actual fingerprint out-of-band, a phone call or a separately verified channel, rather than trusting whatever arrived in the metadata file itself; confirm its expiry date and calendar a renewal reminder well ahead of it; and validate the certificate chain if the partner's certificate is issued by an intermediate certificate authority |
| Required attributes | Agree in writing, before go-live, exactly which attributes the partner will send and their expected format, map them onto your internal canonical schema, and identify any genuinely missing required attribute before it surfaces as a production login failure |
| Scopes and claims | For OIDC specifically, agree the exact scopes being requested and the claims returned for each, resisting the temptation to request a broader scope "just in case," since the scope negotiation itself deserves the same least-privilege discipline as any other access grant |
| Test cases | A written set of scenarios run before go-live: a successful login with a valid test account, an attempt using an expired or near-expiry certificate that should fail, a login missing a required attribute that should fail, a login from a deactivated test account that should fail, and the logout or session-termination flow if the partner supports one |
| Operational contacts | A named technical or security contact on each side, not a generic support inbox, covering at minimum an urgent security issue, a routine maintenance heads-up, and certificate-rotation coordination, stored alongside the trust record itself |
| Trust lifecycle management and periodic validation | A recurring re-validation of the whole relationship on a defined schedule: confirming the contact list is still accurate, confirming the certificate hasn't changed outside the agreed process, and confirming the partner still needs the level of access originally granted |
| Revocation procedures | A documented, rehearsed process for immediately disabling a partner's federation trust: who has the authority to invoke it, how quickly it actually takes effect, and what happens to that partner's legitimate users the moment trust is revoked |
Worked example
Onboarding "BrightPath Logistics" via SAML federation into a shared shipping portal, tracked as a completed checklist:
| Checklist item | Outcome for BrightPath |
|---|---|
| Metadata exchange | Metadata retrieved over a mutually authenticated channel from BrightPath's published federation endpoint, confirmed by both teams to be the current version dated the same week as onboarding |
| Certificate validation | Fingerprint confirmed by phone with BrightPath's identity team; certificate expires in 18 months, and a renewal reminder is calendared 60 days ahead of that date |
| Required attributes | Agreed set: employee_id, department, shipping_region; BrightPath's identity provider (IdP) initially omits shipping_region from its assertion, caught during this step and fixed before any test login was attempted |
| Scopes and claims | Scope limited to portal:read and shipment:read; BrightPath's initial request also asked for shipment:write, which the portal team declined since no BrightPath workflow in scope for this onboarding actually needs to create or modify shipments |
| Test cases | Five scenarios run and passed: valid test-account login, expired-certificate login correctly rejected, login missing shipping_region correctly rejected per the agreed required-attribute list, deactivated test-account login correctly rejected, and single-logout correctly terminating the session on both sides |
| Operational contacts | BrightPath's identity lead and the portal team's on-call security contact are recorded directly in the trust record, with a rotation-coordination contact listed separately from the incident contact |
| Trust lifecycle management | Scheduled for a semi-annual review; the first review date is set six months from go-live |
| Revocation procedures | A tested, config-level flag exists to disable BrightPath's trust within minutes if needed; the fallback experience for BrightPath's users during a revocation is a clear error message directing them to BrightPath's own support, agreed in advance rather than left undefined |
The one real gap this process caught, the missing shipping_region attribute, would otherwise have surfaced as a confusing production failure the first time a BrightPath user's login succeeded but the portal couldn't determine which shipping region to show them.
Trade-offs and pitfalls
Skipping out-of-band verification of the certificate fingerprint and trusting whatever arrives in the metadata file is a real risk, not a formality: if the metadata exchange channel itself is compromised, a substituted certificate would pass every check that only looks at the file contents themselves. Treating this checklist as a one-time gate at onboarding, rather than a recurring lifecycle, is how stale trust accumulates: the contact who was correct on day one moves teams, the certificate is quietly rotated on the partner's side without the agreed process being followed, or the partner no longer actually needs the access originally granted, and none of that is caught without a scheduled review. Under-specifying revocation procedures until the day they're actually needed, in the middle of a live security incident, is a costly and avoidable gap; the process needs to be pre-tested during calm conditions, not improvised under pressure. Finally, requesting broader scopes than the current integration needs "to avoid asking again later" quietly defeats least privilege and widens the blast radius if that partner's own systems are ever compromised; asking again later, when there's an actual need, is a small cost compared to that risk.
Design a secure cross-domain SSO architecture that minimizes trust exposure: token exchange patterns, limiting claims and scopes, short token lifetimes, partner-scoped client credentials, automated metadata validation, and contractual/operational controls for partner access.
Sample Answer
Direct answer
A cross-domain SSO integration with an external partner should minimize trust by exchanging a narrowly scoped, short-lived token at the trust boundary rather than propagating the partner's own broad internal token inward, limiting the claims and scopes that token carries to exactly what the integration needs, and validating the partner's federation metadata automatically rather than trusting a one-time manual configuration. Because a partner is, by definition, outside the organization's own security control, every technical control here (token exchange, scope limiting, short lifetimes, partner-scoped credentials, metadata validation) needs a matching contractual and operational control, an agreement defining what the partner is responsible for, and a way to detect when they fail to meet it, since technical controls alone cannot compel a partner's own security practices.
Structured elaboration
Token exchange patterns
- Never let a partner's own identity-provider-issued token, whatever its native claim set, flow directly into internal systems. Use a token exchange step (the OAuth 2.0 Token Exchange pattern is the standard shape) at the trust boundary: the partner authenticates using their own credentials, and a boundary service MINTS a new, internally scoped token carrying only the claims internal systems need, rather than passing the partner's original token through unmodified.
- This boundary step is the single place a compromised or overly broad partner-issued token gets reduced down to exactly what's needed; without it, a misconfiguration on the partner's own identity provider (an overly broad claim set on their side) becomes your internal systems' problem directly.
Limiting claims and scopes
- Define the narrowest possible scope for a partner integration, read-only access to a specific resource type, not "whatever the partner's own users can do", and enforce it at the token-exchange boundary above, not as a downstream service's own responsibility to filter after the fact.
- Claims minimization applies symmetrically: the token issued to the partner-facing boundary should carry only what a downstream service actually needs (partner organization id, granted scope, expiry), not a full internal user-profile shape reused for convenience.
Short token lifetimes
- Partner-facing tokens should have a materially shorter lifetime than internal user tokens, minutes, not hours, since the entire point is bounding the exposure window if the partner's own systems are compromised. A short-lived token that must be refreshed frequently is deliberately more operationally demanding than a long-lived one; that friction is the security property, not a bug to engineer away.
- Pair short access-token lifetimes with server-side-only refresh handling at the boundary, never handing a long-lived refresh token to the partner's own client code, so the partner's ongoing access depends on the boundary service's own health and policy, not on a credential the partner holds indefinitely.
Partner-scoped client credentials
- Issue a distinct client id and credential PER PARTNER, never a shared credential across partners, so one partner's credential compromise or offboarding can be revoked without affecting any other partner's integration.
- Rotate partner credentials on a defined schedule and immediately on any signal of compromise, treating rotation as a routine operational capability that must work quickly, not a rare emergency procedure improvised under pressure.
Automated metadata validation
- Fetch and validate federation metadata, the partner identity provider's signing keys, supported endpoints, certificate expiry, automatically on a schedule, not once manually with the assumption it stays correct forever. A partner rotating their own signing key without your side automatically picking up the change is a common, entirely preventable outage and, if handled by silently falling back to an old key, a security risk.
- Validate the metadata's own signature and certificate chain, and alert on unexpected changes, a new, unrecognized signing key appearing without an expected rotation notice, rather than silently accepting any update, since an attacker capable of tampering with metadata delivery could otherwise substitute their own key.
Contractual and operational controls for partner access
- Technical controls alone cannot compel a partner's own security practices; the integration agreement should specify the partner's own obligations, credential handling, breach-notification timelines, minimum security requirements, as an enforceable backstop for exactly the failure modes technical controls can't directly prevent.
- Monitor partner-scoped token usage for anomalies, a partner's traffic suddenly spiking, or accessing resource types outside their normal pattern, as its own dedicated alerting category, since a compromised partner integration often looks, from the partner-scoped credential's own perspective, like completely valid, correctly authorized traffic.
Worked example
A concrete claims-reduction trace at the token-exchange boundary:
- The partner authenticates against their own identity provider and receives a partner-issued token with claims
{partner_user_id: "u-4471", partner_roles: ["admin", "billing", "support"]}, a broad claim set reflecting the partner's own internal role structure. - The boundary service validates that token's signature against automatically fetched partner metadata, and separately checks that the calling client id is this specific partner's own scoped credential, not a generic shared one.
- The boundary service mints an internal token containing ONLY
{partner_org_id: "acme-partner", scope: "read:shipment-status", exp: now + 5min}, explicitly droppingpartner_user_idandpartner_rolesentirely, since internal systems never need to know the partner's internal role structure, only the scope this integration was actually granted.
Internal services downstream of the boundary never see, and cannot depend on, the partner's own internal roles at all, exactly the reduction that limits the blast radius of anything going wrong on the partner's own identity-provider side.
Trade-offs and pitfalls
- The token-exchange boundary is an additional service and an additional point of operational responsibility; skipping it, letting partner tokens flow through directly "just this once" for a fast integration, is exactly how claims-minimization discipline erodes over time, since removing an already-shipped shortcut is much harder than never shipping it.
- Short token lifetimes push refresh-failure handling onto the boundary service; if it goes down, every partner integration depending on it fails closed simultaneously, the correct default for a trust-boundary component, but it means the boundary service needs the same high-availability discipline as any other critical-path authentication component.
- Automated metadata validation that alerts with equal urgency on every legitimate scheduled key rotation and every real attack attempt trains operators to ignore the alert; distinguish an announced, scheduled rotation from an unannounced key change so the alert that matters doesn't get lost in routine noise.
- Contractual controls are only as good as their enforcement; an agreement specifying a partner's breach-notification timeline is worthless without the operational monitoring above to catch a breach the partner themselves doesn't notice or doesn't disclose promptly.
Compare common Multi-Factor Authentication (MFA) approaches : TOTP (time-based OTP), SMS OTP, push-based approval, and hardware-backed/U2F/WebAuthn tokens : in terms of security, usability, deployability, and attack surface. For each method, list typical threats (e.g., SIM swapping, phishing, device theft) and describe when you would choose or avoid that method for a user-facing application.
Sample Answer
Direct answer
The four common multi-factor authentication (MFA, proving identity with more than one independent factor) methods trade off along the same two axes: how resistant the method is to phishing, and how much friction and cost it adds. Time-based one-time password (TOTP) apps and hardware-backed passkeys (WebAuthn/FIDO2) sit at the strong end, SMS one-time passwords sit at the weak end because the delivery channel itself can be hijacked independent of anything the user does wrong, and push-based approval sits in between: easy to use, but vulnerable to a specific social-engineering pattern (repeatedly prompting the user until they tap approve by habit or fatigue) that neither of the code-based methods share.
Structured elaboration
| Method | Security (phishing resistance) | Usability | Deployability | Attack surface / typical threats |
|---|---|---|---|---|
| SMS one-time password (OTP) | Weakest: the delivery channel itself can be subverted independent of the user | Highest: no app required, universally understood | Depends on telecom SMS gateways; cost and delivery reliability vary by region | SIM swapping (a carrier is socially engineered into porting the victim's number), interception at the telecom-network level, real-time phishing relay of the code |
| TOTP (authenticator app) | Good: the code itself is never transmitted over a network channel an attacker can pass through | High: requires installing and checking an app, minor typing friction | Cheap, standards-based, works offline once enrolled | Real-time phishing relay (a fake login page that immediately forwards the code the user typed), theft of the enrollment secret from a compromised device or backup |
| Push-based approval | Medium: removes manual code entry, but the approval action itself can be induced | Highest of the code/prompt-based methods: one tap | Requires the vendor's own app and network connectivity; not a cross-vendor standard | MFA fatigue or prompt bombing (sending repeated approval requests until the user taps approve out of habit or annoyance), device theft if the device is unlocked |
| Hardware-backed / WebAuthn (FIDO2) | Strongest: cryptographically bound to the site's own origin, so a look-alike phishing domain simply cannot obtain a valid signature | High once enrolled (tap or biometric), but requires a compatible key or platform authenticator | Hardware cost, and enrollment/recovery process complexity if a user's only authenticator is lost | Physical theft of the token (mitigated by requiring a PIN or biometric on the key itself), gaps in the recovery process |
Why the phishing-resistance ranking holds. SMS and TOTP both ultimately depend on the user (or an attacker impersonating the site) having a code that a phishing page can capture and immediately relay to the real site in real time (an adversary-in-the-middle relay); TOTP is still meaningfully better than SMS because it removes the telecom-layer interception risk (SIM swapping, network-level interception) that has nothing to do with the user's own behavior at all. Push notifications remove the "type a code" step but introduce a different failure mode: repeated, low-friction approval prompts that a user can eventually tap through without reading. WebAuthn is qualitatively different, not just incrementally better, because the cryptographic protocol itself checks the requesting site's origin before it will produce a valid signature, so the phishing page cannot obtain a usable credential regardless of how convincing it looks to the human.
When to choose or avoid each, for a user-facing application. TOTP is a strong, low-cost default for a broad consumer audience: free to implement, no telecom dependency, and meaningfully better than SMS for a modest amount of added friction. SMS OTP should be avoided as the only factor for anything of real value; it is best reserved for a recovery or fallback path for users without a smartphone, not the primary method, since its weaknesses live in infrastructure the application does not control. Push-based approval fits an enterprise or internal workforce application where the user population is known and already carries a managed device; it should be paired with a context check (showing the requesting device, location, or a number the user must match, rather than a bare "approve or deny" prompt) specifically to blunt prompt-bombing. Hardware-backed WebAuthn is the right default for privileged or high-value accounts (administrators, executives, anyone likely to be individually targeted), where phishing resistance matters enough to justify the enrollment friction and hardware cost, even if it is not yet practical to require for every user in a large consumer base on day one.
Worked example
A consumer web application decides its MFA policy by user tier rather than one policy for everyone: ordinary users are offered TOTP as the default second factor (cheap to support, meaningfully better than nothing, and better than SMS) with SMS OTP available only as an account-recovery fallback for a user who cannot install an authenticator app. Administrative accounts with access to production data or billing are required to enroll a hardware-backed WebAuthn key, since those accounts are the ones most likely to be individually targeted with a convincing spear-phishing attempt, and the origin-binding property of WebAuthn is what actually stops that attack, not merely the account holder's own vigilance. Internal support staff, who already carry a company-managed phone, use push-based approval with number matching (the login page displays a two-digit number the user must enter into the push prompt), specifically to close the fatigue-attack path a bare "approve/deny" prompt would leave open.
Trade-offs and pitfalls
The recurring pitfall is treating "we require MFA" as a single fact rather than a spectrum: an application that lets every account, including administrators, satisfy MFA with SMS alone has not meaningfully raised the bar against a targeted attacker willing to attempt a SIM swap, even though it can honestly claim MFA is enabled. A second pitfall is deploying push-based approval without any context or number-matching step; a bare approve/deny prompt is exactly what makes prompt-bombing effective, since the user has no information to distinguish a legitimate login attempt from an attacker's repeated requests. A third pitfall is over-indexing on WebAuthn's security strength while under-investing in its recovery flow: a user whose only hardware key is lost or broken needs a well-designed, equally secure recovery path, or the strongest method in the table becomes the one most likely to lock a legitimate user out.
Design note: cloud console access versus service identities. The comparison above assumes a human is present to complete an interactive challenge. That assumption does not hold for automated API calls made by a service identity (a machine or workload credential, not a person), which cannot tap a push prompt or read a TOTP code. The right policy split is to require MFA, ideally hardware-backed, for interactive console access by human administrators, while protecting service identities through mechanisms built for non-interactive use instead: short-lived, automatically rotated credentials, or workload identity federation that lets a service prove who it is without a standing long-lived secret at all. Treating "no interactive MFA" as a gap to fill with a weaker human-facing method (like requiring a service account to somehow "complete" SMS OTP) is the wrong instinct; the equivalent protection for a service identity is a fundamentally different, non-interactive credential lifecycle.
You must integrate on-prem Active Directory with a cloud IdP to support SSO for cloud services and legacy apps. Describe the architecture patterns for directory synchronization versus federation, including security trade-offs (password hash sync vs pass-through auth vs federation), account provenance, how to synchronize groups and nested groups, and how to handle password policy differences.
Sample Answer
Direct answer
There are two fundamentally different architecture patterns for connecting on-premises Active Directory (AD, Microsoft's on-prem directory service for user, group, and computer objects) to a cloud identity provider (IdP): directory synchronization, which copies user and group objects into the cloud directory so authentication happens entirely there, and federation, which keeps authentication on-premises and has the cloud IdP redirect every login back to an on-prem federation server for a signed assertion. Synchronization itself splits into two sign-in modes, password hash sync and pass-through authentication, each trading availability against how much credential material ever leaves the premises. Getting this right also means tracking where each identity actually originated, correctly expanding nested group membership during sync, and reconciling two systems' password policies so a user is never told their password is valid by one and rejected by the other.
Structured elaboration
Architecture patterns: synchronization versus federation. In synchronization, a sync agent running on-premises periodically pushes user and group objects from AD into the cloud directory; once synced, the cloud IdP is authoritative for its own sign-ins, and legacy on-prem apps that authenticate directly against local AD via SAML or Kerberos keep working unmodified since only the cloud-facing identity gets copied out. In federation, the cloud IdP stores no validated credential at all: every sign-in is redirected, via SAML or a similar protocol, to an on-prem federation server that authenticates the user against live AD and hands back a signed assertion the cloud IdP trusts. The architectural difference that matters: synchronization makes the cloud directory authoritative over synced data, while federation keeps on-premises authoritative and makes every single cloud login dependent on the on-prem federation server's availability.
Security trade-offs: password hash sync vs. pass-through authentication vs. federation. Password hash sync (PHS) sends AD's password hash, already one-way hashed, re-hashed again for cloud storage, to the cloud directory, which then validates sign-ins entirely on its own; cloud services keep working even if the on-premises network is unreachable. The cost: an on-prem password change, account lockout, or disablement can lag behind the actual state until the next sync cycle, and hash material now exists in two systems instead of one, giving a cloud-side breach something to attack that would not exist under the other two patterns. Pass-through authentication (PTA) has a lightweight on-prem agent validate each cloud sign-in against live AD in real time, so no password hash ever leaves the premises, but cloud sign-in now depends on that agent and its network path being available; an on-prem outage takes cloud sign-in down with it. Federation stores no credential material in the cloud at all and instead trusts a signed assertion from an on-prem federation server, which gives the strongest "nothing leaves on-prem" guarantee but the heaviest operational load (federation servers, their certificates, and their own high-availability design) and the hardest on-prem dependency of the three. In short: PHS optimizes for availability at the cost of hash material existing in two places; PTA and federation optimize for credential material never leaving on-premises at the cost of making on-premises a hard dependency for every cloud login.
Account provenance. For every identity in the cloud directory, you need to know whether it originated on-premises (synced from AD) or was created natively in the cloud, and govern the two differently. A synced object should be treated as effectively read-only in the cloud for anything AD already owns, password state, group membership, disabled status, because editing it cloud-side gets silently overwritten by the next sync cycle or produces a split-brain state where the two systems disagree about the same account. The practical mechanism is tagging every synced object with an immutable source-anchor value derived from its on-prem object identifier, and building offboarding and access-review automation around that tag, so disabling a user in AD reliably disables their cloud access on the next sync, and so a cloud-native account (a contractor or partner provisioned directly in the cloud, never in AD) is never mistaken for an AD-governed one during an audit.
Synchronizing groups and nested groups. Syncing only a user's direct group memberships silently drops access for anyone whose permission actually comes through a group nested two or three levels deep, which is routine in a large AD forest built up over years. The sync agent has to expand nested membership during sync, or the receiving cloud directory has to evaluate nested groups natively, or the migration quietly breaks exactly the access it was meant to preserve. Deep or circular nesting is a genuine operational hazard on top of that: cap how many levels of nesting the sync will expand, and audit periodically for nesting cycles, because a naive recursive expansion can either loop indefinitely or produce a flattened membership list large enough to exceed the cloud directory's per-object limits.
Handling password policy differences. On-prem AD's password policy (complexity rules, lockout thresholds, password history) and the cloud IdP's own default policy will not automatically agree, and under password hash sync or pass-through authentication, letting the cloud enforce a second, conflicting policy on the same credential is how users end up told their password is fine by one system and rejected by the other. The working pattern is to make on-prem AD the single source of truth for password policy in a hybrid design, and either disable the cloud directory's native password-policy enforcement for synced accounts, or configure it to mirror AD's rules exactly. Anywhere the cloud IdP legitimately needs a stronger control than on-prem enforces, most commonly multi-factor authentication (MFA, requiring a second proof of identity beyond the password) on risky sign-ins, that control should layer on top of the existing password check rather than replace or duplicate it, so the two systems compose instead of contradicting each other.
Worked example
"Northwind," a company with a three-domain on-prem AD forest built over a decade of acquisitions, adopts password hash sync for cloud sign-in and federation-free simplicity, since its cloud services need to stay available even during on-prem maintenance windows.
flowchart TB
subgraph Sync["Synchronization: password hash sync or pass-through auth"]
AD1[On-prem Active Directory]
Agent[Sync agent]
CloudDir[Cloud directory]
AD1 -->|sync objects, hash or live check| Agent
Agent --> CloudDir
end
subgraph Fed["Federation, the alternative Northwind rejected"]
AD2[On-prem Active Directory]
FS[On-prem federation server]
CloudIdP[Cloud IdP, trusts assertion only]
CloudIdP -->|redirect| FS
FS --> AD2
AD2 -->|signed assertion| FS
FS -->|assertion| CloudIdP
end
During sync setup, Northwind discovers its "Finance-AllAccess" group is nested four levels deep (Finance-AllAccess contains Finance-Regional-Leads, which contains Finance-EMEA, which contains Finance-EMEA-Payables, and the individual users actually sit in that innermost group). A flat, direct-membership-only sync would have shown zero members of Finance-AllAccess in the cloud directory, silently breaking access for every finance analyst whose permission depended on that chain. Northwind's sync agent is configured to expand nested membership up to five levels and alert if it ever detects a cycle. Every synced user object carries a source-anchor value tied to its AD object identifier, so when a departing employee is disabled in AD, the next sync cycle disables their cloud access automatically, and a security review can immediately tell that account apart from the twelve contractor accounts Northwind provisioned directly in the cloud IdP, which have no AD source anchor at all. Finally, Northwind finds its on-prem policy requires 14-character passwords with no reuse of the last 24, while the cloud IdP's default policy allows 8-character passwords; rather than let the cloud enforce its weaker default (which would let a user set a password AD's own policy would have rejected) or its own separate stricter rule (which could reject a password AD already accepted), Northwind disables the cloud directory's native password-policy checks for every synced account and leaves AD as the sole authority, layering step-up MFA in the cloud IdP only for sign-ins flagged as high-risk.
Trade-offs and pitfalls
Choosing password hash sync purely for its availability benefit, without accounting for the sync interval, means an account disabled or locked out on-prem can still authenticate successfully in the cloud for as long as one sync cycle, a gap that matters in an active offboarding or compromise scenario and needs its own compensating control (a fast, on-demand sync trigger for exactly those events) rather than being accepted silently. Federation's strongest selling point, that no credential material ever reaches the cloud, is also its biggest operational liability: it makes every cloud sign-in depend on an on-prem service that now needs its own redundancy, certificate rotation, and monitoring, and an outage there takes down cloud access even though nothing in the cloud itself failed. The most common nested-group pitfall is discovering the broken access only after go-live, because a small pilot group rarely reaches the deeper end of an old AD forest's nesting; the fix is to explicitly test the sync against the forest's actual deepest nesting chains, not just a handful of well-behaved top-level groups. On password policy, a frequent wrong turn is letting both systems enforce independently "for defense in depth," which sounds safer but produces exactly the contradictory-rejection experience the single-source-of-truth pattern above is designed to prevent; additional strength belongs in an additional control like MFA, not in a second, uncoordinated password policy.
Design a cryptographic key management and signing infrastructure for tokens (JWT/SAML) that supports key rotation, HSM-backed storage, cross-region replication, graceful rollover (support old keys for token lifetime), and fast compromise recovery. Describe key metadata (kid/version), rotation cadence, signer/verifier patterns, publishing of public keys (JWKS), and how services discover and cache key material securely.
Sample Answer
Direct answer
Design this as three separable concerns. First, where the private key material physically lives and who can invoke it: HSM-backed (hardware security module), never exported in the clear. Second, how the public half reaches every verifier: a JWKS (JSON Web Key Set) endpoint, fetched and cached, keyed by kid (key ID). Third, a rotation lifecycle that always keeps at least one still-valid old key published alongside a new one, so tokens signed before a rotation remain verifiable until they naturally expire. Treat compromise recovery as the same rotation machinery run on an emergency timeline, not a separate system.
Structured elaboration
HSM-backed storage. Private signing keys are generated inside, and never leave, a hardware security module or an equivalent cloud KMS (key management service) with HSM-backed key material. The signing service asks the HSM to perform the sign operation and gets back only the signature, never the key itself. This bounds the blast radius of a host compromise: an attacker who compromises the signing service's host can invoke signing operations only while they retain that access, but cannot exfiltrate the private key to use elsewhere or later. Every sign invocation should be logged through the HSM or KMS's own audit trail, since that log is exactly what lets you scope a future incident quickly.
Cross-region replication. Both signing capability and public-key material need to reach every region that issues or verifies tokens.
- For signing, either (a) use a managed KMS's multi-region key replication if the provider offers a mature one, so any region's issuer can sign with what is logically the same key, or (b) run independent region-local signing keys, each with its own
kid, all published to one shared JWKS. Option (b) avoids depending on cross-region KMS replication maturity, at the cost of a slightly larger key set to track, and is a defensible default absent strong confidence in option (a). - For verification, JWKS content is public and non-sensitive, so it replicates trivially behind a CDN (content delivery network) or globally-replicated object storage.
Graceful rollover. Never flip the signer to a brand-new key the instant it's generated. Publish the new public key to JWKS first, and wait long enough for verifiers to have fetched and cached it (bounded by the JWKS cache TTL, time-to-live, plus a safety margin) before switching the signer over. Keep the old public key published and accepted for at least the longest outstanding token lifetime in the system before removing it.
Fast compromise recovery. The same rotation primitive, run without the soak period:
- Generate and publish a new key immediately, and force verifiers to pick it up right away rather than waiting on the normal cache TTL, for example via a push notification or a short-lived emergency flag they poll.
- Flip signing to the new key immediately.
- Remove the compromised key from JWKS immediately, accepting that some legitimately-issued, still-unexpired tokens signed under the compromised key will now fail verification. During an active compromise, rejecting legitimate sessions is the correct trade against accepting forged tokens.
- Force step-up re-authentication for the affected users or services, since their existing tokens are now invalid.
This means the design needs a "push" or fast-invalidate path for JWKS caching, in addition to the normal TTL-based caching used for routine rotation, purely to support the emergency case.
Key metadata (kid and version). Every key entry carries a kid (an opaque identifier; a scheme like {purpose}-{algorithm}-{sequence}, for example token-sign-es256-007, helps operators reading logs, even though the value itself doesn't need to encode meaning) plus lifecycle metadata tracked in the key-management system: status (pending, active-signing, active-verify-only, retired), created-at, activated-at, algorithm, and, for HSM-backed keys, a reference to the HSM key handle, never the raw key material.
Rotation cadence. Rotate routinely on a fixed schedule, commonly quarterly for high-value signing keys, sometimes monthly for higher-risk services. Bound it above by your tolerance for the blast radius of a slow-to-detect compromise, and below by your minimum safe soak-plus-overlap window: if the rotation cadence is shorter than the time it takes to safely roll a new key out and retire an old one, you never finish retiring old keys before starting the next rotation, and JWKS grows without bound.
Signer and verifier patterns. The signer is a narrow, privileged component, ideally the only thing with HSM or KMS sign permission, exposing a minimal interface. Verifiers are numerous, low-privilege, and only ever need read access to public key material, never signing access. Because most operations in this system are verifications, not signings, this asymmetry (few signers, many verifiers) is exactly why publishing public keys widely via JWKS is safe, and why centralized signing is the actual security boundary worth defending.
Publishing JWKS. The standard shape is a document like {"keys": [{"kid", "kty", "use", "alg", "n", "e" (for RSA), or "crv", "x", "y" (for elliptic curve)}]} served over TLS at a well-known, versioned URL. Multiple keys can coexist (old and new, during rollover), and a verifier selects by kid from the incoming token's header.
Discovery and secure caching. Verifiers fetch JWKS over TLS. Confidentiality of the fetch itself isn't the point, since the keys are public, but authenticity of the source matters: fetching over plaintext HTTP would let a machine-in-the-middle substitute a malicious key set. Cache with a TTL, and on encountering an unrecognized kid in an incoming token, trigger an immediate out-of-band re-fetch rather than waiting for the TTL, since that's exactly the signal that a rotation just happened. Rate-limit that "refetch on unknown kid" path, so an attacker sending tokens with garbage kid values can't use it to hammer the JWKS endpoint.
flowchart LR
HSM[HSM or cloud KMS: private signing key] --> Signer[Signer service]
Signer -->|publishes new public key| Publisher[JWKS publisher]
Publisher --> CDN[Public JWKS endpoint]
CDN -->|fetch and cache by kid| V1[Verifier: Service A]
CDN -->|fetch and cache by kid| V2[Verifier: Service B]
CDN -->|fetch and cache by kid| V3[Verifier: Service N]
Signer -->|signed token, kid = new| V1
Worked example
Say the current active key is kid=es256-003, and it's time for the scheduled quarterly rotation (an illustrative walkthrough, not a measurement). Day 0: generate es256-004 inside the HSM, and publish it to JWKS alongside es256-003, so JWKS now lists two keys. Verifiers cache JWKS for 10 minutes; to be safe, wait 30 minutes (comfortably above that TTL) before doing anything else. Day 0 + 30 minutes: switch the signer to sign new tokens with es256-004. Tokens already signed under es256-003 remain valid, and verifiers still have that key published. If access tokens live 1 hour, then by Day 0 + 1 hour 30 minutes (the 30-minute wait plus the 1-hour token lifetime), every token ever signed under es256-003 has expired, and it can safely be removed from JWKS.
Trade-offs and pitfalls
Storing private keys outside the HSM "just for this one debugging session" is the single most common way HSM guarantees get silently defeated.
Rotating on a cadence shorter than the soak-plus-overlap window needed means JWKS never actually shrinks, and stale keys accumulate untracked.
Relying solely on TTL-based JWKS caching, with no fast-invalidate path, means a key you've "revoked" during a compromise stays accepted by every verifier for up to a full cache TTL afterward.
Mixing region-local signing keys in some regions with true cross-region key replication in others, without a deliberate choice, creates asymmetric trust that's hard to reason about during an incident.
Skipping audit logging on HSM sign invocations removes the one thing that lets "fast compromise recovery" actually determine scope, namely which tokens were genuinely forged during the compromise window.
Unlock Full Question Bank
Get access to all Identity, Authentication, and Access Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.