Identity, Authentication, and Access Management Questions
Designing and operating identity and access control systems. Covers authentication protocols and standards (OAuth, SAML, OIDC, MFA), authorization models (RBAC, ABAC), identity lifecycle and privilege management, IAM architecture and automation, and access control across cloud and on-premises environments. The 'who can do what' control plane, distinct from cryptographic key management.
Design a continuous or adaptive authorization system that adjusts session trust in real time using signals such as device posture, IP reputation, behavioral anomalies, and recent authentication events. Explain how to represent and propagate trust level to microservices, how services should enforce it, and storage/caching strategies to keep decisions timely while auditable.
Sample Answer
Direct answer
Adaptive, continuous authorization computes a trust level, a tiered or continuous score rather than a single valid-or-invalid flag, from real-time signals: device posture (is the device compliant and up to date), IP reputation (is the source network associated with abuse), behavioral anomalies (does this request pattern match the user's own history), and the recency of authentication events (how long ago, and how strongly, did the user last prove who they are). That trust level is propagated to every microservice so each can make its own proportionate decision, for example allowing a read while requiring step-up authentication before a sensitive write, rather than one global allow-or-deny gate applied uniformly to the whole session.
Structured elaboration
Representing and propagating trust level to microservices. Represent trust as a small number of named tiers (say, high, medium, low) rather than a raw numeric score, carried as a signed claim alongside the access token, together with the individual signal values that produced that tier: device compliance status, an IP-reputation category, and minutes since the last strong authentication. The tiers give services a simple decision surface ("allow at medium or higher"), while the underlying signals let a service apply its own finer policy where the standard tiers do not fit its risk tolerance, such as a payments service requiring high trust specifically combined with multi-factor authentication (MFA) completed in the last five minutes. Because trust can change within a session's lifetime (a periodic device-posture check can flip from compliant to non-compliant; a threat-intelligence feed can flag the current network mid-session), it cannot simply be baked into the token at issuance and left static. It needs its own continuous-evaluation propagation path, similar in shape to the event-driven mechanism used for outright revocation elsewhere in identity systems, but carrying a graded score rather than a binary event.
How services should enforce it. Each microservice's own policy defines what trust tier, or specific signal combination, a given action requires, rather than one blanket threshold applied to the whole system. Low-risk actions may proceed at low or medium trust; high-risk actions require high trust and often an explicit, recent step-up authentication, not merely an inferred high score from ambient signals. This is risk-proportionate enforcement: the same session, at the same moment, can be permitted to do one thing and required to step up before doing another, because the trust requirement belongs to the action, not to the session as a whole. When enforcement rejects an action for insufficient trust, it should return a specific, actionable "step up required" response the client can act on, such as re-prompting for MFA, rather than a generic denial, so a legitimate user whose trust genuinely dipped (new network, aging authentication) has a clean path back rather than a confusing lockout.
Storage and caching strategies to keep decisions timely while auditable. Timeliness and auditability pull in different directions and need different storage. The low-latency path wants a small, fast cache of each session's current trust tier at the edge, updated by a central trust engine so individual microservices never each re-derive trust from raw signal sources on every request, which would be both slow (each service paying its own external-lookup cost) and inconsistent (two services could compute slightly different trust for the same session if the underlying signals changed between their two independent evaluations). The audit requirement wants the opposite: a durable, append-only, timestamped history of every trust-level change and which specific signal drove it, since a later investigation needs to reconstruct what the trust level actually was at a past moment, not only what it is now. Reconcile the two from a single update event: write the fast current-value cache and append to the durable history log from the same trust-recomputation event, so the audit write can lag by a small, bounded amount without ever blocking the real-time enforcement path, since the audit record is inherently retrospective rather than part of the live decision.
flowchart LR
Device[Device posture signal]
IP[IP reputation signal]
Behavior[Behavioral anomaly signal]
Auth[Recency of authentication signal]
Engine[Trust engine: recomputes tier from current signals]
EdgeCache[(Edge cache: current trust tier per session)]
History[(Append-only trust history log)]
SvcA[Service: view dashboard, low bar]
SvcB[Service: transfer funds, high bar + recent MFA]
Device --> Engine
IP --> Engine
Behavior --> Engine
Auth --> Engine
Engine -->|push update| EdgeCache
Engine -->|append change record| History
EdgeCache --> SvcA
EdgeCache --> SvcB
Worked example
A user logs in from their usual laptop and receives trust_level: high (compliant device, recognized network, MFA completed two minutes ago). That tier is written to the edge cache, keyed by session, and appended to the durable trust-history log with a timestamp.
Twenty minutes later, mid-session, the user's apparent network address changes to one a threat-intelligence feed flags as a VPN (virtual private network) exit node recently associated with account-takeover attempts elsewhere. The trust engine recomputes the session's tier down to low, pushes the update to the edge cache, and appends a history entry: {session_id, timestamp: 14:32:00, old_tier: "high", new_tier: "low", driving_signal: "ip_reputation_flag"}.
At 14:33, still under low trust, the same session requests view_dashboard, a low-risk action that service's own policy permits at low trust or above, so it succeeds without interruption; the design's proportionality means one negative signal does not lock the user out of everything. Shortly after, the session attempts transfer_funds, whose policy requires high trust combined with MFA within the last five minutes; enforcement reads the current cached tier (low, from 14:32) and rejects with a step-up-required response rather than a generic denial. The user completes a fresh MFA challenge; the trust engine recomputes from the full current signal set, not by simply overwriting the one signal that just changed, so if the flagged network is still in use, the recomputed tier may land at medium rather than snapping straight back to high, and the transfer proceeds only once the recomputed tier actually clears the service's specific bar.
Weeks later, an unrelated fraud review can query this session's trust-history log and reconstruct exactly when and why trust dropped, and how enforcement responded at that moment, entirely from the durable record, independent of whatever the session's trust level happens to be by the time anyone looks.
Trade-offs and pitfalls
- Collapsing this design back into one global trust threshold for the whole system defeats its entire purpose. A single blanket bar either over-restricts low-risk actions, annoying users with step-up prompts for a public dashboard, or under-protects high-risk ones, letting a borderline-trust session complete a sensitive transfer; the value is specifically in per-action, risk-proportionate thresholds.
- Letting each microservice independently re-derive trust from raw signal sources multiplies both latency and inconsistency. Every service pays its own external-lookup cost, and two services can legitimately compute different trust levels for the same session at the same moment if signals shifted between their separate evaluations; a shared, centrally-computed and cached assessment avoids both problems.
- Keeping only the current trust value, with no history, makes the exact investigation this system exists to support impossible: whether a specific access was appropriately gated at the time it happened. The append-only history is not a nice-to-have; it is what actually makes the design auditable, which the requirement explicitly asks for, not merely fast.
- Treating a successful step-up as fully restoring trust, rather than recomputing from the complete current signal set, can paper over a still-active negative signal, such as a still-suspicious network, simply because the user proved one different positive signal. Trust should be recomputed from everything currently known, not patched by overwriting just the signal that most recently changed.
- Returning a generic denial on a trust drop, instead of a specific step-up-required response, turns an ordinary, innocent situation (new network, aging authentication) into a confusing dead end, and tends to push users toward workarounds, such as disabling security features, rather than toward the clean recovery path the design is meant to offer.
Design a Continuous Access Evaluation (CAE) system that enables near-immediate revocation of access when credentials are compromised or roles change. Describe how change events are detected, how they are securely propagated to enforcement points (API gateways, microservices, mobile clients), options for push vs pull invalidation, securing the propagation channel, and methods to minimize latency while scaling to tens of thousands of active sessions.
Sample Answer
Direct answer
Continuous Access Evaluation (CAE) closes the gap a normal short-lived-token model leaves open: the window between a credential being compromised or a role changing, and the token that still reflects the old state finally expiring on its own. The design has three moving parts: detect a small, deliberately limited set of critical change events at their natural source, propagate those events to every enforcement point through a channel that is fast without becoming a bottleneck or a single point of failure, and enforce by revoking or re-evaluating the affected session before its token would otherwise have expired.
Structured elaboration
The diagram below shows the shape the rest of this answer explains: change events originate at their natural sources, flow through one signed event bus, and reach enforcement points primarily by push, with periodic reconciliation as the backstop for anything a push missed.
flowchart LR
IdP[Identity provider: password/MFA change, disable]
Risk[Risk engine: fraud/compromise signal]
AuthZ[Authorization system: role/privilege change]
Bus[[Signed critical-event bus]]
Gateway[API gateway: local revocation cache]
Micro[Microservice: local revocation cache]
Mobile[Mobile client: short-lived token only]
Recon[(Periodic reconciliation pull)]
IdP --> Bus
Risk --> Bus
AuthZ --> Bus
Bus -->|push| Gateway
Bus -->|push| Micro
Gateway -.->|poll on reconnect| Recon
Micro -.->|poll on reconnect| Recon
Recon -.-> Bus
Bus -.->|mobile: not relied on for push| Mobile
How change events are detected. React in near-real-time to a small, deliberately limited set of event types, not every attribute change: password or multi-factor authentication (MFA) change or removal, account disable, an explicit admin-initiated revocation, a risk engine's fraud or compromise signal (impossible travel, a known-bad network address), and a role or privilege change above a defined severity threshold. Each is detected at its natural source, the identity provider (IdP) for authentication events, the risk engine for fraud signals, the authorization system for privilege changes, and converted into one normalized critical-change-event format so downstream propagation never needs to understand each source's own internal details.
How events are securely propagated to enforcement points. A lightweight publish/subscribe event bus, or a standards-based mechanism such as the OpenID Foundation's Shared Signals Framework and its Security Event Token format, carries events to every subscribed enforcement point, whether that is an API gateway, a microservice, or (with the caveat below) a mobile client. Events are keyed by subject or session identifier, so an enforcement point can cheaply check whether an incoming event concerns any session it currently trusts, without processing irrelevant events meant for other subjects.
Push versus pull invalidation. Push delivery, the event bus actively sending the event to every subscriber as soon as it exists, has the lowest latency, since propagation time is roughly the transport and processing time alone, but it requires every enforcement point to maintain a live, healthy subscription; one that briefly disconnects misses events during that window with no way to know it missed anything, unless it separately reconciles afterward. Pull, where each enforcement point periodically asks a central service for new events concerning its current sessions, is simpler and self-healing, since a temporarily offline point simply catches up on its next poll, but it introduces a latency floor equal to the poll interval, and shortening that interval to reduce the floor increases aggregate polling load on the central service regardless of whether anything actually changed. The practical answer at this scale is a hybrid: push as the primary, low-latency path, backed by periodic pull-based reconciliation as a safety net, so a missed push becomes a bounded delay until the next reconciliation, not a silent, unbounded gap.
Securing the propagation channel. The event itself needs the same rigor as any other security-critical assertion in the system, since a forged "ignore that event" message, or the reverse, a forged flood of fake revocations used as a denial-of-service against legitimate users, would defeat or abuse the whole mechanism. Events should be signed by their originating source so an enforcement point can verify authenticity rather than trusting whichever system happened to deliver the message, transported over an encrypted, mutually authenticated channel (mutual TLS fits naturally here, using the same machine-identity mechanism already governing service-to-service authentication elsewhere in the system), and carry a monotonic sequence number or timestamp per subject so an enforcement point can detect and correctly discard a replayed or out-of-order event rather than processing a stale one after a fresher one has already arrived.
Minimizing latency while scaling to tens of thousands of sessions. Route events so an enforcement point only receives the ones relevant to sessions it is actually serving, for example by partitioning the event bus by subject, rather than broadcasting every event to every node globally and filtering locally; this keeps per-node processing cost roughly constant as both session count and node count grow. At each enforcement point, apply received events to a small, local, in-memory revocation cache that every request checks first, a cheap lookup that stays fast even as the cache grows into the thousands of entries. Bound that cache's growth by pruning any revoked entry once the token it would have invalidated has expired on its own anyway, since the entry is then redundant; this keeps the steady-state cache size proportional to event rate times token lifetime, not to the total historical count of every event ever generated, and definitely not to the total number of active sessions in the system. Mobile clients need a different treatment: a device that is offline will not receive a push no matter how well the propagation system is built, so mobile clients should rely on a short, frequently-refreshed token lifetime as their real exposure bound, rather than assuming push delivery will reach them promptly.
Worked example
Take 40,000 active sessions served across a fleet of enforcement-point instances, and assume, illustratively (not a given figure), that the system observes an average of 50 critical change events per minute across the whole population, combining password resets, MFA changes, admin revocations, and risk-engine signals.
With an access-token lifetime of 15 minutes, the steady-state size of each enforcement point's local revocation cache is bounded by event rate times token lifetime: 50 events per minute x 15 minutes = 750 entries needing to be retained at any moment, since any entry older than 15 minutes corresponds to a token that has already expired on its own and can be pruned. That figure stays at roughly 750 regardless of whether the system has 40,000 active sessions or ten times that, because the cache tracks only currently-revoked entries, not the full active-session population, which is exactly the property that lets this design scale.
For worst-case latency: push delivery's actual transit time depends on the transport and is not something to assert as a fixed number here, but the reconciliation interval is a chosen design parameter, not a matter of luck. If reconciliation runs every 30 seconds, then even a session behind a momentarily disconnected gateway instance is guaranteed to be caught within 30 seconds of reconnecting, giving the system a defensible, stated worst-case revocation latency, derived directly from a parameter the design controls, rather than an unbounded or merely hoped-for figure.
Trade-offs and pitfalls
- Pure push with no reconciliation leaves an open-ended gap. An enforcement point disconnected during a rolling deployment or a network blip misses events silently and has no way to know anything was missed unless it separately re-syncs; the hybrid design exists specifically to bound this otherwise unbounded exposure.
- Pure pull at a short interval to chase lower latency scales the central service's load linearly with the number of enforcement points times the poll frequency, whether or not anything changed. Taken far enough, the central service becomes the very bottleneck and single point of failure the design was meant to avoid.
- Broadcasting every event to every enforcement point globally does not scale as both session count and node count grow. Partitioning by subject is what keeps per-node cost roughly constant instead of growing with total system size.
- Assuming a mobile client will receive a push promptly is a common and consequential mistake. A disconnected device simply will not, regardless of how well the propagation system is engineered, so mobile's real security posture rests on its token lifetime, not on push latency, and designing as though the reverse were true leaves a much larger real exposure window than intended.
- Treating the event channel as inherently trustworthy because it is internal plumbing ignores that it is itself now a security-critical surface. An unauthenticated channel can be abused to suppress a real revocation or to flood legitimate users with fake ones, so it needs the same signing and authentication discipline as any other credential-bearing path in the system.
List TLS/HTTPS best practices to protect authentication credentials and tokens in transit for API services. Include minimum protocol versions, cipher suite recommendations, HSTS, certificate lifecycle management, and when mutual TLS is appropriate for stronger client authentication.
Sample Answer
Direct answer
Protecting authentication credentials and tokens in transit means requiring TLS (Transport Layer Security, the protocol that encrypts and authenticates a connection) 1.2 as an absolute minimum, with TLS 1.3 preferred for new deployments; enforcing HTTPS with HSTS (HTTP Strict Transport Security, a response header that tells the browser to never fall back to plain HTTP for this site); using only modern, forward-secret cipher suites; and treating certificate issuance, renewal, and revocation as an ongoing operational process, not a one-time setup step. Mutual TLS (mTLS, where both the client and the server present and verify certificates, not just the server) adds a stronger, certificate-based client authentication layer on top, appropriate for service-to-service traffic and other cases where a password or token alone is not sufficient assurance of the caller's identity.
Structured elaboration
Minimum protocol version. TLS 1.0 and TLS 1.1 are formally deprecated (IETF RFC 8996, March 2021) and should not be accepted at all; TLS 1.2 is the practical minimum for any API handling credentials or tokens today, and TLS 1.3 (which removed several legacy, weaker mechanisms outright, including static RSA key exchange and CBC-mode ciphers) should be preferred wherever the client and server stack both support it.
Cipher suite recommendations. Prefer AEAD (authenticated encryption with associated data) ciphers such as AES-GCM or ChaCha20-Poly1305, negotiated with an ephemeral (forward-secret) key exchange like ECDHE, so that even if a server's long-term private key is later compromised, past captured traffic cannot retroactively be decrypted. Avoid legacy ciphers (RC4, 3DES) and static (non-ephemeral) RSA key exchange, which lacks forward secrecy entirely; TLS 1.3 simplifies this by only offering a small, modern, forward-secret cipher set by design, removing the older options as a configuration mistake to make in the first place.
HSTS. Sending the Strict-Transport-Security response header (with a meaningful max-age, and includeSubDomains where applicable) instructs the browser to refuse any future plain-HTTP connection to the site for the specified duration, closing the window for an SSL-stripping attack (an on-path attacker silently downgrading a user's first request from HTTPS to HTTP before a redirect would normally correct it). A preload directive, submitted to the browser vendors' HSTS preload list, closes even that very first request's exposure, at the cost of being effectively permanent and hard to reverse once submitted.
Certificate lifecycle management. This is no longer a "set it up once a year" task: the CA/Browser Forum's maximum publicly-trusted TLS certificate validity period is now on a mandated shrinking schedule (398 days previously, reduced to 200 days as of March 2026, dropping to 100 days in March 2027 and 47 days by March 2029), which makes manual certificate renewal impractical going forward and effectively requires automation, most commonly via the ACME protocol (the mechanism Let's Encrypt and most modern certificate authorities support) to issue, renew, and deploy certificates without manual intervention, alongside expiry monitoring and a working revocation path (OCSP or a certificate revocation list) for the rare case a private key is compromised before natural expiry.
When mutual TLS is appropriate. Standard (server-only) TLS authenticates the server to the client but not the reverse; the server still relies on a separate mechanism (a password, an API key, a bearer token) to authenticate the client. Mutual TLS is appropriate when that separate mechanism is not strong enough assurance on its own, most commonly service-to-service traffic inside a backend (where each service holds its own certificate and the calling service's identity should be cryptographically verified, not just asserted by a token that could be copied), or a small set of high-assurance external integrations (a banking or payments partner) where both parties need strong mutual identity assurance beyond what a bearer credential provides. It is generally not the right fit for ordinary browser-to-server consumer traffic, where issuing and managing a client certificate per end user is operationally heavy for the assurance gained over a well-implemented token-based scheme.
Worked example
An API team is deciding how to protect a service that issues and validates OAuth2 access tokens. Baseline: enforce TLS 1.2 minimum (reject any TLS 1.0/1.1 handshake attempt outright), prefer TLS 1.3, restrict cipher suites to AEAD-with-ECDHE only, and send Strict-Transport-Security: max-age=63072000; includeSubDomains; preload (a two-year max-age, common for a domain being submitted to the preload list). Certificates are issued via ACME with automated renewal well before each certificate's (now shorter) validity window ends, with alerting if a renewal fails. For the token-issuing endpoint's own service-to-service callers (internal microservices that need to mint tokens on behalf of users), the team adds mutual TLS on top of TLS itself: each internal caller presents its own service certificate, verified against an internal certificate authority, so a stolen API key alone is not sufficient to call the token-issuance endpoint as an impersonated internal service. Public-facing browser traffic to the same API, by contrast, stays on standard server-only TLS with strong bearer-token authentication, since issuing browser users individual client certificates would be significant operational overhead for comparatively little additional assurance in that context.
Trade-offs and pitfalls
HSTS preload is effectively a one-way door: removing a domain from the browser preload lists takes a long time to fully propagate to all users, so a mistake made while HSTS-only infrastructure is still being finished can lock out legitimate traffic for an extended period; stage it (start with a short max-age, verify HTTPS is solid everywhere including all subdomains, then increase the duration and add preload) rather than jumping straight to the strongest setting. The most common certificate-lifecycle pitfall, especially as the industry-mandated maximum validity period keeps shrinking, is treating renewal as a manual, calendar-reminder process; that approach was already fragile at 398-day certificates and becomes untenable at the shorter validity periods now being phased in, making automated issuance and renewal the only practical long-term approach rather than an optional convenience. Mutual TLS's main pitfall is operational: it requires a working internal certificate authority, distribution mechanism, and rotation process for every participating service, and adopting it without that operational maturity in place creates more outage risk (an expired or misconfigured client certificate breaking service-to-service calls) than the security benefit is worth for a team not ready to run it.
Explain Cross-Site Request Forgery (CSRF): how it works, typical attack chains, and at least four mitigation strategies for web applications. Include differences in mitigation when auth state is stored in cookies versus when Authorization headers are used.
Sample Answer
Direct answer
Cross-Site Request Forgery (CSRF) exploits the fact that browsers automatically attach a user's stored credentials, most commonly session cookies, to every request sent to a site, regardless of which page actually initiated that request; an attacker who gets a logged-in victim to load a crafted page can trigger a state-changing request to the real site, and the browser helpfully attaches the victim's valid session cookie along with it. At least four independent mitigations exist: CSRF tokens (the synchronizer-token pattern), the SameSite cookie attribute, verifying the Origin or Referer header on state-changing requests, and requiring a custom request header that a simple cross-site form cannot set. The attack only works at all because the browser auto-attaches cookies; when authentication instead lives in an Authorization header that the application code must explicitly set, a cross-site page has no built-in mechanism to attach it, which removes most of the classic CSRF threat model outright, at the cost of shifting the risk toward how that token is stored and whether it can be stolen via other means (most notably Cross-Site Scripting).
Structured elaboration
How CSRF works, mechanically. A user logs into bank.example.com and receives a session cookie. Without logging out, they visit attacker.example, which contains a hidden HTML form (or an auto-submitting fetch/image tag for GET-based variants) that targets bank.example.com/transfer with attacker-chosen parameters. When the victim's browser loads that page, it submits the form to bank.example.com, and because cookies are attached per-domain regardless of which page triggered the request, the victim's real session cookie rides along. If the server's only check is "is this session cookie valid," the request succeeds as if the victim had submitted it themselves, transferring funds, changing an email address, or performing whatever the endpoint does, with the attacker never touching the victim's credentials directly.
Typical attack chains. The classic chain is: victim is authenticated to a target site in one browser tab -> victim is lured (phishing link, malicious ad, compromised third-party page) to visit an attacker-controlled page in another tab or the same session -> that page auto-submits a form or fires a cross-site request to a sensitive endpoint on the target site -> the browser attaches the victim's cookie automatically -> the server performs the action believing it was a legitimate request from the authenticated user. A GET-based variant is even simpler and needs no form at all: an <img src="https://bank.example.com/transfer?to=attacker&amount=500"> tag fires a GET request the instant the page loads, which is exactly why state-changing operations should never be reachable via GET.
Mitigation 1: CSRF tokens (synchronizer-token pattern). The server generates a random, session-bound token, embeds it in the legitimate page's forms (as a hidden field) or exposes it for the client to attach as a header, and rejects any state-changing request that does not include a matching token. An attacker's cross-site page cannot read this token (it is not set as a readable cookie the attacker's origin can access, and same-origin policy blocks reading the real page's content from a different origin), so it cannot forge a request that includes the correct value.
Mitigation 2: SameSite cookie attribute. Setting the session cookie to SameSite=Lax or SameSite=Strict tells the browser itself not to attach the cookie on qualifying cross-site requests (all non-GET, non-top-level requests for Lax; effectively all cross-site requests for Strict), which blocks the classic hidden-form attack at the browser level before it ever reaches the server. This is a strong, low-effort layer, but should not be relied on exclusively, since it depends on the browser enforcing it correctly and does not cover every legitimate cross-site flow a real app might need.
Mitigation 3: Verifying Origin/Referer on state-changing requests. Modern browsers reliably send an Origin header on cross-origin requests (and often same-origin ones too for non-GET requests); a server can reject any state-changing request whose Origin does not match its own, which catches forged cross-site requests even from a client that ignores SameSite. Referer-based checking is an older, less reliable fallback (it can be stripped by privacy settings or proxies), so Origin is the preferred signal where available.
Mitigation 4: Requiring a custom header. A plain HTML form submitted cross-site can only set a limited set of headers and content types; requiring, say, an X-Requested-With header or a custom X-CSRF-Token header on state-changing requests means a bare cross-site form POST cannot satisfy the requirement at all, since forms have no mechanism to add arbitrary headers. This is a genuine partial defense on its own and a natural pairing with mitigation 1, since the CSRF token is commonly delivered exactly this way.
Cookie-based vs Authorization-header-based auth. CSRF as a threat model exists specifically because the browser attaches cookies automatically and without the page's cooperation. When authentication instead uses a bearer token sent in an Authorization header, the browser has no built-in mechanism to attach that header to a request initiated by a different origin's page; JavaScript on the attacker's page cannot set an Authorization header on a simple cross-site form submission at all, and a cross-site fetch call that tried to add one would need the token value in the first place, which the attacker's origin does not have access to under the same-origin policy. So switching authentication from cookies to an explicitly-attached Authorization header removes the classic CSRF attack surface almost entirely; the trade-off is that the application code must now securely obtain and store that token itself (commonly in memory rather than local storage), which reopens the door to XSS-based token theft if that storage decision is made carelessly. CSRF and XSS-based token theft are different threats with different root causes, cookie auto-attachment for one, script execution in a trusted origin for the other, and moving to header-based auth trades one for needing to manage the other more carefully, not a free win.
Worked example
Trace the classic transfer-forgery scenario against a server with no CSRF defenses at all:
- Victim logs into
bank.example.com, receivesSet-Cookie: session=abc123with noSameSiterestriction. - Victim, still logged in, visits
attacker.example, which contains:<form action="https://bank.example.com/transfer" method="POST" id="f"><input name="to" value="attacker_account"><input name="amount" value="5000"></form><script>document.getElementById('f').submit()</script>. - The browser submits this form to
bank.example.com. Per-domain cookie storage meanssession=abc123is attached automatically, since the browser does not consider which page triggered the request, only which domain it targets. bank.example.comsees a request with a valid session cookie and processes the transfer.
Now apply all four mitigations and re-trace: with SameSite=Lax on the session cookie, step 3 never attaches the cookie at all for this cross-site POST, and the request arrives unauthenticated. Even if the cookie somehow attached anyway, the server's CSRF-token check (mitigation 1) rejects the request because the attacker's form has no way to include the per-session token that was only ever delivered to the real page. The Origin header check (mitigation 3) independently catches it, since the request's Origin is https://attacker.example, not https://bank.example.com. And the requirement for a custom header (mitigation 4) means the plain HTML form in step 2 could never have satisfied the request shape at all, since forms cannot set arbitrary headers, forcing the attacker toward more complex (and more blockable) techniques.
Trade-offs and pitfalls
- A common mistake is implementing exactly one of these four mitigations and considering CSRF solved; each has a specific gap (
SameSitedepends on browser enforcement,Originchecking can occasionally be complicated by legitimate proxies, CSRF tokens require correct per-session issuance and validation, custom-header requirements can be bypassed on endpoints that also accepttext/plainbodies without careful configuration), so production systems typically layer at least two of these rather than relying on one. - Allowing state-changing operations via GET is a design mistake independent of CSRF defenses, since the simplest CSRF payload (an image tag) is a GET request and bypasses form-based reasoning entirely.
- Moving to
Authorization-header auth removes CSRF but is not a strict security upgrade on its own; it shifts the burden to secure token storage and handling, and a careless implementation (token in local storage, readable by any injected script) can end up more exposed than a well-configured cookie withHttpOnlyandSameSite, not less. - CSRF tokens must be tied to the session and validated server-side on every state-changing request, not merely present; a common shortcut of checking only that "a token field exists" without verifying it matches the session's issued value defeats the purpose entirely.
Compare OAuth 2.0, OpenID Connect (OIDC), and SAML for solving authentication and authorization problems. For each protocol explain primary use cases (e.g., web SSO, mobile apps, enterprise federation), how authentication statements are conveyed, and typical deployment considerations (mobile vs enterprise SSO). Provide criteria you would use to choose one protocol over the others.
Sample Answer
Direct answer
OAuth 2.0 is an authorization framework: it lets a user grant a third-party application scoped access to an API or resource without handing over a password. On its own it has no standard concept of "who logged in," only "what access was granted." OpenID Connect (OIDC) is a thin identity layer built on top of OAuth 2.0 that adds a standardized token proving who authenticated, not just what the app can now touch. SAML (Security Assertion Markup Language) is an older, XML-based protocol built specifically for browser-based single sign-on (SSO), most often used to federate identity into enterprise web applications. All three answer "who is this, and can we trust that claim across a network boundary," but they target different client types and eras of the web.
Structured elaboration
| OAuth 2.0 | OIDC | SAML | |
|---|---|---|---|
| Primary use case | Delegated authorization: "let this app read my calendar" | Web and mobile login (SSO): "let this app know who I am" | Enterprise SSO: federating identity into an organization's web apps |
| How the trust statement is conveyed | An access token authorizes API calls; the token itself does not certify who authenticated | A signed ID token (a JSON Web Token, or JWT) carrying claims such as sub, iss, aud, exp, and the time of authentication | An XML assertion containing an authentication statement, signed by the identity provider and posted to the application via a browser redirect |
| Typical deployment fit | Any client needing scoped API access: mobile apps, single-page apps, machine-to-machine calls | The modern default for new consumer and enterprise login integrations | Legacy and regulated enterprise software, where the identity provider is often an on-prem or cloud directory (for example Active Directory Federation Services, or an identity provider like Okta configured for SAML) that the buyer already standardized on |
Criteria for choosing between them:
- Building a login experience for a modern web or mobile app that also needs to call an API on the user's behalf: use OIDC. It sits on top of OAuth 2.0, so you get delegated authorization and a verified identity from the same flow.
- Only need delegated API access, with no identity concept for the calling app itself (a backend job reading a user's calendar): plain OAuth 2.0 is sufficient and simpler.
- Integrating with an enterprise's existing identity provider, and that provider or the target application only speaks SAML: you use SAML even though it is heavier to implement than a JWT-based approach, because the counterpart has no OIDC endpoint to talk to.
- Mobile deployment pushes the decision toward OIDC: SAML's browser-redirect-and-XML-post pattern is awkward inside a native app, while OIDC's authorization code flow was purpose-built for exactly that client type.
- Enterprise SSO deployment sometimes forces SAML regardless of preference: large enterprise buyers frequently standardize on a SAML identity provider for audit and compliance reasons, and the vendor's application may expose only a SAML integration point.
Worked example
Trace "Log in with Google," which demonstrates both the OAuth delegation layer and the OIDC identity layer built on top of it:
- The user clicks "Log in with Google"; the app redirects to Google's authorization endpoint requesting
scope=openid email profile. - Google authenticates the user through its own login screen and the user consents to the requested scopes.
- Google redirects back to the app with a short-lived authorization code.
- The app exchanges that code at Google's token endpoint for an ID token (the OIDC-specific artifact, a JWT) and an access token (the underlying OAuth artifact).
- The app validates the ID token's signature and claims (
issequals Google's issuer,audequals the app's own client id,exphas not passed) to learn who logged in, from thesubandemailclaims. - If the app also wants to read the user's Google Calendar, it uses the separate access token for that call. That is the original OAuth layer doing its job, distinct from step 5.
This split is exactly why using bare OAuth to implement login was a historical mistake, before OIDC existed: an access token alone does not certify identity (it authorizes calls to a specific API, and it isn't required to be a verifiable, self-contained token at all), so an app inspecting only an access token to decide "who is logged in" could be fooled by a token that was legitimately issued, just for a different purpose or audience. OIDC's ID token exists specifically to close that gap.
Trade-offs and pitfalls
- SAML assertions are XML-based and require careful canonicalization and signature validation. Implementing that by hand is a well-known source of signature-wrapping vulnerabilities; always use a maintained library rather than parsing and verifying the XML yourself.
- Treating an OAuth access token as proof of identity, instead of using OIDC's ID token, is the single most common protocol-selection mistake in this space; it works in testing and fails once a token issued for a different audience gets presented to the wrong service.
- Bridging protocols (a SAML-only enterprise identity provider fronting an OIDC-only application, or the reverse) is a common real integration need, but it adds an extra hop and an extra trust boundary; treat that bridge as its own design problem rather than assuming one protocol trivially substitutes for the other.
Unlock Full Question Bank
Get access to all Identity, Authentication, and Access Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.