API Security, Authentication and Authorization Questions
Controlling who can call an API, what they may do, and defending it against abuse. Covers the access-control mechanics: API keys, OAuth 2.0 flows, OpenID Connect, JWT issuance/validation, session vs. token auth, scopes/roles for fine-grained authorization, token lifetime and refresh, mutual TLS, and machine-to-machine vs. user-delegated access. Also covers the adversarial hardening view: input validation, injection and deserialization risks, broken object-level authorization (BOLA), mass assignment, secrets handling, and the OWASP API Security Top 10, plus securing data in transit, preventing enumeration/scraping, and testing APIs for vulnerabilities.
Compare opaque tokens versus self-contained JWTs for service-to-service authentication in a distributed microservices environment. Discuss performance (validation latency), revocation complexity, payload confidentiality, token size, caching implications, and give recommendations for hybrid approaches that balance stateless validation with revocation needs.
Sample Answer
Direct answer
Opaque tokens trade a network round trip (an introspection call to the issuer) for instant, centralized revocation. Self-contained JSON Web Tokens (JWTs) trade that centralized control for fast, local, offline validation. In a distributed machine-to-machine (M2M) environment at real scale, most teams land on a hybrid: short-lived signed JWTs as the default, since the overwhelming majority of calls should need zero network hop to validate, backed by a narrow, fast-path mechanism for the rare "kill this credential right now" case.
Structured elaboration
| Dimension | Opaque token | Self-contained JWT |
|---|---|---|
| Validation latency | Every validating service calls the issuer's introspection endpoint, adding a network hop and making the issuer a shared, latency-sensitive dependency at scale | Local signature check only, using a cached public key, no network call, scales horizontally with the caller itself |
| Revocation | Trivial: the issuer marks the token invalid, the very next introspection call reflects it | Hard: the token is valid until its exp claim regardless of what the issuer now wants, unless you add an auxiliary mechanism |
| Payload confidentiality | The token itself carries no information; a captured token reveals nothing without querying the issuer | Claims are base64url-encoded, not encrypted, by default, readable by anyone holding the token, including any intermediary that logs or proxies it |
| Token size | Small and fixed, often just an opaque identifier | Grows with claim count, adding real header overhead on every call at high queries-per-second (QPS) if claims bloat |
| Caching implications | Services cache introspection results, reintroducing a staleness-versus-load trade-off on top of the revocation problem itself | Services cache the issuer's signing keys (via a JWKS, JSON Web Key Set, endpoint), which change rarely, so this cache is cheap and carries no correctness trade-off |
Why revocation is the real fork in the road: because a JWT is self-contained, "revoking" it does not mean anything to the token itself, the signature still verifies and the claims are still readable. Real revocation for JWTs requires bolting on state: short time-to-live (TTL) values as the default mitigation (bounding the exposure window), a deny-list checked at validation time (which reintroduces a lookup, partially undoing the statelessness benefit you adopted JWTs for), or issuer-side generation/versioning per client.
Worked example
Consider a service mesh where an order service calls an inventory service, which calls a pricing service, on every checkout request. With opaque tokens, each of those two internal hops adds a full round trip to the issuer's introspection endpoint before the call is even allowed to proceed, meaning a single external request now pays for multiple serial introspection calls stacked on top of its own logic. With signed JWTs, each hop verifies the caller's token locally against a cached public key, no extra network dependency introduced by authentication itself. This qualitative difference, not a specific millisecond figure, is why introspection-per-call does not scale cleanly as call chains get deeper.
Recommendations for a hybrid approach
- Default to short-lived signed JWTs (minutes, not hours, since machine-to-machine calls do not need a human to reauthenticate) so the common case stays local and fast.
- Pair that with a lightweight, narrow deny-list checked only for the rare case that actually needs it (a compromised service account, a discovered key leak), not on every single request; expire each deny-list entry once the underlying JWT would have expired anyway, so the extra state stays bounded rather than growing forever.
- Support issuer-side token-family versioning (bumping a minimum-issued-at cursor per client) so you can cheaply invalidate every future token for a specific caller, checked at the same low frequency as the JWKS cache refresh, without a per-request lookup.
- Reserve opaque tokens plus introspection for genuinely high-stakes, low-volume paths, such as issuing a brand-new elevated-privilege credential, where the extra round trip is an acceptable cost for tighter control.
Trade-offs & pitfalls
Choosing JWTs everywhere and ignoring the revocation gap is a common answer that sounds clean but quietly assumes a compromised credential is an acceptable risk for its full remaining lifetime. Choosing opaque tokens everywhere under-appreciates that introspection becomes the single most-called, most latency-sensitive endpoint in the whole system as call volume grows. A subtler pitfall: forgetting that JWT claims are readable, not encrypted, and putting anything genuinely secret into a claim, where it is visible to every intermediary that ever sees the token.
Perform a threat model for an external-facing API. Identify threats such as injection attacks, broken authentication, excessive data exposure, rate-limiting bypass, and DDoS. As a Solutions Architect, propose mitigation strategies including validation, least-privilege, rate limits, WAF, and API-level quotas, and discuss trade-offs and monitoring approaches.
Sample Answer
Structure the threat model with a systematic method, the STRIDE categories (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege), applied to each trust boundary the API crosses: client to gateway, gateway to backend, backend to database. Then map every identified threat to a concrete mitigation and a way to monitor for it in production, since a threat model that stops at a list without controls and monitoring doesn't change what ships.
Scope and trust boundaries
External-facing REST API, untrusted public internet clients, an API gateway (trust boundary 1), backend services (trust boundary 2), and a database or storage layer (trust boundary 3).
Systematic identification, by category and boundary
| STRIDE category | Example threat at this API | Primary mitigation |
|---|---|---|
| Spoofing | attacker forges a client identity or replays a stolen token | strong authentication (OAuth 2.0 or mutual TLS), short token lifetimes |
| Tampering | request or response payload modified in transit or by a compromised intermediary | TLS everywhere, request signing for high-value operations |
| Repudiation | client denies having made a state-changing call | structured, tamper-evident audit logs tied to an authenticated identity |
| Information disclosure | the API returns more fields than the caller should see | schema-driven response filtering, per-scope field allowlists |
| Denial of service | one client or an expensive query exhausts shared capacity | per-client rate limits and quotas at the gateway, query cost limits |
| Elevation of privilege | a regular-user token reaches an admin-only endpoint | function-level authorization enforced server-side on every route, never inferred from the client's UI |
Mapped onto the threats the question specifically calls out: injection attacks (SQL, NoSQL, command) are a tampering-category failure, since they let an attacker change what a backend query or command actually does; broken authentication is a spoofing failure, letting an attacker pass as someone they're not; excessive data exposure is this API's information-disclosure failure, returning fields the caller shouldn't see; rate-limiting bypass and distributed denial-of-service (DDoS) are both denial-of-service failures, differing only in mechanism, one exploits a gap in how limits are enforced, the other simply throws enough volume at the API to matter regardless of any single limiter.
Mitigations, organized by where they run
Design time: enumerate every object identifier in a URL or body and write down, per identifier, how ownership is verified; use parameterized queries or an ORM to close injection paths. Gateway or edge: OAuth 2.0 or API key authentication, a web application firewall (WAF) with managed and custom rules, per-client rate limiting, TLS termination with modern cipher suites. Application layer: least-privilege service accounts, explicit authorization checks per action (not just per resource), input validation against a strict schema. Data layer: encryption at rest, database accounts scoped to only the tables and operations that service actually needs.
Worked example
Consider GET /v1/accounts/{accountId}/statements. The threat-model pass asks: can caller A read caller B's statements by changing accountId (information disclosure)? Can a client request a huge date range and force a full table scan (denial of service)? Does the handler check a verified role claim, or does it trust a client-suppliable header (elevation of privilege)?
The mitigation trace: add a server-side check statement.account.owner_id == token.sub before returning any data, closing the disclosure gap; cap the date range parameter server-side (for example, reject ranges over 366 days and force pagination), closing the expensive-query denial-of-service path; derive the caller's role only from the verified token's claims, never from a client-suppliable header, closing the privilege-escalation path. The monitoring trace: alert if a single token generates 403s on accountId values it doesn't own more than five times in a minute (an enumeration signature); alert on requests repeatedly hitting the date-range cap from one client (probing for the limit).
Trade-offs and pitfalls
A full STRIDE pass on every endpoint before every release doesn't scale, reserve the deep pass for new trust boundaries or high-value endpoints, and rely on automated contract tests to catch regressions on endpoints already modeled. WAF rules cut noise but generate false positives that erode trust in alerts if left untuned, budget time to tune them, not just deploy them. These same categories map closely onto the current OWASP (the Open Web Application Security Project, a nonprofit that publishes ranked lists of common vulnerability categories) API Security Top 10 (the 2023 edition): the information-disclosure and elevation-of-privilege threats above correspond to API3:2023 Broken Object Property Level Authorization and API5:2023 Broken Function Level Authorization respectively. Using STRIDE to find threats systematically and the OWASP list to sanity-check nothing common was missed is stronger than relying on either alone, but the two lists are not interchangeable: they are two separate documents with their own independent numbering, and the API-specific OWASP list (API1:2023-API10:2023) is a separate, narrower edition than the general OWASP Top 10 for web applications broadly (currently the 2025 edition), don't cite a category number from one list as if it belongs to the other.
You need an access control model for an API that supports fine-grained permissions (resource-level, action-level) and can scale to millions of principals and resources. Discuss evaluation latency, caching of permissions, hierarchical roles, attribute-based access control, and how to keep revocation latency low.
Sample Answer
Direct answer
Separate the authorization decision from the authorization data: build a dedicated policy-evaluation layer that answers "can this principal take this action on this resource" against a cached, denormalized view of the permission graph, rather than computing that answer with a live join across primary relational tables on every request. At millions of principals and resources, the live-join approach is both too slow and too tightly coupled to core application schema.
Structured elaboration
Modeling resource-level and action-level permissions
Model permissions as tuples of principal, action, and resource, or principal, action, and resource pattern, rather than a single flat role per user. Real APIs need both "can Alice read document 42," resource-level and specific, and "can any Editor create documents in project X," action-level and pattern-based. A model that can only express one of these two shapes forces awkward workarounds for the other.
Hierarchical roles
Define roles that inherit from other roles, Viewer as a subset of Editor as a subset of Admin, so permission grants are authored once per role rather than once per user-permission pair. Combine this with resource hierarchies, a permission granted at a folder or project level implicitly applying to everything nested underneath, so you are not writing millions of individual grants for millions of individual resources.
Attribute-based access control (ABAC)
Layer attribute-based rules on top of the role hierarchy for decisions that depend on context a static role cannot express: time of day, a resource's sensitivity tag, whether a principal's department matches the resource's owning department, or the request's originating IP range. The practical pattern at scale is "roles for the coarse default, ABAC for the exceptions," not replacing roles entirely; a pure-ABAC system where every decision is a fresh rule evaluation becomes far harder to reason about and audit as the rule set grows.
Evaluation latency
The authorization check sits directly on the request hot path, so it has to be fast. Achieve that by pre-computing or denormalizing the parts of the decision that do not change often, role-to-permission mappings and resource hierarchy, into a form the evaluation layer can check in memory or via a single cache lookup, reserving genuinely dynamic per-request evaluation for the smaller set of cases that actually need ABAC attribute matching.
Caching of permissions
Cache the evaluated decision, or a principal's effective permission set, close to the service making the check, with a short time-to-live (TTL). Cache the underlying role and hierarchy graph, which changes far less often than individual grants, more aggressively. These two caches need different invalidation strategies, since they change at very different rates.
Keeping revocation latency low
This is the direct tension with caching: a cached "yes" that should now be "no" is a live authorization bug, not merely staleness, so revocation needs an active invalidation path, pushing a targeted cache-bust for the specific principal, resource, or role that changed, rather than relying on TTL expiry alone to eventually catch up. A common pattern combines a short default TTL, seconds, not minutes, with an active invalidation event fired on any grant or role change, so the cache stays fresh in the common case and the invalidation event closes the remaining gap immediately rather than waiting out the TTL.
Worked example
A document-sharing API has a folder hierarchy. Granting "Editor" on a top-level folder to a principal applies to every document created under it going forward, with no new grant row written per document. The evaluation layer first checks the principal's cached effective-permission set, the fast path, doing no real work at all on a cache hit. On a miss, it walks the resource's hierarchy chain upward to the nearest ancestor with an explicit grant, then caches that result with a short TTL. Revoking the folder-level grant fires an invalidation event that busts the cached decision for every principal-resource pair touched by that grant, rather than waiting for each individual cache entry's TTL to expire on its own.
Trade-offs & pitfalls
Hierarchical inheritance is powerful but makes "why does this principal have this permission" hard to answer without good tooling; a permission-explain or trace capability is not optional at this scale, it is how the system actually gets debugged and audited. ABAC's flexibility is also its risk: rules that reference many attributes become hard to reason about and can interact in unintended ways, so keep the ABAC rule set small and reviewed rather than letting it grow into an ever-larger, ever-harder-to-audit pile. Treating revocation latency as "eventually consistent is fine" is the single most common mistake at this scale; a stale cache that grants access after it should have been revoked is a security incident, not a minor user-experience issue, so the invalidation path deserves as much engineering attention as the fast-path cache itself.
Build an authentication and authorization scheme for a multi-tenant API that supports bearer tokens for user auth and API keys for service-to-service calls. Include per-tenant rate limits, key rotation, revocation, secure key storage, and how to represent tenant scoping in tokens or claims.
Sample Answer
Represent the tenant as a first-class claim (tid) inside every bearer token, so the API gateway can enforce tenant isolation without an extra database lookup on the hot path, and treat API keys for service-to-service calls as a separate credential type: hashed at rest, bound to exactly one tenant at issuance, rotated by versioning rather than in-place mutation.
Two credential types, one enforcement point
- User bearer tokens: OAuth 2.0 Authorization Code flow (the user logs in and consents at the identity provider, which redirects back with a one-time code that the app then exchanges server-side for a token) issues a short-lived JSON Web Token (JWT), 5-15 minutes, with claims
sub(user id),tid(tenant id), androles/scopes. - Service API keys: a high-entropy key id plus secret, sent as
Authorization: ApiKey <id>.<secret>, stored hashed server-side (for example with Argon2id), with a metadata row mappingkey_id -> tenant_id, scopes, status. - Both paths terminate at one API gateway, which extracts a normalized identity and tenant context (
X-Tenant,X-Scopesheaders) and forwards it downstream, so backend services never re-implement auth.
Tenant scoping in tokens
Whether the credential is a JWT's tid claim or an API key's tenant_id metadata row, both resolve to the same enforcement check: every downstream query filters WHERE tenant_id = :tid, and the gateway rejects any request whose payload references a different tenant_id than the token's own, before it ever reaches business logic. This is the tenant-boundary version of the same defense used against broken object-level authorization (a caller reaching another user's specific record just by guessing or changing its id) at the user boundary.
Per-tenant rate limits
Use a token bucket keyed by tenant_id, not by individual API key, because a tenant with several service accounts should share one quota. A per-key-only limit lets a tenant multiply its effective throughput just by minting more keys. Store buckets in a shared cache (for example Redis) keyed tenant:{tid}, with the refill rate set per pricing tier.
Key rotation and revocation
API keys are versioned (key_id:v1, v2, ...). Rotation creates a new version ACTIVE, marks the old one DEPRECATED for a grace window, then REVOKED. The revocation list is cached at the gateway and refreshed on a short interval or via an event push. Signing keys for JWTs rotate through a published JWKS (JSON Web Key Set) endpoint with overlapping key IDs, so tokens signed with the outgoing key still validate until they naturally expire.
Secure key storage
API key secrets are hashed (never stored or logged in plaintext) and shown in cleartext exactly once, at creation. Signing keys live in a key management service (KMS) or hardware security module (HSM); application code never touches the private key material directly.
Worked example
Tenant acme (tid=acme-042) runs a service account calling a billing API at 40 requests/second peak. Rate-limit config: 50 requests/second sustained, burst 100. The token bucket starts at 100 tokens and refills at 50/second, each request consuming one token. At a sustained 40 req/sec, the bucket never empties (refill 50 exceeds consumption 40), so no throttling occurs. If acme mints a second service account and the two together push 90 req/sec, the shared tenant-level bucket still throttles correctly at 50 req/sec sustained, whereas a per-key-only limit of 50 each would have let the tenant reach 100 req/sec simply by adding a key, defeating the purpose of a tenant-level quota.
Trade-offs and pitfalls
Putting tid only in application logic (not the token) forces a lookup on every request; putting it in the token is faster, but a tenant reassignment then requires reissuing tokens rather than just updating a database row. API keys are simpler for automation but riskier if leaked, since they carry no built-in expiry, mitigate with short rotation cadence and IP allowlisting per key. The common pitfall shown above is rate-limiting per API key instead of per tenant, which lets a tenant multiply its effective quota by minting more keys.
Architect a security model for a large-scale microservices platform (~1000 services) that uses a service mesh (e.g., Envoy/Istio) and an API gateway. Goals: enforce strong service-to-service authentication and authorization, minimize blast radius, centralize policy where sensible but avoid bottlenecks, ensure observability and incident response. Provide key components, identity model, policy enforcement points, rollout plan and scaling considerations.
Sample Answer
Direct answer
At roughly 1,000 services, service-to-service authentication and authorization has to be
enforced by infrastructure (a service mesh sidecar, such as Envoy, paired with a control plane
like Istio) rather than by each service's own application code, because you cannot reliably
audit or update security logic duplicated across a thousand codebases owned by many different
teams. The mesh's sidecar proxies handle mutual TLS (mTLS) and per-request authorization
uniformly, the control plane distributes identity and policy to every sidecar, and the API
gateway stays a separate, thinner layer that only handles edge concerns (external traffic
authentication, coarse rate limiting) rather than trying to be the single point enforcing every
internal rule.
Structured elaboration
Identity model. Every service instance gets a workload identity, most commonly a SPIFFE
(Secure Production Identity Framework For Everyone) identity encoded into a short-lived X.509
certificate, tied to what the workload is (its Kubernetes service account and namespace, for
example) rather than a static shared secret. The mesh's control plane (Istio's istiod, for
example) issues and continuously rotates these certificates automatically; no individual service
team manages its own certificate lifecycle.
Enforcing service-to-service auth and authorization. The sidecar proxy next to each service
terminates and originates mTLS transparently, so two services communicate over an encrypted,
mutually-authenticated channel without either one's application code implementing TLS itself.
Authorization (which services may call which other services, and for which operations) is
expressed as policy (Istio's AuthorizationPolicy resources, for example) and enforced by the
sidecar before a request ever reaches the application, giving every service the same
authorization enforcement mechanism regardless of what language or framework it is written in.
Policy enforcement points. At this scale there are exactly two places policy actually gets
enforced, and each has a distinct job: the API gateway is the enforcement point for traffic
entering the mesh from outside (external clients, partners), handling coarse checks like
authentication and rate limiting once at the edge; each service's own sidecar is the enforcement
point for everything after that, deciding, per call, whether one internal service may reach
another. Keeping these two enforcement points separate, rather than routing all internal traffic
back through the gateway, is what keeps the gateway from becoming a bottleneck for traffic that
never needed to leave the mesh.
Minimizing blast radius. Default-deny between services (a service can only call, or be
called by, the specific services its policy explicitly allows) turns a compromised service into
a contained incident instead of a pivot point to the rest of the fleet; without this, one
compromised service with network reachability to everything else effectively compromises the
whole mesh.
Centralizing policy without creating a bottleneck. Policy is authored and versioned
centrally (so the security posture is auditable in one place, as code) but distributed to
every sidecar so each one can make allow/deny decisions locally, at line rate, without a
synchronous call back to a central decision service on every request. This is the key design
move for enforcing at 1,000-service scale: centralize the authoring and distribution of
policy, not the evaluation of it.
Observability and incident response. Every sidecar can emit consistent access logs, metrics,
and distributed traces for the traffic passing through it, giving uniform visibility across all
1,000 services without depending on each service team to instrument request-level auth logging
themselves. This uniformity is what makes incident response at this scale tractable: a security
team traces a suspicious request across service boundaries using the mesh's own telemetry rather
than reconciling a thousand different logging formats.
Rollout plan. Introduce the mesh incrementally rather than flipping mTLS enforcement on
everywhere at once: start in permissive mode (the sidecar accepts both plaintext and mTLS
traffic while metrics show which callers have and have not migrated), onboard services namespace
by namespace or team by team, and only flip to strict mTLS enforcement for a given service once
its traffic is confirmed fully migrated. A big-bang cutover at 1,000 services risks a
simultaneous outage across the fleet if any meaningful fraction of callers have not yet adopted
the sidecar.
Scaling considerations. The sidecar's own resource footprint (CPU and memory per pod) and the
control plane's config-push fan-out both need capacity planning at this scale; a control-plane
change that pushes new policy to 1,000 sidecars simultaneously needs to be paced (canaried,
rate-limited) rather than broadcast all at once, since a bad policy pushed everywhere
simultaneously turns a policy bug into a fleet-wide outage instead of a contained one.
Worked example
graph TD
Ext[External Traffic] --> GW[API Gateway - Edge AuthN and AuthZ]
GW --> SM[Service Mesh Ingress]
SM --> SVA[Service A plus Envoy Sidecar]
SM --> SVB[Service B plus Envoy Sidecar]
SVA -->|mTLS plus SPIFFE ID| SVB
CP[Istio Control Plane] -.->|distributes policy and certs| SVA
CP -.-> SVB
SVA --> OBS[Telemetry: access logs, traces]
SVB --> OBS
The gateway handles only the edge boundary (authenticating and rate-limiting external traffic
before it enters the mesh at all); everything after that, service A calling service B, for
example, goes sidecar-to-sidecar over mTLS with an authorization decision made locally by
service B's own sidecar, based on policy the control plane already pushed to it. If the control
plane is briefly unavailable, existing sidecars keep enforcing their last-known policy and
certificates until they expire, since evaluation never depended on a live call to the control
plane; only new certificate issuance and policy updates pause during that window, which is
the direct payoff of "centralize distribution, not evaluation."
Trade-offs and pitfalls
- A service mesh adds real operational complexity and per-request latency overhead (an extra
network hop through the sidecar in each direction); this is a deliberate trade against the
alternative of every service reimplementing TLS and authorization independently, which does not
scale to 1,000 services being maintained correctly and consistently over time. The trade-off
is worth stating explicitly rather than presenting the mesh as free. - Permissive mode during rollout is a temporary state, not a target state. Leaving services
in permissive mode indefinitely (common when a migration stalls) means mTLS is not actually
being enforced for those services even though the infrastructure exists, which is easy to miss
in metrics that only measure "sidecar deployed" rather than "strict mode enforced." - Common wrong turn: routing all internal traffic back through the central API gateway to
reuse its authorization logic. This defeats the scaling argument for a mesh in the first place
and turns the gateway into both a latency bottleneck and a single point of failure for traffic
that never needed to leave the mesh's own service-to-service path. - Common wrong turn: pushing a policy change to all 1,000 sidecars simultaneously without a
canary. A syntactically valid but logically wrong authorization policy (an overly broad deny
rule, for example) pushed everywhere at once can cause a fleet-wide outage in the time it takes
the control plane to distribute the update, which is why staged rollout applies to policy
changes, not just to the initial mesh adoption.
Unlock Full Question Bank
Get access to all 25 API Security, Authentication and Authorization interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.