API Security, Authentication and Authorization Questions
Controlling who can call an API, what they may do, and defending it against abuse. Covers the access-control mechanics: API keys, OAuth 2.0 flows, OpenID Connect, JWT issuance/validation, session vs. token auth, scopes/roles for fine-grained authorization, token lifetime and refresh, mutual TLS, and machine-to-machine vs. user-delegated access. Also covers the adversarial hardening view: input validation, injection and deserialization risks, broken object-level authorization (BOLA), mass assignment, secrets handling, and the OWASP API Security Top 10, plus securing data in transit, preventing enumeration/scraping, and testing APIs for vulnerabilities.
What is an API gateway, and what security responsibilities does it typically take on for the services sitting behind it?
Sample Answer
Direct answer
An API gateway is a single entry point that sits in front of a set of backend services and handles cross-cutting concerns, routing, traffic management, and security, before a request ever reaches business logic. On the security side, its main value is enforcing cheap, universal checks exactly once, instead of every individual service having to duplicate that logic.
Structured elaboration
Typical security responsibilities a gateway takes on:
- Transport security: terminating TLS (Transport Layer Security) and enforcing HTTPS-only, including a minimum supported TLS version.
- Coarse authentication: validating a token's signature and expiry, or checking an API key, before forwarding the request at all.
- Coarse authorization: confirming this identity is allowed to call this route in general, not the fine-grained "does this caller own this specific record" check, which still belongs inside the service.
- Rate limiting and throttling: protecting the whole platform from abuse or an accidental traffic spike from any single caller.
- Request validation: rejecting malformed, oversized, or wrong-content-type requests early, before they cost any downstream service compute.
- IP allow/deny lists and basic filtering: blocking known-bad sources or request patterns before they reach anything meaningful.
- Centralized audit logging: one place to see every request that entered the system, useful for both security review and debugging.
- Identity propagation: injecting a verified identity (for example, a user-id header) into the forwarded request, so downstream services trust the gateway's verification rather than each re-parsing raw tokens themselves.
Worked example
A POST /orders request arrives with no authentication token. The gateway rejects it with a 401 response before the orders service, inventory service, or payment service ever see it, so none of them spend any compute on traffic that was never going to be allowed. A validly authenticated request for the same endpoint is forwarded along with an X-User-Id header the gateway added after verifying the token, so the orders service can trust that identity without re-validating the raw token itself.
Trade-offs & pitfalls
A gateway should take on responsibilities that are universal and cheap to check without business context, not fine-grained, per-resource authorization, which needs domain knowledge only the service actually has. The most common pitfall is treating the gateway as the only layer of defense: it is a first filter that removes obviously bad traffic early, not a substitute for a service independently checking that a specific caller is allowed to touch a specific resource.
Propose a policy-as-code solution (for example using Open Policy Agent - OPA) to enforce attribute-based access control across APIs. Describe where policies should evaluate (sidecar, gateway, service), policy distribution/versioning, performance strategies (caching/compile-time optimizations), CI testing of policies, and how to secure attribute sources and mitigate stale attributes.
Sample Answer
Direct answer
A policy-as-code approach to attribute-based access control (ABAC, where access decisions depend
on attributes of the user, resource, and context rather than a fixed role list) centers on Open
Policy Agent (OPA), a general-purpose policy engine, evaluating policies written in its Rego
language as close to the request as latency allows, most often colocated in a sidecar next to
each service, with policies distributed as versioned bundles from a central server and pulled
(or pushed) to every OPA instance on a schedule.
Structured elaboration
Where policies should evaluate. There are three realistic placement options, and they are
not mutually exclusive:
- Gateway: good for coarse, cross-cutting checks (is this caller authenticated at all, is
this API version deprecated) that apply uniformly regardless of which backend service handles
the request. - Sidecar (colocated with each service): the most common choice for fine-grained
authorization, since it evaluates in-process (a local network call at most) rather than
requiring a round trip to a shared service, and each service's sidecar can be configured with
only the policy bundle relevant to that service. - In-service (embedded as a library): lowest latency of all (no network hop, not even to a
sidecar) but couples the policy engine's lifecycle to the service's own deploys, which fits
latency-critical paths but loses some of the operational uniformity a sidecar gives you.
The general principle: evaluate as close to the request as the latency budget allows, and prefer
the sidecar pattern by default since it keeps policy evaluation decoupled from application code
without paying a network round trip to a remote service.
Policy distribution and versioning. Author policies as code, version them in the same way
application code is versioned (a git repository, code review, CI checks), and package them into
signed, versioned bundles that a bundle server serves to every OPA instance. Each OPA instance
polls for and pulls new bundle versions on an interval (or the server pushes them), so a policy
change rolls out without redeploying the services that consume it; the bundle's version is
itself an auditable artifact, so "which policy was in effect for this request" is answerable
after the fact.
Performance strategies. OPA compiles Rego policies before evaluation and can serve most
authorization decisions in well under a millisecond once a bundle is loaded, since the decision
is a local, in-memory evaluation, not a network call to an external service. For very
high-throughput paths, further techniques include partial evaluation (precompiling a policy
against known-at-compile-time inputs to shrink the runtime decision), and caching decisions for
identical inputs over a short, explicitly bounded window when the policy and the underlying
attributes are not expected to change within that window.
CI testing of policies. Treat Rego policies exactly like application code under test: write
unit tests (OPA's own test framework, opa test) asserting that specific input attribute
combinations produce the expected allow/deny decision, including deliberately adversarial and
edge-case inputs (a request with a missing or malformed attribute, a role that should never gain
a specific permission), and run those tests in CI before a policy bundle is built and published,
the same gate application code changes go through.
Securing attribute sources and mitigating stale attributes. ABAC decisions are only as
trustworthy as the attributes feeding them (a user's department, a resource's sensitivity
classification, a device's posture); if OPA is fed attributes forwarded from an upstream caller
without verifying their source, a caller can forge a favorable attribute directly. Attributes
should come from a trusted source the policy engine (or the service feeding it) independently
verifies, a signed token's claims, a lookup against a system of record, rather than an
unauthenticated header. Staleness is a related but separate risk: an attribute cached or fetched
minutes ago (a user's role, changed since) can produce a decision based on outdated reality;
mitigate it by keeping high-consequence attributes (role, employment status) on a short refresh
interval or fetched fresh for sensitive operations, while accepting a longer, cheaper refresh
interval for low-consequence, slow-changing attributes (a user's display name).
Worked example
graph TD
Req[API Request] --> GWEnf[Gateway - coarse policy]
Req --> Sidecar[Sidecar - per service policy]
Req --> App[In app - fine grained ABAC]
GWEnf --> OPA[OPA Engine - local decision]
Sidecar --> OPA
App --> OPA
Bundle[Policy Bundle Server] -.->|versioned push| OPA
Attr[Attribute Sources: user, resource, context] --> OPA
A concrete Rego-shaped decision (illustrative, not executed): a request to approve an expense
report carries attributes {user.department: "finance", user.role: "manager", resource.amount: 4500, resource.owner_department: "finance"}. A policy rule such as allow if input.user.role == "manager" and input.user.department == input.resource.owner_department and input.resource.amount <= 5000 evaluates entirely from attributes already present on the
request, with no network call needed at decision time, which is what keeps this pattern fast
enough to sit on the request path even at the sidecar. Changing the approval ceiling from 5000 to
a new value is a policy bundle change, published and rolled out to every sidecar, with no service
code redeployed.
Trade-offs and pitfalls
- ABAC's flexibility is also its main operational risk. Because rules combine arbitrary
attributes instead of a fixed role list, it is easy to write a policy that is logically correct
for the cases you tested and subtly wrong for a combination you did not, which is exactly why
CI testing of policies (including adversarial edge cases) is not optional here the way it might
feel optional for a small, fixed RBAC (role-based access control) rule set. - Colocated evaluation trades a small memory and CPU footprint per service for latency and
availability. A centralized policy-decision service is simpler to operate as a single thing,
but makes every authorization check depend on that service's availability and adds a network
hop to every request; the sidecar pattern avoids both costs at the price of running many small
OPA instances instead of one larger one. - Common wrong turn: trusting client-supplied or upstream-forwarded attributes without
verifying their source. A well-tested policy is still exploitable if an attacker can simply
set the attribute the policy checks; verify attribute provenance, not just the policy logic
itself. - Common wrong turn: refreshing all attributes on the same schedule. Treating a user's
display name and a user's active-employment status as equally safe to cache for, say, an hour
ignores that the second one directly gates access; calibrate refresh/staleness tolerance per
attribute by how consequential it is, not uniformly.
Implement a thread-safe in-memory token-bucket rate limiter in Python. Provide functions:
- set_rate(key: str, tokens_per_sec: float, burst: int)
- acquire(key: str) -> bool
Requirements: allow bursts up to 'burst', refill at tokens_per_sec, be concurrency-safe (use threading.Lock), and avoid unbounded memory growth (evict idle keys).
Sample Answer
Direct answer
A token-bucket limiter keyed by client gives each key its own bucket that starts full (up to
burst tokens) and refills continuously at tokens_per_sec; acquire() lazily refills the
bucket based on elapsed time since it was last touched, then takes one token if at least one is
available. Concurrency safety comes from guarding all bucket reads and writes with a single
lock (or one lock per bucket, for less contention under many distinct keys); unbounded growth is
avoided by tracking each bucket's last-used time and periodically evicting buckets that have
been idle past a TTL.
Structured elaboration
Data model. Each key maps to a small record: current tokens, the configured rate and
burst, and two timestamps: last_refill (last time tokens were topped up) and last_used
(last time the key was touched at all, which drives eviction). Storing floating-point tokens
rather than integers lets sub-second refills accumulate correctly instead of always rounding to
zero.
Lazy refill. Rather than running a background thread that ticks every bucket on a timer
(wasteful for keys nobody is calling), acquire() computes elapsed time since the bucket's
last_refill, adds elapsed * rate tokens capped at burst, and only then checks whether a
token is available. This keeps idle keys cheap: a key that goes untouched for an hour costs
nothing until the next acquire() call recomputes its refill in one step.
Thread safety. All bucket state (create-if-missing, refill, and decrement) must happen
under one critical section per key, or two threads racing on the same key could both read
"1 token available" and both decrement, granting two requests off of one token. A single
process-wide threading.Lock around all bucket operations is simplest and correct; if lock
contention across many distinct keys becomes a bottleneck, a lock striped by key (or a lock
embedded in each bucket's own record) reduces contention while keeping each individual key's
operations atomic.
Bounded memory. Without eviction, every distinct key you ever see creates a permanent
dictionary entry, which is unbounded if keys are, for example, per-IP or per-API-key with high
cardinality. Track last_used per bucket and provide a sweep (evict_idle) that drops any
bucket untouched for longer than an idle TTL; call this sweep periodically (on a timer, or
opportunistically on a fraction of acquire() calls) rather than on every single call, so the
sweep cost is amortized.
Reconfiguration. set_rate() on an existing key should not reset an in-flight bucket back to
full, since that would let a client dodge a rate cut by triggering a reconfiguration; it should
keep the current fill level but clamp it down to the new burst ceiling if the new burst is
smaller than the current token count.
Worked example
import threading
import time
class TokenBucketLimiter:
# Thread-safe in-memory token-bucket rate limiter, keyed by client id.
# Each key gets its own bucket holding up to `burst` tokens, refilling
# at `tokens_per_sec`. An idle bucket (no acquire() past `idle_ttl`
# seconds) is evicted on the next sweep so memory does not grow forever.
def __init__(self, idle_ttl=300.0, clock=time.monotonic):
self._clock = clock # injected so a test can control time deterministically
self._idle_ttl = idle_ttl
self._lock = threading.Lock()
self._buckets = {} # key -> {tokens, rate, burst, last_refill, last_used}
def set_rate(self, key: str, tokens_per_sec: float, burst: int) -> None:
with self._lock:
now = self._clock()
bucket = self._buckets.get(key)
if bucket is None:
self._buckets[key] = {
"tokens": float(burst), "rate": tokens_per_sec, "burst": burst,
"last_refill": now, "last_used": now,
}
else:
bucket["rate"] = tokens_per_sec
bucket["burst"] = burst
bucket["tokens"] = min(bucket["tokens"], burst)
def acquire(self, key: str) -> bool:
with self._lock:
bucket = self._buckets.get(key)
if bucket is None:
return False # unknown key: no rate configured, deny by default
now = self._clock()
elapsed = now - bucket["last_refill"]
if elapsed > 0:
refill = elapsed * bucket["rate"]
bucket["tokens"] = min(bucket["burst"], bucket["tokens"] + refill)
bucket["last_refill"] = now
bucket["last_used"] = now
if bucket["tokens"] >= 1.0:
bucket["tokens"] -= 1.0
return True
return False
def evict_idle(self) -> int:
with self._lock:
now = self._clock()
stale = [k for k, b in self._buckets.items() if now - b["last_used"] > self._idle_ttl]
for k in stale:
del self._buckets[k]
return len(stale)
def bucket_count(self) -> int:
with self._lock:
return len(self._buckets)
# deterministic fake clock so the demo is reproducible: no real sleeping
fake_time = [0.0]
def clock():
return fake_time[0]
limiter = TokenBucketLimiter(idle_ttl=10.0, clock=clock)
limiter.set_rate("user-A", tokens_per_sec=2.0, burst=3)
# burst=3: first 3 calls at t=0 should succeed, the 4th should be denied
results_at_t0 = [limiter.acquire("user-A") for _ in range(4)]
print("t=0 acquires (burst=3):", results_at_t0)
# advance the fake clock by 1.0s -> refill = 1.0 * 2.0 tokens/sec = 2 tokens
fake_time[0] += 1.0
results_after_1s = [limiter.acquire("user-A") for _ in range(3)]
print("t=1.0 acquires (2 tokens refilled):", results_after_1s)
# concurrency check: 50 threads race for a bucket with burst=10, rate=0
limiter.set_rate("user-B", tokens_per_sec=0.0, burst=10)
granted = []
glock = threading.Lock()
def worker():
ok = limiter.acquire("user-B")
with glock:
granted.append(ok)
threads = [threading.Thread(target=worker) for _ in range(50)]
for t in threads: t.start()
for t in threads: t.join()
print("concurrent acquires granted out of 50 (burst=10, rate=0):", sum(granted))
# eviction: user-C goes idle past the 10s ttl while user-A stays active
limiter.set_rate("user-C", tokens_per_sec=1.0, burst=1)
print("bucket_count before idle advance:", limiter.bucket_count())
fake_time[0] += 11.0
limiter.acquire("user-A") # touches user-A so it stays "used" at the new time
evicted = limiter.evict_idle()
print("evicted idle buckets:", evicted)
print("bucket_count after eviction:", limiter.bucket_count())
Output:
t=0 acquires (burst=3): [True, True, True, False]
t=1.0 acquires (2 tokens refilled): [True, True, False]
concurrent acquires granted out of 50 (burst=10, rate=0): 10
bucket_count before idle advance: 3
evicted idle buckets: 2
bucket_count after eviction: 1
The first 3 calls consume the full burst and the 4th is denied, exactly matching burst=3.
After advancing the fake clock by 1 second at tokens_per_sec=2.0, exactly 2 tokens refill, so 2
of the next 3 calls succeed. With rate=0 (no refill) and burst=10, 50 threads racing
concurrently still only grant exactly 10 tokens total, which is the direct evidence that the
lock is preventing a race where more than burst requests get through. Finally, after advancing
past the 10-second idle TTL and touching only user-A, eviction removes the 2 buckets
(user-B, user-C) that went untouched, leaving only the 1 active bucket.
Trade-offs and pitfalls
- A single global lock is simple but becomes a bottleneck under many distinct keys with high
concurrency. If profiling shows lock contention, the fix is per-key locking (aLockinside
each bucket record, taken only for that bucket's own read-modify-write) or sharding the key
space across several independently-locked dictionaries, not removing the lock. - In-memory state means the limiter is per-process. Behind a load balancer with multiple
application instances, each instance enforces its own independent limit, so the effective
limit across the fleet is roughlyconfigured_limit * instance_count. That is often
acceptable for a soft, per-instance guard, but a hard global limit needs a shared store
(Redis, for example) with the same token-bucket logic implemented atomically there instead. - Common wrong turn: resetting
tokenstoburston everyset_rate()call. That lets a
client bypass a rate reduction simply by triggering any reconfiguration; only clamp the
existing token count down to the new burst ceiling, never reset it upward. - Pitfall: forgetting the "unknown key" case. Returning
True(allow) for a key with no
configured rate is a fail-open default that silently disables rate limiting for anything not
explicitly configured; the implementation above fails closed (denies) instead, which is the
safer default for a security-relevant control.
You need an automated emergency revocation and credential rotation plan for compromised client credentials affecting thousands of clients. What would you build? Include detection triggers, mass-revocation mechanics, phased rotation, client notification strategies, fallback modes to preserve critical functionality, and automated rollback if revocations cause unintended outages.
Sample Answer
Direct answer
At thousands of affected clients, a single big-bang revocation is itself a risk: if the
detection was a false positive, or if revocation triggers an unexpected downstream failure, you
have just caused a self-inflicted mass outage. The right design is a phased rollout (a small
canary batch revoked and monitored first, then the rest) with an automated error-rate check
gating the next phase, paired with a client-notification channel and a fallback mode that keeps
critical functionality alive during the rotation window.
Structured elaboration
Detection triggers. The plan starts before revocation: what actually declares "these
credentials are compromised." Realistic triggers include a leaked-secret scanner finding a key
committed to a public repo, an anomaly-detection alert showing a batch of credentials being used
from unexpected, correlated locations simultaneously, or a third-party breach disclosure naming
your credentials among leaked data. The trigger's confidence level should influence the response
speed: a directly confirmed leak (found in a public repo) justifies immediate action, while a
lower-confidence anomaly might justify starting the canary phase rather than a full mass
revocation.
Classify blast radius before acting. Determine which specific credentials, and which
clients, are actually affected, rather than revoking broadly "to be safe." Over-broad revocation
turns a contained incident into a bigger outage than the original compromise would have caused,
and makes the phased rollout's canary step meaningless if "the canary" is actually the entire
affected population already.
Mass-revocation mechanics and phased rotation. Revoke in stages: a small canary slice first
(for example 5% of affected clients), with automated monitoring of the immediate downstream
effect (authentication error rates, support ticket volume, dependent-service health) before
proceeding. If the canary phase looks clean, proceed to the remaining clients, ideally still in
batches rather than one further big step, so a problem discovered at 30% is cheaper to stop and
diagnose than one discovered at 100%.
Client notification strategies. Notify affected clients through more than one channel
(email, an in-dashboard banner, a status-page entry, and for programmatic integrations, an API
response that clearly signals "this credential is revoked, obtain a new one here" rather than a
generic authentication failure) since a silent revocation just looks like an outage from the
client's side and generates support load instead of self-service recovery.
Fallback modes to preserve critical functionality. For any client or integration where a hard
cutoff would cause serious harm (a payments integration going dark mid-transaction, for example),
consider a bounded grace period where the old credential is accepted for a small set of
lower-risk operations only (read-only calls, say) while write or sensitive operations require the
new credential immediately; this is a deliberate, scoped exception, not a general delay of the
whole revocation.
Automated rollback if revocations cause unintended outages. The phased design's real payoff
is here: the automated check gating each phase (error rate crossing a threshold, dependent-service
health degrading) should be able to pause or reverse the rollout automatically, re-enabling the
just-revoked batch's old credentials temporarily while the team investigates, rather than
requiring a human to notice the outage and intervene manually before further damage accrues.
Worked example
graph TD
Det[Detection: leak trigger] --> Class[Classify blast radius]
Class --> Phase1[Phase 1: revoke 5 percent canary]
Phase1 --> Check1{Error rate ok?}
Check1 -->|yes| Phase2[Phase 2: revoke remaining 95 percent]
Check1 -->|no| Rollback[Automated rollback plus alert]
Phase2 --> Notify[Client notification plus new credential issuance]
Notify --> Verify[Verify uptake, monitor auth failures]
With 10,000 affected client credentials: Phase 1 revokes 500 clients (5%) and holds for a defined
bake time (for example 15 minutes) while monitoring authentication error rates specifically among
that revoked cohort's expected traffic, not the whole system's aggregate rate, since a 5%
increase in a tiny slice can be invisible in an aggregate metric. If that cohort's error rate
stays within the expected "credential revoked, client has not rotated yet" range, and no
unrelated service's health degrades, Phase 2 proceeds to revoke the remaining 9,500. If instead
the canary phase shows an unexpected spike (say, a shared downstream dependency that was not
accounted for in the blast-radius classification starts failing), the automated gate halts the
rollout at 500 revoked rather than proceeding to 10,000, and the rollback path re-enables that
canary batch's old credentials temporarily while the team investigates what the classification
step missed.
Trade-offs and pitfalls
- A canary phase adds real time to full remediation, which is a genuine cost when the
underlying leak is actively being exploited; the trade-off against a faster full revocation is
explicit and should be a judgment call based on the trigger's confidence level (a confirmed
active exploitation may justify skipping straight to broader revocation despite the added
risk, where a lower-confidence anomaly does not). - A fallback mode that keeps old credentials working for "low-risk" operations is itself an
attack-surface decision, not a free safety net. If the attacker who obtained the leaked
credential can still use it for anything at all during the grace period, define precisely what
"low-risk" means for your system, since a wrong classification here (treating a read operation
that returns sensitive data as low-risk, for example) undermines the point of revoking at all. - Common wrong turn: treating rollback as a manual, human-triggered step. At the scale of
thousands of affected clients, by the time a human notices a rollout-caused outage and decides
to intervene, meaningful damage has usually already accrued; the automated gate needs the
authority to pause or reverse on its own, with the human notified, not consulted first. - Common wrong turn: measuring rollout health against a system-wide aggregate metric instead of
the specific affected cohort. A canary batch's problems are easy to miss in a global error
rate exactly because the canary is, by design, a small slice of total traffic.
You operate a mixed monolith + microservices environment. For security controls (authentication, authorization, rate limiting, input/schema validation, transport security), decide which responsibilities should be enforced at the API gateway/proxy and which should remain inside services. Justify choices with availability, security, and performance trade-offs and propose testing and observability to validate enforcement.
Sample Answer
Direct answer
Push the checks that are cheap, universal, and identity-only to the gateway (authentication verification, transport security, coarse rate limiting, structural request validation), and keep the checks that need business or ownership context inside the service (fine-grained authorization, business-aware rate limiting, semantic validation). Never treat the gateway's decision as the only check: services should still re-verify authorization even for traffic they assume already passed the gateway, because a bypass or misconfiguration at the gateway should not be the single point of failure for data access.
Structured elaboration
| Control | At the gateway | Inside the service | Why the split |
|---|---|---|---|
| Authentication | Verify token signature and expiry once, forward a verified identity to downstream services | Trust the forwarded identity, do not re-parse raw credentials | Cheap and identical for every route; centralizing it avoids every service wiring its own token validation |
| Authorization | Coarse: is this identity allowed to call this route at all | Fine-grained: does this specific caller own this specific resource | The gateway has no idea which record ID belongs to which user; that context lives only in the service |
| Rate limiting | Global/per-key throttling to protect the whole platform from volume abuse | Business-aware limiting (e.g. failed login attempts per account) that needs domain state | The gateway can count requests; it cannot reason about "too many wrong passwords for this specific account" |
| Input/schema validation | Structural: does the payload match the declared shape, reject malformed junk early | Semantic: is this SKU real, is there enough stock, do these fields make business sense together | Structural checks are cheap and universal; semantic checks require the service's own data |
| Transport security | TLS termination for client-to-gateway traffic | Mutual TLS or a service mesh for gateway-to-service and service-to-service hops | Encrypting only the outer hop leaves the internal network unauthenticated between services |
Worked example
A checkout request hits the gateway without a valid token: the gateway rejects with 401 before the order service, inventory service, or payment service ever see it, saving three services from processing traffic that was never going to be allowed. A validly authenticated request for PATCH /orders/482 is forwarded with a verified user-id header; the order service still checks that the caller actually owns order 482 before applying the patch, because the gateway only confirmed "this is a real, authenticated user," not "this user owns this specific order." If that service-side ownership check were removed on the assumption the gateway already handled it, any authenticated user could edit any order by changing the ID in the URL, which is exactly the Broken Object Level Authorization pattern, one of the most reliably tested authorization failures in API interviews at any level.
Trade-offs & pitfalls
Availability: a gateway that is horizontally scaled and stateless is not a bigger single point of failure than any other tier, but a gateway that grows to hold business logic becomes a deploy bottleneck for every service behind it, so keep it thin on purpose.
Performance: the extra hop at the gateway adds latency, but it is repaid by filtering malformed or unauthenticated traffic before it costs any downstream service compute, which is a net win under load.
Security: defense in depth means the service re-checks authorization even though the gateway already made a coarse allow decision. A common pitfall is deleting service-side authorization checks because "the gateway already checks it," which turns a misrouted internal call, a compromised adjacent service, or a gateway misconfiguration into a full authorization bypass.
Testing and observability: run contract tests that assert the gateway's allowed routes and the service's own authorization rules agree, so they cannot silently drift apart. Run synthetic BOLA probes directly against services in a staging environment, bypassing the gateway entirely, to confirm services do not rely on the gateway as their only defense. Correlate logs across the gateway and every service hop with a shared request ID, track 401/403 rates and rate-limit rejections per route as standing metrics, and alert when the gateway's allow decision and a service's deny decision disagree for the same request, since that disagreement is itself a signal of policy drift between the two layers.
Unlock Full Question Bank
Get access to all 17 API Security, Authentication and Authorization interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.