API Security, Authentication and Authorization Questions
Controlling who can call an API, what they may do, and defending it against abuse. Covers the access-control mechanics: API keys, OAuth 2.0 flows, OpenID Connect, JWT issuance/validation, session vs. token auth, scopes/roles for fine-grained authorization, token lifetime and refresh, mutual TLS, and machine-to-machine vs. user-delegated access. Also covers the adversarial hardening view: input validation, injection and deserialization risks, broken object-level authorization (BOLA), mass assignment, secrets handling, and the OWASP API Security Top 10, plus securing data in transit, preventing enumeration/scraping, and testing APIs for vulnerabilities.
Build an automated test to detect timing side-channel leaks in an authentication API where the response time might differ based on username validity. Explain how you would reduce noise, choose sample sizes, apply statistical analysis, and automate detection in CI without producing many false positives due to network variance.
Sample Answer
Direct answer
A timing side-channel test compares the distribution of response times for two request
classes (valid username vs. invalid username), not a single measurement, because one request
is dominated by network jitter. Build it as: collect many interleaved paired samples, trim
outliers, run a distribution-comparison statistic that does not assume clean symmetric noise,
and require the result to repeat across several independent CI runs before failing a build, so
one noisy run cannot trip a false positive.
Structured elaboration
Why the leak exists. An endpoint that does extra work only when a username is real (looks
up a stored password hash, runs a bcrypt/argon2 compare) systematically takes longer for valid
usernames, even when the HTTP response body and status code look identical. That gap lets an
attacker enumerate real usernames purely from timing, with no error message needed.
Reduce noise before comparing anything:
- Warm up the service and discard the first N requests (cold caches, JIT/interpreter warmup,
TCP slow start all inflate early samples). - Interleave the two request types (alternate valid, invalid, valid, invalid...) instead of
running one whole batch then the other. Load drift, garbage collection pauses, and neighbor
noisy-tenant effects then hit both groups equally instead of biasing one. - Trim outliers (for example the top and bottom 5%) before computing a summary statistic.
Network jitter and retransmits produce occasional huge spikes that would otherwise swamp a
genuine sub-millisecond gap.
Choosing a sample size. The bigger the real timing gap relative to the noise, the fewer
samples you need to detect it reliably. A gap of a few milliseconds against light noise might
separate cleanly with a few dozen samples; a sub-millisecond gap buried in several milliseconds
of network jitter (the realistic case for a well-behaved API) typically needs several hundred
to a few thousand samples per group. Start around 300 to 500 per group, and increase it if
independent runs disagree with each other.
Statistical analysis. Do not compare raw means with a standard two-sample t-test (a statistic that compares two groups' averages, assuming the noise around each is roughly symmetric and bell-curve-shaped): network
latency is right-skewed (a long tail of slow outliers, not a symmetric bell curve), which
violates the t-test's assumptions and makes it unreliable here. A permutation test on the
trimmed data is a better fit: it makes no assumption about the shape of the noise, works
directly on the trimmed samples, and gives an empirical p-value (the probability of seeing a gap at least this large purely by chance, if there were no real timing difference) by asking "how often would a
gap this large appear if the labels were random?"
Automate in CI without false positives from network variance:
- Fix an explicit significance threshold (a low one, since you will run the test often) and
require several independent runs to agree before failing the build. A single p < 0.05 result
will happen by chance on noisy timing data with some regularity. - Run on a dedicated or otherwise quiet CI runner where possible, so background load is not
itself a confound. - Include a negative control in the same pipeline: the same statistical test comparing
valid-vs-valid samples should almost never fire. If it does, your pipeline's noise floor is
too high to trust the real test's result that day.
Worked example
The script below simulates response times for 400 requests per group: a shared 12ms base cost,
a real 0.6ms leak injected only into the "invalid username" branch (modeling an extra hashed
comparison), and network jitter modeled as a right-skewed exponential distribution (mean 3ms)
rather than a symmetric one, since that is closer to real network noise. It trims the top and
bottom 5% of each group, then runs a one-sided permutation test. It checks 5 independent
"CI runs" (5 different random seeds) and only flags a leak if at least 4 of 5 agree at
p < 0.01, then repeats the whole pipeline with no injected leak as a negative control.
import random
import statistics
random.seed(42)
N = 400 # samples per group
def sample_response_times(base_ms, leak_ms, n):
# base_ms: shared processing cost. leak_ms: the side-channel gap this
# branch adds (0 for the "valid username" baseline). Jitter is modeled
# as a right-skewed exponential, closer to real network noise than Gaussian.
out = []
for _ in range(n):
jitter = random.expovariate(1 / 3.0) # mean 3ms right-skewed jitter
core = random.gauss(base_ms, 0.4) # small stable variance in app code
out.append(core + leak_ms + jitter)
return out
def trim(samples, pct=0.05):
# drop the top/bottom pct of samples to blunt network-spike outliers
s = sorted(samples)
k = int(len(s) * pct)
return s[k: len(s) - k] if k > 0 else s
def permutation_test(a, b, n_perm=2000, seed=0):
# one-sided permutation test for mean(b) > mean(a)
rng = random.Random(seed)
observed = statistics.mean(b) - statistics.mean(a)
pooled = a + b
na = len(a)
count_ge = 0
for _ in range(n_perm):
rng.shuffle(pooled)
perm_a = pooled[:na]
perm_b = pooled[na:]
diff = statistics.mean(perm_b) - statistics.mean(perm_a)
if diff >= observed:
count_ge += 1
p_value = (count_ge + 1) / (n_perm + 1) # avoids an impossible p=0
return observed, p_value
def run_one_ci_check(leak_ms, run_seed):
random.seed(run_seed)
valid = sample_response_times(base_ms=12.0, leak_ms=0.0, n=N)
invalid = sample_response_times(base_ms=12.0, leak_ms=leak_ms, n=N)
observed_ms, p = permutation_test(trim(valid), trim(invalid), n_perm=2000, seed=run_seed)
return observed_ms, p
# a real 0.6ms leak (e.g. an extra bcrypt compare only run when the username
# exists) buried under ~3ms mean network jitter, checked across 5 independent
# CI runs (5 seeds) instead of trusting a single run
alpha = 0.01
significant_runs = 0
print("run observed_diff_ms p_value significant(p<0.01)")
for i, seed in enumerate([1, 2, 3, 4, 5]):
observed_ms, p = run_one_ci_check(leak_ms=0.6, run_seed=seed)
sig = p < alpha
significant_runs += int(sig)
print(f"{i+1:>3} {observed_ms:>16.3f} {p:>7.4f} {sig}")
print(f"\nsignificant in {significant_runs}/5 runs (require >=4/5 to flag in CI)")
print("FLAG TIMING LEAK" if significant_runs >= 4 else "no consistent leak detected")
# negative control: no injected leak, same pipeline, to show the gate does
# not fire on jitter alone
print("\nnegative control (leak_ms=0.0):")
significant_runs_ctrl = 0
for i, seed in enumerate([11, 12, 13, 14, 15]):
observed_ms, p = run_one_ci_check(leak_ms=0.0, run_seed=seed)
sig = p < alpha
significant_runs_ctrl += int(sig)
print(f"{i+1:>3} {observed_ms:>16.3f} {p:>7.4f} {sig}")
print(f"significant in {significant_runs_ctrl}/5 runs (require >=4/5 to flag in CI)")
print("FLAG TIMING LEAK" if significant_runs_ctrl >= 4 else "no consistent leak detected (correct: no leak was injected)")
Output:
run observed_diff_ms p_value significant(p<0.01)
1 0.433 0.0045 True
2 0.696 0.0005 True
3 0.627 0.0005 True
4 0.771 0.0005 True
5 0.318 0.0170 False
significant in 4/5 runs (require >=4/5 to flag in CI)
FLAG TIMING LEAK
negative control (leak_ms=0.0):
1 -0.014 0.5402 False
2 -0.215 0.8966 False
3 -0.068 0.6727 False
4 -0.070 0.6737 False
5 -0.051 0.6207 False
significant in 0/5 runs (require >=4/5 to flag in CI)
no consistent leak detected (correct: no leak was injected)
With a real 0.6ms leak, 4 of 5 simulated CI runs cross the p < 0.01 threshold, so the gate
correctly fires. With no injected leak, all 5 runs come back non-significant, so the negative
control confirms the pipeline does not fire on jitter alone. Run 5 in the first block (p =
0.017) illustrates exactly why a single-run threshold is unsafe: that one run alone would have
been called "not significant" at the 0.01 level even though the leak was real, which is why the
gate requires 4 of 5 runs to agree rather than trusting any single run.
Trade-offs and pitfalls
- Statistical detection is not the same as measuring real-world exploitability. This test
tells you a leak likely exists locally; it does not tell you how many requests an attacker
over the public internet, with far more noise than your CI network, would need to exploit it.
Those are two different measurements, and a leak that is statistically detectable in a quiet
CI environment may be much harder to exploit remotely. - Common wrong turn: comparing raw means on unfiltered data with a standard t-test. Network
jitter is heavy-tailed, so a handful of retransmits can swing a raw mean far more than a
genuine sub-millisecond leak, producing both false positives and false negatives. - Common wrong turn: trusting a single p-value. Noisy timing measurements will occasionally
produce a "significant" result by chance; requiring several independent runs to agree, as in
the worked example, controls that risk in a way a single run cannot. - Pitfall: "fixing" the endpoint's happy path but not its neighbors. If you make the
invalid-username branch add a matching artificial delay but a cache layer, a database index,
or connection pooling upstream still behaves differently for existing vs. non-existing users,
the leak often survives at a smaller magnitude. Re-run the same statistical test as a
regression gate after any fix, not just once at discovery time.
Explain Cross-Origin Resource Sharing (CORS): the headers involved, the browser enforcement model, and what security guarantees CORS does and doesn't actually provide. Then walk through why a wildcard Access-Control-Allow-Origin combined with Access-Control-Allow-Credentials: true is a dangerous configuration, and how you'd safely configure CORS on an API that uses cookies or bearer tokens.
Sample Answer
Direct answer
Cross-Origin Resource Sharing (CORS) is a browser-enforced relaxation of the same-origin policy that lets a web page on one origin ask a server on another origin to opt in to being called from JavaScript. It is enforced entirely client-side by the browser, so it protects browser-based callers from a malicious page, but it does nothing to stop a non-browser caller, curl, another server, a script, since none of those ever consult CORS headers at all.
Structured elaboration
Headers involved
- Request side:
Origin, the browser telling the server where the request came from. - Response side:
Access-Control-Allow-Origin(which origin or origins may read the response),Access-Control-Allow-Credentials(whether cookies or HTTP auth may be included), andAccess-Control-Allow-Methods/Access-Control-Allow-Headers(what a preflight check actually permits). - For non-simple requests, custom headers, methods like
PUTorDELETE, certain content types, the browser first sends anOPTIONSpreflight request and only proceeds with the real request if the preflight response allows it.
Browser enforcement model
This is the most commonly misunderstood part: for many requests, the server still processes and can still respond to a cross-origin call even when the origin isn't allowed. What CORS actually blocks is the browser handing that response back to the calling page's JavaScript. CORS is a response-reading gate enforced by browsers, not a request-blocking firewall, and it provides zero protection against server-to-server calls or any tool that simply does not implement browser CORS rules.
What CORS does and doesn't guarantee
- Does: stop a malicious website from using a logged-in victim's browser and its ambient cookies to read data back cross-origin via JavaScript, when configured correctly.
- Doesn't: authenticate or authorize anyone, protect an API from direct non-browser calls, or substitute for real server-side access control.
Why the wildcard-plus-credentials combination is dangerous
Access-Control-Allow-Origin: * together with Access-Control-Allow-Credentials: true is explicitly disallowed by the CORS specification itself, browsers refuse to honor that exact literal combination. But a very common misconfiguration achieves the same dangerous effect without triggering that block: reflecting whatever Origin header the request sent back as the allowed origin, for every single request, while also setting Access-Control-Allow-Credentials: true. That passes spec validation (it is technically not a literal wildcard) but functionally means any website in the world can make a credentialed, cookie-carrying request to this API and read the response, which lets a malicious site silently ride a logged-in user's session and exfiltrate their data.
Safe configuration for cookie or bearer-token APIs
Maintain an explicit allowlist of known, trusted origins, your own frontend domains and named partner domains, and validate the incoming Origin header against that list, only echoing it back, never a wildcard, when it actually matches. Set Access-Control-Allow-Credentials: true only on that narrow, allowlist-matched response, never combined with a reflect-anything policy. For bearer-token APIs that do not rely on cookies at all, sending the token in an Authorization header instead, you can often skip credentialed CORS entirely, since a stolen or reflected CORS configuration cannot ambiently attach a header a malicious page does not know to send.
Worked example
Dangerous configuration: a request arrives with Origin: https://evil.example, and the server responds with Access-Control-Allow-Origin: https://evil.example plus Access-Control-Allow-Credentials: true, for every incoming origin, with no allowlist check at all.
Safe configuration: the server checks the incoming Origin against an explicit list, ["https://app.example.com", "https://partner.example.com"], echoes it back only on a match, and either returns a 4xx or simply omits the CORS headers (which the browser then treats as a block) when there is no match.
Trade-offs & pitfalls
"Just reflect the Origin header, it's easier than maintaining an allowlist" is precisely the dangerous shortcut described above. CORS misconfiguration is a browser-side control, so testing it with curl or Postman will not reveal the vulnerability the way it actually manifests in the real world; you have to test from an actual disallowed origin inside a browser, or reason directly about the response headers, since curl never enforced CORS in the first place and a working curl test creates a false sense of security.
Walk me through how you'd build a scalable pipeline to detect API abuse (credential stuffing, scraping, fraud) across hundreds of services and millions of requests per minute. Include data collection and enrichment (geo, ASN, device fingerprint), real-time detection and scoring (streaming feature aggregation, ML models), alerting to SIEM/SOAR, automated blocking/lists and the feedback loop for model updates, while preserving low latency on request paths.
Sample Answer
Direct answer
At the scale described (hundreds of services, millions of requests per minute) the pipeline has
to split into a fast synchronous path that adds only a few milliseconds per request, and a
slower asynchronous path that does the expensive enrichment and model scoring off to the side
and feeds its verdicts back as a cache the fast path can check cheaply. You cannot run full
machine-learning scoring inline on every request at that volume without blowing the latency
budget, so the design's central decision is what stays synchronous (a cheap reputation lookup)
versus what happens asynchronously and only changes future requests (enrichment, scoring, model
updates).
Structured elaboration
Sizing the problem first. "Millions of requests per minute" is on the order of tens of
thousands of requests per second; for example 5,000,000 requests/minute is about 83,000
requests/second sustained. Any synchronous, per-request check has to fit inside a latency budget
of a few milliseconds at that rate, which rules out anything that calls out to a heavyweight
model or an external enrichment API inline.
Data collection and enrichment. Capture request metadata at the edge (source IP, user agent,
TLS fingerprint, authenticated identity if any) and enrich it: geolocation and ASN (autonomous
system number, which identifies the network/ISP a request came from) lookups from a local,
periodically-refreshed dataset rather than a live network call; device fingerprinting from
client-side signals where available. Do the enrichment lookups against an in-memory or
local-cache copy of the reference data, not a network round trip per request, since a network
call per request at 83,000 requests/second is its own outage risk.
Real-time detection and scoring. Split scoring into two tiers:
- A cheap, synchronous tier that checks a precomputed verdict (IP or identity already on a
blocklist or flagged as high-risk from a shared, low-latency cache) and applies simple rules
(velocity thresholds) that need no model inference. - An asynchronous tier that aggregates streaming features (request velocity per identity, ratio
of failed to successful auth attempts, geographic dispersion of a single credential's usage)
over sliding windows, and periodically runs those aggregated features through a scoring model.
The model's output updates the shared verdict cache that the synchronous tier reads, so the
request path never blocks on model inference; it only ever blocks on a cache read.
Alerting to SIEM/SOAR. Route confirmed and borderline detections to the security team's SIEM
(security information and event management system, which centralizes security logs for
investigation) and SOAR (security orchestration, automation and response, which can trigger
automated response playbooks) rather than only auto-blocking. High-confidence, high-severity
patterns can trigger automated blocking directly; medium-confidence patterns should raise an
alert for a human or a scripted playbook to act on, since automated blocking on a weak signal
risks blocking real users (a false positive that costs revenue and trust).
Automated blocking and the feedback loop. Blocking decisions (IP bans, credential lockouts,
CAPTCHA challenges) need to be revisable: log every automated action with the signal that caused
it, and feed confirmed false positives (a legitimate user who got blocked and later proved it,
for example by successfully completing account recovery) back into the model's training data or
into rule exceptions, so the system's precision improves rather than accumulating permanent
mistakes.
Preserving low latency on request paths. The single most important architectural rule is
that nothing on the synchronous request path may depend on the availability or latency of the
detection pipeline's slower components. If the verdict cache is unreachable, the request path
should fail open to "no additional friction" (log the miss, do not block), not fail closed to
"block everything," unless the product's risk tolerance explicitly demands the opposite for a
specific high-value action (initiating a payment, for example).
Worked example
graph LR
R[API Request] --> E[Enrichment: geo, ASN, device fingerprint]
E --> F[Streaming Feature Aggregation]
F --> ML[Real time ML Scoring]
ML -->|high risk| B[Auto Block or Challenge]
ML -->|medium risk| SIEM[Alert to SIEM and SOAR]
ML -->|low risk| P[Pass Through]
B --> FB[Feedback Loop]
SIEM --> FB
FB --> ML
Reading the diagram left to right: the synchronous request path is only the leftmost box, since
everything from "Streaming Feature Aggregation" onward runs asynchronously against buffered
data, not inline with the request. At 83,000 requests/second, if the synchronous enrichment plus
a verdict-cache read together cost 2ms, that is a fully absorbable addition to a typical API's
latency budget; if the same path instead waited on the "Real time ML Scoring" box per request,
the pipeline would need that box to sustain 83,000 scoring calls per second with sub-millisecond
latency each, which is why that box is drawn as feeding a cache asynchronously rather than
sitting inline.
Trade-offs and pitfalls
- Fail-open vs. fail-closed under detector outage. Failing open (letting requests through
when the detection pipeline is down) protects availability and revenue but temporarily loses
abuse protection; failing closed protects against abuse but can turn a detection-pipeline
outage into a full product outage. State this trade-off explicitly per endpoint rather than
picking one default for the whole system, since a login endpoint and a public read-only search
endpoint have very different risk profiles. - False positives have a real cost, not just a technical one. An overly aggressive
auto-block tier degrades the experience for legitimate users and generates support load; this
is why the design routes medium-confidence signals to an alert instead of an automatic block,
and why the feedback loop exists at all. - Common wrong turn: scoring every request synchronously "to be safe." This looks more
thorough on paper but does not scale to the stated volume and turns the fraud pipeline into
the system's latency bottleneck; the asynchronous, cache-backed design is not a shortcut, it is
the only version of this architecture that survives the request rate. - Common wrong turn: treating the feedback loop as optional. Without it, the model's
precision decays as attackers adapt and as the false-positive rate silently rises, since
nothing in the system is measuring or correcting for it.
Given a GraphQL mutation that accepts deeply nested input to create users and related resources, perform a threat model that focuses on injection, excessive data exposure, denial-of-service via complex nested queries, and authorization bypass. Propose precise mitigations such as sanitization, field-level authorization hooks, depth/complexity limiting, persisted queries and cost estimation.
Sample Answer
Direct answer
For a GraphQL mutation that accepts deeply nested input to create users and related resources,
the threat model has four connected risks: injection through any field that reaches a data
store or downstream system, excessive data exposure through the response shape a single query
can request, denial-of-service through the cost of resolving deeply nested or high-fan-out
selections, and authorization bypass at the object and field level within the nested structure,
not just at the mutation's top-level entry point. Because a single GraphQL request can touch
many resources and relationships in one call, each of these risks compounds with nesting depth
in a way a single flat REST endpoint does not.
Structured elaboration
Injection. Every leaf value in the nested input (a related resource's name, an address
field several levels deep) is still untrusted input reaching business logic and, eventually,
storage; the nesting does not change the injection risk, it just multiplies how many fields need
the same discipline applied consistently: parameterized queries, never string-built ones, plus
strict schema validation on every nested object (not only the mutation's top-level arguments,
which is where a reviewer's eye is naturally drawn) and correct output encoding wherever any of
this data is later rendered back out.
Excessive data exposure. Because the same mutation can request a return shape that includes
newly created and related resources, a response can overfetch: it can return internal or
sensitive fields on those related resources that the caller was never meant to see, simply
because the schema allows selecting them and nobody scoped the mutation's response shape as
carefully as its input shape was scoped. Mitigate this the same way field-level authorization
is enforced on queries: role- and ownership-aware field resolvers on the response type, not an
assumption that "this is a write endpoint, so read-side authorization doesn't apply here."
Denial-of-service via complex nested queries. A mutation with deeply nested input, or a
follow-up query selecting deeply nested relationships, can force the server to do
exponentially more work than the request's size on the wire suggests, since each level of
nesting can multiply the resolvers invoked by the requested page size at that level. Mitigate
with depth limiting (rejecting a query or mutation whose selection nests deeper than a configured
maximum) and query complexity/cost analysis (assigning each field a cost, multiplying by
requested list sizes at each level, and rejecting a request whose total cost exceeds a budget
before executing any of it).
Authorization bypass. A nested mutation creating multiple related resources in one call needs
authorization checked at every object being created or referenced, not only at the top-level
"can this user call this mutation at all" gate; a caller authorized to create their own user
profile is not automatically authorized to attach that profile to an organization they do not
belong to, simply because the organization reference is buried three levels into the nested
input rather than being the mutation's top-level argument.
Concrete mitigations, tied to the risks above:
- Sanitization and parameterization at every leaf field, same discipline as any other
user-supplied input reaching a data store. - Field-level authorization hooks on both the input side (can this caller set this field or
reference this related object) and the response side (can this caller see this field on the
result), not only on the mutation as a whole. - Depth and complexity limiting, rejecting requests that exceed a configured nesting depth
or computed cost before execution begins, so a malicious request is rejected cheaply rather
than partially executed before being caught. - Persisted queries (the client sends a reference to a pre-registered, pre-approved query
shape rather than an arbitrary ad-hoc query string), which is a strong mitigation against
unexpected malicious query shapes specifically, since the server only ever executes shapes it
already reviewed and approved, though it does not by itself replace per-request authorization
checks on the data those approved shapes touch. - Cost estimation, the mechanism underlying complexity limiting: assign a numeric cost to
each field (higher for fields that fan out via a list) so the total cost of a specific request
can be computed and compared against a budget before execution.
Operational testing to validate the mitigations. These protections need to be verified the
same way any other security control is verified, with automated tests exercising the negative
cases: a test that submits a mutation nested one level past the configured depth limit and
asserts it is rejected before any database write occurs; a test computing a request's expected
cost against the cost model and asserting the server's rejection threshold matches; and a
field-level authorization test that runs the identical nested mutation as two different callers
(one entitled to set or see a given nested field or object, one not) and asserts the
unauthorized caller's attempt is rejected or the field comes back redacted, since a passing
depth-limit test says nothing about whether the authorization checks inside that depth are
actually being enforced.
Overfetching and introspection abuse, as related but distinct concerns. Overfetching (a
client requesting far more of the graph than its use case needs, even without malicious intent)
is a milder version of the excessive-data-exposure risk above and is best addressed by the same
field-level authorization plus reasonable default response shapes; introspection abuse (an
attacker using the schema's own introspection query to map out the entire graph, including
fields or types not intended for public discovery, as reconnaissance for a later attack) is
mitigated by disabling introspection in production, backed by a checked-in schema snapshot so
tooling and tests still have a field list to work from, or, where some introspection must stay
available, at minimum excluding internal-only types and fields from it.
Worked example
Consider a nested mutation: user { id, posts(first: 20) { id, comments(first: 20) { id } } }.
Using a simple cost model where a scalar field costs 1 point per row in scope and a list field
multiplies the cost of everything beneath it by the requested page size:
user.id : 1 (scope multiplier 1)
user.posts (node) : 1 (scope multiplier 1, before its own multiplier applies)
posts.id : 20 (scope multiplier 20, from posts' page size)
posts.comments (node) : 20 (scope multiplier 20)
comments.id : 400 (scope multiplier 20 * 20, nested page sizes compound)
-----------------------------------------------
total cost : 442
If the server enforces a cost budget of, say, 300 per request, this query is rejected before
execution, purely from its declared shape, with no database call made. Without complexity
limiting, a client (or attacker) could push first: 100 at each level instead of first: 20,
pushing the comments.id term alone to 100 * 100 = 10,000, a strictly wire-cheap request that
would force the server to resolve tens of thousands of rows.
Trade-offs and pitfalls
- Depth and complexity limits are blunt instruments that can reject legitimate, unusually
shaped requests, not just malicious ones; calibrate the budget against your real schema's
legitimate worst-case use cases, and expose a documented way for a legitimate client with a
genuinely larger need to request a higher limit, the same way an API rate-limit exception
process should exist. - Persisted queries strongly constrain query shape but do nothing about authorization on the
data a shape touches; a persisted query approved months ago can still be misused by a
caller who should not have access to the specific objects it happens to reference this time, so
persisted queries reduce, but do not replace, per-request field- and object-level authorization. - Common wrong turn: authorizing only the mutation's top-level entry point. A nested mutation
is not one flat authorization check, it is potentially many, one per object being created or
referenced within the nested input, and skipping the nested ones is exactly the authorization-
bypass risk this threat model calls out. - Common wrong turn: validating shape (depth, complexity) but never testing it. A depth limit
that was implemented but never covered by a test asserting it actually rejects an over-depth
request is a control that looks present in code review and absent in practice the first time it
matters.
Propose detection and mitigation strategies for credential stuffing and automated account takeover attempts on login endpoints. Propose telemetry signals (IP velocity, failed-login patterns, device fingerprinting), anomaly detection heuristics, progressive throttling and challenge mechanisms (CAPTCHA, MFA step-up), and techniques to minimize false positives while blocking automated abuse.
Sample Answer
Direct answer
Credential stuffing (attackers replaying stolen username/password pairs from other breaches
against your login endpoint) and automated account takeover are detected primarily through
behavioral telemetry, not a single rule: velocity per IP and per credential, failed-login
patterns, and device fingerprinting feed an anomaly signal that drives progressive friction
(throttling, then a CAPTCHA, then a step-up to multi-factor authentication, MFA) rather than an
all-or-nothing block. The same abuse-defense posture extends past the login form: public API
endpoints that let a client scrape or excessively read data need their own rate and reputation
controls, and a confirmed compromise needs a remediation path (revoking affected credentials,
notifying affected customers), not just a detection alert.
Structured elaboration
Telemetry signals to collect.
- IP velocity and reputation: request rate per IP, and whether the IP is a known VPN,
residential proxy pool, or Tor exit node, since credential-stuffing tooling routes through
large proxy pools specifically to defeat simple per-IP rate limits. - Failed-login patterns: the tell-tale shape of credential stuffing is a low
success-to-attempt ratio spread across many distinct usernames, each tried only once or
twice, as opposed to a brute-force attack hammering one account with many password guesses; a
detector that only watches "failed attempts against one account" misses stuffing entirely. - Device fingerprinting: signals derived from the client (browser/TLS fingerprint, screen
and font characteristics for a web client) that let you notice the same automated client
reappearing under many different credentials or IPs.
Anomaly detection and progressive response. Combine the signals above into a risk score per
login attempt rather than a single hard rule, and respond proportionally:
- Low risk: allow normally.
- Medium risk: add friction, a CAPTCHA challenge or a short delay, cheap enough that it barely
affects a real user but meaningfully slows down high-volume automation. - High risk: require an MFA step-up even if the password was correct, since a correct password
from a breached credential list does not prove the caller is the legitimate account owner. - Confirmed automation: block or throttle hard at the network layer (the IP or proxy pool).
Minimizing false positives. Progressive challenges exist specifically so that a real user
having a bad day (typo'd password, unfamiliar network) experiences mild friction rather than a
hard lockout, while only sustained, clearly automated patterns escalate to a block. Tune
thresholds against a labeled sample of known-legitimate traffic (users on shared corporate
NATs, for example, who will always look like "many login attempts from one IP") before rolling
out a stricter rule broadly, and alert on the block rate itself, since a spike in legitimate
users getting blocked is a signal the thresholds are miscalibrated, not that abuse suddenly
increased.
Beyond the login endpoint: broader public-API abuse defense. The same abuse-defense posture
applies past login, since a public API that is not itself an auth endpoint can still be scraped
or read excessively (a competitor bulk-harvesting a product catalog or pricing data, for
example): rate limiting and quotas per API key or per identity, the same IP-reputation signals
described above applied to read-heavy endpoints, and CAPTCHA gating specifically on suspicious
UI flows (a signup form, a password-reset form) where a CAPTCHA is tolerable friction, rather
than on every API call, where it would break legitimate programmatic clients entirely. When a
compromise or sustained abuse campaign is confirmed, remediation extends beyond detection:
revoke or force-rotate the specific credentials or API keys involved, and notify affected
customers when their accounts or data were plausibly touched, since detection without a
remediation and notification path leaves the actual harm unaddressed even after the technical
attack is stopped.
Worked example
A login endpoint sees, over one hour: 40,000 login attempts from 6,000 distinct source IPs,
targeting 35,000 distinct usernames, with a 0.6% success rate. Contrast that shape with normal
traffic, where the success rate on real login attempts is typically well above 90%, and failures
cluster on a small number of accounts (someone mistyping their own password) rather than
spreading almost one-to-one across usernames. The "many usernames, each tried once or twice,
overall success rate far below normal" shape is the credential-stuffing signature; a system
tracking only "failed attempts per account" would see nothing unusual, since almost no single
account crosses a per-account threshold. A velocity-plus-fingerprint detector, by contrast, would
flag the aggregate pattern across usernames and IPs, trigger progressive CAPTCHA challenges on
the highest-risk slice of that traffic, and, once confirmed, feed the responsible IP ranges into
a longer-lived reputation block while triggering forced password resets for the small number of
accounts where a stuffing attempt actually succeeded.
Trade-offs and pitfalls
- Per-account thresholds alone miss credential stuffing by design, since the attack's whole
point is spreading load thin across many accounts specifically to stay under any single
account's threshold; detection has to look at the aggregate pattern across accounts, not just
within one. - CAPTCHA fatigue is a real cost. Applying CAPTCHA broadly, including to every API call
rather than only to suspicious UI flows, degrades the experience for legitimate users and
breaks legitimate automated integrations; reserve it for the specific flows and risk tiers
where it is proportionate. - IP reputation is a leaky signal on its own. Shared NATs, corporate networks, and mobile
carrier-grade NAT mean many legitimate users can share one IP; combine IP reputation with
device and behavioral signals rather than blocking on IP alone. - Common wrong turn: treating detection as the finish line. A confirmed stuffing success
(an attacker actually got into an account) still requires the remediation steps: forced
credential rotation for affected accounts, session invalidation, and customer notification
where applicable; a detector that only logs an alert leaves the compromised accounts open.
Unlock Full Question Bank
Get access to all 7 API Security, Authentication and Authorization interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.