API Security, Authentication and Authorization Questions
Controlling who can call an API, what they may do, and defending it against abuse. Covers the access-control mechanics: API keys, OAuth 2.0 flows, OpenID Connect, JWT issuance/validation, session vs. token auth, scopes/roles for fine-grained authorization, token lifetime and refresh, mutual TLS, and machine-to-machine vs. user-delegated access. Also covers the adversarial hardening view: input validation, injection and deserialization risks, broken object-level authorization (BOLA), mass assignment, secrets handling, and the OWASP API Security Top 10, plus securing data in transit, preventing enumeration/scraping, and testing APIs for vulnerabilities.
You're building a secure webhook receiver for third-party partners. Requirements: authenticate payloads, prevent replay attacks, support retries while ensuring idempotency, scale to high volume, and allow secret rotation. Describe signing schemes (HMAC vs asymmetric), replay defenses (timestamps, unique IDs), idempotency handling and operational patterns for secret rotation.
Sample Answer
Direct answer
Sign every webhook payload so the receiver can prove it really came from the partner (a shared-secret HMAC, or an asymmetric signature if there is no shared secret), bind a timestamp and a unique event ID into that signature so a captured request cannot be replayed later or against a different event, and record processed event IDs so retried deliveries are absorbed as no-ops instead of double-processed. Handle secret rotation by accepting two valid secrets during a short overlap window instead of a hard cutover.
Structured elaboration
Signing schemes: HMAC vs asymmetric
- HMAC (Hash-based Message Authentication Code) is the common default: both sides share one secret, the sender computes a keyed hash (typically HMAC-SHA256) over the request, and the receiver recomputes the same hash and compares it. Fast, simple, and what most webhook providers (Stripe, GitHub) use.
- Asymmetric signing (the sender signs with a private key, the receiver verifies with the sender's public key) removes the need to ever share a secret at all, which matters when you cannot trust a channel to distribute a shared secret safely, or when many receivers need to verify the same sender's signature without each holding a copy of a secret that could leak. It costs more CPU per verification and adds public key distribution/rotation as its own problem.
Replay defenses
- A timestamp header, signed as part of the payload (not sent alongside unsigned), so an attacker who captures a valid request cannot replay it after altering the timestamp without invalidating the signature.
- The receiver rejects any request whose timestamp is outside a short tolerance window (a few minutes), which bounds how long a captured request stays useful even if replayed unmodified.
- A unique event ID per webhook delivery, checked against a short-lived dedup store (a cache keyed by event ID with a TTL slightly longer than the timestamp tolerance), so an identical request replayed within the tolerance window is still caught.
Idempotency for retries
- Partners retry on timeout or a 5xx response, which means the receiver will see the same event ID more than once by design, not by attack. The fix is the same dedup store used for replay defense: before doing any side effect, check whether this event ID has already been processed; if so, return success without repeating the work.
- The "already processed" check and the "mark as processed" write need to be atomic (a unique constraint in the datastore, or an atomic check-and-set in the cache) so two near-simultaneous deliveries of the same event cannot both pass the check before either marks it done.
- Respond fast: verify signature, check the dedup store, enqueue the actual business work, and return 200 immediately. Doing the real processing synchronously inside the request is what makes retries expensive and dedup races more likely under high volume.
Secret rotation
- Never a hard cutover. Accept signatures computed with either the current secret or the previous secret for a defined overlap window, and include a key identifier in the request headers so the receiver knows which secret to try first instead of testing both on every request.
- Communicate the rotation window to partners, retire the old secret only after the window closes, and treat "a partner is still signing with the retired secret past the deadline" as an operational alert, not a silent failure.
Worked example
A concrete HMAC-SHA256 signature, computed and verified exactly as shown (copy, run, and you will get the same digest since every input is pinned):
import hmac, hashlib
secret = b"whsec_5f8a3c9e2b7d4a1f6e0c8b2d9a4f7c1e"
timestamp = "1735689600"
body = b'{"event":"payment.succeeded","id":"evt_8f2a","amount":4999}'
signed_payload = timestamp.encode() + b"." + body
signature = hmac.new(secret, signed_payload, hashlib.sha256).hexdigest()
print(signature)
Output:
2e702838565a081c906544b3c8d22f6ce61ab8cedeb9babb9f7ae1245e596b14
On the receiving side, verification recomputes the same digest and compares in constant time, then applies the replay and idempotency checks:
def verify_and_should_process(headers, body, secret, seen_event_ids, now, tolerance_s=300):
timestamp = headers["X-Timestamp"]
if abs(now - int(timestamp)) > tolerance_s:
return False, "timestamp outside tolerance"
expected = hmac.new(secret, (timestamp + ".").encode() + body, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, headers["X-Signature"]):
return False, "bad signature"
event_id = headers["X-Event-Id"]
if event_id in seen_event_ids:
return True, "already processed, ack without reprocessing"
seen_event_ids.add(event_id)
return True, "process"
This is not run in this answer since it depends on request-time state (now, a shared dedup store); the HMAC computation above is the part with fully pinned inputs and a verifiable output.
Trade-offs & pitfalls
Using a plain string comparison (==) instead of a constant-time comparison (hmac.compare_digest) leaks timing information an attacker can use to guess the signature byte by byte. Trusting a client-supplied timestamp that is not itself signed into the HMAC input lets an attacker pair an old, still-valid signature with a fresh timestamp. If the idempotency store is local to one server instance instead of shared, two receiver replicas behind a load balancer can each independently "not have seen" the same event and both process it. Rotating a secret without an overlap window causes a hard outage for every in-flight or slightly-delayed delivery signed with the old secret.
How would you handle authentication token lifecycle and rotation for long-lived API clients such as IoT devices? Include token issuance, refresh, rotation, revocation, offline device handling, heartbeat strategies, secure storage on device, and methods for detecting and responding to token compromise.
Sample Answer
Direct answer
Treat an Internet of Things (IoT) device like a machine-to-machine client with a materially worse threat model than a server: it can be physically accessed, it cannot always phone home, and it typically has limited secure-storage hardware. The design centers on a strong, device-bound identity established once at provisioning, short-lived operational tokens refreshed opportunistically rather than on a fixed clock, and a revocation path that still works for a device that happens to be offline right now.
Structured elaboration
Issuance
At manufacturing or provisioning time, each device gets a unique, device-bound credential, ideally a private key generated on a hardware security element or secure enclave on the device itself, so the private key material never exists outside that chip, plus a certificate or registration record tying that key to a specific device ID. This is the root of trust everything else is refreshed from; it is not the day-to-day operational token.
Refresh
The device uses its long-lived provisioning credential to periodically obtain short-lived operational access tokens, brief, signed JSON Web Tokens (JWTs) that a receiving service can verify locally against a cached signing key rather than a long-lived static credential, refreshed well before expiry during normal connectivity so a brief network blip does not strand the device without a valid token.
Rotation
Operational tokens rotate frequently, short time-to-live values. The underlying device credential itself should also be rotatable on a much slower cadence, a scheduled re-provisioning or over-the-air credential rotation, so that even the root device identity is not a forever-secret, though this is the hardest piece operationally for a fleet with unreliable connectivity.
Revocation
Maintain a device-identity-level revocation list keyed on device ID, not on individual tokens, since a compromised device should lose all access, not just one token, checked at the point a device requests a new operational token. Because operational tokens are already short-lived, revocation does not need to reach an already-issued token instantly; it only has to stop the next refresh from succeeding, which bounds the compromise window to that token's remaining lifetime.
Offline device handling
A device disconnected for an extended period will have an expired operational token by the time it reconnects. Design the re-provisioning flow to work from the device's still-valid root credential, not requiring a human to physically re-touch the device, while still checking that root credential against the revocation list on every reconnect, so an offline period never becomes a loophole that skips the revocation check.
Heartbeat strategies
Devices send periodic, low-cost heartbeat or check-in calls, distinct from full operational traffic, that double as both a liveness signal for fleet management and a natural, low-frequency point to check the device's credential against the revocation list and pull a fresh operational token, without requiring constant high-frequency calls that would drain a battery-powered device.
Secure storage on device
The root credential's private key should live in hardware-backed secure storage, a Trusted Platform Module (TPM), secure enclave, or dedicated secure element chip, that resists extraction even with physical access, since an IoT device, unlike a typical server, can plausibly end up physically in an attacker's hands. If the hardware cannot support that, at minimum encrypt credentials at rest with a key derived from device-specific hardware characteristics, understanding that is a materially weaker guarantee than true hardware-backed storage.
Detecting and responding to compromise
Baseline expected behavior per device or device class, call frequency, data volume, which endpoints it normally calls, and alert on deviation. A sensor that normally reports once an hour suddenly polling every second, or calling an endpoint it has never called before, is a strong compromise signal. On a confirmed or suspected compromise, revoke at the device-identity level immediately, and for a fleet-wide vulnerability, a shared firmware bug rather than one stolen device, be prepared to force re-provisioning of the entire affected device class.
Worked example
A fleet of smart thermostats, each provisioned with a device-bound key at the factory, sends a heartbeat every fifteen minutes and refreshes its operational token alongside it. A firmware vulnerability is discovered that allows key extraction from a physically accessed unit. The response is to revoke at the affected batch or device-ID range level, force those specific devices through re-provisioning on their next heartbeat, and treat the vulnerable firmware version itself as a compromise indicator to actively hunt for across fleet telemetry, since any device still running it is a live exposure regardless of whether it has been physically tampered with yet.
Trade-offs & pitfalls
The most common and costly mistake is baking one long-lived, static API key into every device's firmware image: cheap to build, catastrophic to rotate, since it lives inside every unit ever shipped, and revoking it revokes the entire fleet at once. Heartbeat frequency is a genuine trade-off against a battery or bandwidth-constrained device's power budget, so the revocation-check cadence is coupled to a real cost, not a free design choice. Hardware-backed secure storage adds real bill-of-materials cost per unit, so cheaper device tiers sometimes skip it and inherit a materially worse compromise story; that trade-off should be made explicitly by the product and security teams together, not defaulted into silently.
Lay out the API contract and operational controls for partner integrations that will process PII and PCI data. Cover authentication (mutual TLS, OAuth), token lifecycle management, field-level encryption, throttling and quota models, audit trails and logging, data retention and deletion policies, and how you'd demonstrate compliance during pre-sales and onboarding.
Sample Answer
Treat authentication, encryption, and auditability as three independent layers that each have to hold up on their own: mutual TLS plus OAuth 2.0 for authentication, envelope field-level encryption so sensitive fields stay encrypted even from your own logs, and an immutable audit trail. Then package all three into a documented onboarding artifact so compliance is something you can demonstrate to a partner's security team before the contract is signed, not something you merely assert.
Authentication layer
Mandatory mutual TLS (mTLS) at the transport layer: the partner presents a client certificate mapped to a partner id, in addition to any application token, so identity survives even if a token leaks. OAuth 2.0 (Client Credentials for service-to-service calls, Authorization Code for any human-in-the-loop flow) issues short-lived JSON Web Token (JWT) access tokens, TTL under 5 minutes, with narrow scopes such as payments:write or pii:read:limited.
Token lifecycle
Access tokens are short-lived and verified locally. Refresh tokens are single-use, rotated, encrypted at rest, and revocable via an introspection endpoint (a network call to the issuer asking whether a token is still valid) or admin endpoint. Certificates rotate on a fixed cadence, for example quarterly; any remaining client secrets rotate at least as often.
Field-level encryption
Envelope encryption: the partner encrypts sensitive fields client-side with a one-time Data Encryption Key (DEK), and the DEK itself is encrypted with your published Key Encryption Key (KEK), held in a Key Management Service (KMS) or Hardware Security Module (HSM). Only a hardened decryption service inside your trust boundary can unwrap the DEK and read the field, the gateway, load balancers, and general application logs never see plaintext PII (Personally Identifiable Information) or PCI (Payment Card Industry) data. Tokenize card numbers, storing a PCI-DSS-compliant token in place of the real number, rather than letting raw card data flow through more of the system than absolutely necessary.
Throttling and quota model
Multi-tier limits, per-second burst, per-minute sustained, and daily cap, all keyed by partner id, not IP. A soft limit returns 429 with a Retry-After header; a hard limit blocks; anomaly-triggered dynamic throttling reduces a partner's limit automatically if abuse is suspected.
Audit trails and logging
Append-only (write-once, read-many, or WORM) storage for every access event: timestamp, partner id, endpoint, allow/deny decision, certificate fingerprint, token id, and a hash of the payload, never the raw sensitive fields, so the trail is forensically useful without itself becoming a second copy of the sensitive data. Retain per the stricter of contractual or regulatory requirement, PCI DSS often drives a roughly one-year retention floor for certain logs, and make retention length a documented, per-data-class policy rather than a single global number.
Retention and deletion
Retention is classification-driven: PCI data, PII, and metadata each get their own clock. Deletion is cryptographic as well as row-level: mark the record deleted, then destroy the DEK for that record, so even a backup copy of the encrypted blob becomes unrecoverable. This is what lets you honor a deletion request against systems (backups, replicas) you can't reach directly to delete individual rows from.
Demonstrating compliance during pre-sales and onboarding
A documented onboarding package: an API spec with per-field sensitivity classification, an architecture summary (mTLS plus OAuth plus envelope encryption), current compliance artifacts (a PCI Attestation of Compliance, an assessor's certification that you meet card-data-security standards; a recent penetration-test summary, evidence that an independent tester actively tried to break in and documented the results; a SOC 2 Type II report if available, an independent auditor's confirmation that your security controls actually operated correctly over months, not just on paper), and a sandbox environment with synthetic data so the partner's security team can test the flow before any real data moves. A signed Data Processing Addendum (DPA), and a Business Associate Agreement (BAA) if health data is ever in scope, belong in the onboarding checklist, not something discovered mid-integration.
Worked example
A partner submits a payment record containing a card number and a customer's mailing address. The card number is tokenized before it ever reaches your API, using your published tokenization SDK. The mailing address, a PII field, is envelope-encrypted client-side with a per-request DEK, wrapped by your KEK. At the gateway, the mTLS handshake confirms partner identity, the OAuth JWT confirms scope payments:write, and the request is logged with a hash of the still-encrypted address field, never the plaintext. Only the downstream payment-processing service, running inside a separate, narrowly-scoped trust boundary, holds the KMS permission to unwrap that specific DEK and read the address. If a customer later exercises a deletion right, the deletion service destroys that record's DEK; the encrypted blob remains in backups but is now permanently unreadable, satisfying the deletion obligation without reaching into every backup copy individually.
Trade-offs and pitfalls
Field-level envelope encryption is strong but adds real integration complexity for the partner; provide a client SDK so partners aren't implementing DEK/KEK handling by hand, or implementations will be inconsistent and error-prone across partners. Short token TTLs and frequent certificate rotation improve security but raise operational overhead, automate rotation end to end or it will lapse under deadline pressure. Destroying the DEK to satisfy a deletion request is an elegant way to handle backups you can't directly edit, but it means losing that data permanently the moment the key is destroyed, get sign-off on retention windows before deletion becomes irreversible.
Build an automated test to detect timing side-channel leaks in an authentication API where the response time might differ based on username validity. Explain how you would reduce noise, choose sample sizes, apply statistical analysis, and automate detection in CI without producing many false positives due to network variance.
Sample Answer
Direct answer
A timing side-channel test compares the distribution of response times for two request
classes (valid username vs. invalid username), not a single measurement, because one request
is dominated by network jitter. Build it as: collect many interleaved paired samples, trim
outliers, run a distribution-comparison statistic that does not assume clean symmetric noise,
and require the result to repeat across several independent CI runs before failing a build, so
one noisy run cannot trip a false positive.
Structured elaboration
Why the leak exists. An endpoint that does extra work only when a username is real (looks
up a stored password hash, runs a bcrypt/argon2 compare) systematically takes longer for valid
usernames, even when the HTTP response body and status code look identical. That gap lets an
attacker enumerate real usernames purely from timing, with no error message needed.
Reduce noise before comparing anything:
- Warm up the service and discard the first N requests (cold caches, JIT/interpreter warmup,
TCP slow start all inflate early samples). - Interleave the two request types (alternate valid, invalid, valid, invalid...) instead of
running one whole batch then the other. Load drift, garbage collection pauses, and neighbor
noisy-tenant effects then hit both groups equally instead of biasing one. - Trim outliers (for example the top and bottom 5%) before computing a summary statistic.
Network jitter and retransmits produce occasional huge spikes that would otherwise swamp a
genuine sub-millisecond gap.
Choosing a sample size. The bigger the real timing gap relative to the noise, the fewer
samples you need to detect it reliably. A gap of a few milliseconds against light noise might
separate cleanly with a few dozen samples; a sub-millisecond gap buried in several milliseconds
of network jitter (the realistic case for a well-behaved API) typically needs several hundred
to a few thousand samples per group. Start around 300 to 500 per group, and increase it if
independent runs disagree with each other.
Statistical analysis. Do not compare raw means with a standard two-sample t-test (a statistic that compares two groups' averages, assuming the noise around each is roughly symmetric and bell-curve-shaped): network
latency is right-skewed (a long tail of slow outliers, not a symmetric bell curve), which
violates the t-test's assumptions and makes it unreliable here. A permutation test on the
trimmed data is a better fit: it makes no assumption about the shape of the noise, works
directly on the trimmed samples, and gives an empirical p-value (the probability of seeing a gap at least this large purely by chance, if there were no real timing difference) by asking "how often would a
gap this large appear if the labels were random?"
Automate in CI without false positives from network variance:
- Fix an explicit significance threshold (a low one, since you will run the test often) and
require several independent runs to agree before failing the build. A single p < 0.05 result
will happen by chance on noisy timing data with some regularity. - Run on a dedicated or otherwise quiet CI runner where possible, so background load is not
itself a confound. - Include a negative control in the same pipeline: the same statistical test comparing
valid-vs-valid samples should almost never fire. If it does, your pipeline's noise floor is
too high to trust the real test's result that day.
Worked example
The script below simulates response times for 400 requests per group: a shared 12ms base cost,
a real 0.6ms leak injected only into the "invalid username" branch (modeling an extra hashed
comparison), and network jitter modeled as a right-skewed exponential distribution (mean 3ms)
rather than a symmetric one, since that is closer to real network noise. It trims the top and
bottom 5% of each group, then runs a one-sided permutation test. It checks 5 independent
"CI runs" (5 different random seeds) and only flags a leak if at least 4 of 5 agree at
p < 0.01, then repeats the whole pipeline with no injected leak as a negative control.
import random
import statistics
random.seed(42)
N = 400 # samples per group
def sample_response_times(base_ms, leak_ms, n):
# base_ms: shared processing cost. leak_ms: the side-channel gap this
# branch adds (0 for the "valid username" baseline). Jitter is modeled
# as a right-skewed exponential, closer to real network noise than Gaussian.
out = []
for _ in range(n):
jitter = random.expovariate(1 / 3.0) # mean 3ms right-skewed jitter
core = random.gauss(base_ms, 0.4) # small stable variance in app code
out.append(core + leak_ms + jitter)
return out
def trim(samples, pct=0.05):
# drop the top/bottom pct of samples to blunt network-spike outliers
s = sorted(samples)
k = int(len(s) * pct)
return s[k: len(s) - k] if k > 0 else s
def permutation_test(a, b, n_perm=2000, seed=0):
# one-sided permutation test for mean(b) > mean(a)
rng = random.Random(seed)
observed = statistics.mean(b) - statistics.mean(a)
pooled = a + b
na = len(a)
count_ge = 0
for _ in range(n_perm):
rng.shuffle(pooled)
perm_a = pooled[:na]
perm_b = pooled[na:]
diff = statistics.mean(perm_b) - statistics.mean(perm_a)
if diff >= observed:
count_ge += 1
p_value = (count_ge + 1) / (n_perm + 1) # avoids an impossible p=0
return observed, p_value
def run_one_ci_check(leak_ms, run_seed):
random.seed(run_seed)
valid = sample_response_times(base_ms=12.0, leak_ms=0.0, n=N)
invalid = sample_response_times(base_ms=12.0, leak_ms=leak_ms, n=N)
observed_ms, p = permutation_test(trim(valid), trim(invalid), n_perm=2000, seed=run_seed)
return observed_ms, p
# a real 0.6ms leak (e.g. an extra bcrypt compare only run when the username
# exists) buried under ~3ms mean network jitter, checked across 5 independent
# CI runs (5 seeds) instead of trusting a single run
alpha = 0.01
significant_runs = 0
print("run observed_diff_ms p_value significant(p<0.01)")
for i, seed in enumerate([1, 2, 3, 4, 5]):
observed_ms, p = run_one_ci_check(leak_ms=0.6, run_seed=seed)
sig = p < alpha
significant_runs += int(sig)
print(f"{i+1:>3} {observed_ms:>16.3f} {p:>7.4f} {sig}")
print(f"\nsignificant in {significant_runs}/5 runs (require >=4/5 to flag in CI)")
print("FLAG TIMING LEAK" if significant_runs >= 4 else "no consistent leak detected")
# negative control: no injected leak, same pipeline, to show the gate does
# not fire on jitter alone
print("\nnegative control (leak_ms=0.0):")
significant_runs_ctrl = 0
for i, seed in enumerate([11, 12, 13, 14, 15]):
observed_ms, p = run_one_ci_check(leak_ms=0.0, run_seed=seed)
sig = p < alpha
significant_runs_ctrl += int(sig)
print(f"{i+1:>3} {observed_ms:>16.3f} {p:>7.4f} {sig}")
print(f"significant in {significant_runs_ctrl}/5 runs (require >=4/5 to flag in CI)")
print("FLAG TIMING LEAK" if significant_runs_ctrl >= 4 else "no consistent leak detected (correct: no leak was injected)")
Output:
run observed_diff_ms p_value significant(p<0.01)
1 0.433 0.0045 True
2 0.696 0.0005 True
3 0.627 0.0005 True
4 0.771 0.0005 True
5 0.318 0.0170 False
significant in 4/5 runs (require >=4/5 to flag in CI)
FLAG TIMING LEAK
negative control (leak_ms=0.0):
1 -0.014 0.5402 False
2 -0.215 0.8966 False
3 -0.068 0.6727 False
4 -0.070 0.6737 False
5 -0.051 0.6207 False
significant in 0/5 runs (require >=4/5 to flag in CI)
no consistent leak detected (correct: no leak was injected)
With a real 0.6ms leak, 4 of 5 simulated CI runs cross the p < 0.01 threshold, so the gate
correctly fires. With no injected leak, all 5 runs come back non-significant, so the negative
control confirms the pipeline does not fire on jitter alone. Run 5 in the first block (p =
0.017) illustrates exactly why a single-run threshold is unsafe: that one run alone would have
been called "not significant" at the 0.01 level even though the leak was real, which is why the
gate requires 4 of 5 runs to agree rather than trusting any single run.
Trade-offs and pitfalls
- Statistical detection is not the same as measuring real-world exploitability. This test
tells you a leak likely exists locally; it does not tell you how many requests an attacker
over the public internet, with far more noise than your CI network, would need to exploit it.
Those are two different measurements, and a leak that is statistically detectable in a quiet
CI environment may be much harder to exploit remotely. - Common wrong turn: comparing raw means on unfiltered data with a standard t-test. Network
jitter is heavy-tailed, so a handful of retransmits can swing a raw mean far more than a
genuine sub-millisecond leak, producing both false positives and false negatives. - Common wrong turn: trusting a single p-value. Noisy timing measurements will occasionally
produce a "significant" result by chance; requiring several independent runs to agree, as in
the worked example, controls that risk in a way a single run cannot. - Pitfall: "fixing" the endpoint's happy path but not its neighbors. If you make the
invalid-username branch add a matching artificial delay but a cache layer, a database index,
or connection pooling upstream still behaves differently for existing vs. non-existing users,
the leak often survives at a smaller magnitude. Re-run the same statistical test as a
regression gate after any fix, not just once at discovery time.
Propose detection and mitigation strategies for abusive API usage and credential theft at scale. Cover techniques such as per-key behavioral baselines, anomaly detection, per-key throttling and freezing, ephemeral credential issuance, credential rotation, fingerprinting, and forensics-ready logging while balancing privacy and performance.
Sample Answer
Build a per-key behavioral baseline, normal request rate, endpoints touched, geography, for every API key, score each new burst of activity against its own history rather than a single global threshold, and respond with graduated actions, soft throttle, then freeze, then full revoke, so a false positive costs a legitimate client a slowdown, not a hard outage.
Per-key behavioral baselines
Track, per key, over a rolling window (for example, 7 days): request rate, endpoint mix, typical payload size, source IP or geography, and time-of-day pattern. Use a decaying window, weighting recent activity more, so a key's baseline adapts as legitimate usage genuinely changes, without being so reactive that an attacker can slowly retrain the baseline toward their own behavior, cap how fast the baseline itself is allowed to shift per day.
Anomaly detection
Combine simple rules, an absolute rate ceiling, or a geographically impossible gap between two consecutive requests, with a statistical or machine-learning layer, unsupervised outlier detection (flagging unusual activity by comparing it to the shape of normal data, with no need for pre-labeled examples of "bad") over the same behavioral features, so you catch both obvious abuse and subtler drift. Score risk as a combination of signals rather than one metric alone, as shown in the worked example below.
Per-key throttling and freezing
Use a graduated response: soft throttle (a reduced rate limit) at a moderate anomaly score, freeze (block new requests, let existing sessions time out) at a high score, and full revoke only after freeze plus either automated corroboration or human review, so a single noisy signal doesn't take down a legitimate integration outright.
Ephemeral credential issuance and rotation
For high-privilege operations, issue short-lived, single-purpose credentials, minutes rather than months, rather than relying solely on the long-lived key. Rotate long-lived keys on a fixed cadence automatically, with a grace overlap so rotation itself never causes an outage.
Fingerprinting
Collect signals that aren't tied to a person's identity. A TLS client fingerprint is the common starting point, since it's cheap to capture at the TLS layer with no application changes; an HTTP header ordering fingerprint and a hashed device signal are rarer, higher-effort additions worth reaching for only once TLS fingerprinting alone isn't separating attackers from genuine clients. The goal is to distinguish "the same script hitting us from many rotated IPs" from genuinely independent clients, without needing to know who the human behind it is.
Forensics-ready logging, balancing privacy and performance
Keep an append-only, tamper-evident log of every mitigation decision, the score, the signals that fired, the action taken, so an incident can be reconstructed after the fact. Store raw request bodies only when a risk threshold is crossed, sampled or fully captured once triggered rather than captured for every request by default, so both storage cost and privacy exposure scale with actual risk.
Worked example
Key partner-key-77 has a 7-day baseline of 150 requests/hour, 95% from a stable IP range, touching a consistent mix of 4 endpoints. In one hour it makes 900 requests (6 times baseline), from a new IP range, hitting a 5th endpoint it has never called before. Define a risk score:
score = rate_ratio + new_ip_flag + new_endpoint_flag
rate_ratio = 900 / 150 = 6
new_ip_flag = 2 (a range never seen before)
new_endpoint_flag = 1 (one never-seen endpoint)
score = 6 + 2 + 1 = 9
Thresholds: score >= 4 triggers a soft throttle (halve this key's rate limit); score >= 8 triggers a freeze (block new requests, page security on-call); score >= 15 triggers auto-revoke pending review. At a score of 9, the key is frozen and a human is paged, but not auto-revoked, since a new endpoint plus an IP change could still reflect a legitimate deployment change on the partner's side. The human reviewer's job is to distinguish that from theft within the freeze window before deciding whether to restore or revoke.
Trade-offs and pitfalls
Graduated response reduces false-positive damage but adds an operational cost, someone has to staff the review queue between freeze and revoke, size that team for your expected alert volume before relying on it. Behavioral baselines that adapt too quickly can be gamed by an attacker who ramps up slowly, capping the daily rate of baseline drift bounds this risk. Full-fidelity request logging aids forensics but raises both storage cost and privacy exposure if applied to every request, trigger deeper capture only once risk crosses a threshold, and document the retention window for that captured data separately from normal logs.
Unlock Full Question Bank
Get access to all API Security, Authentication and Authorization interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.