Anomaly and Fraud Detection Questions
Detecting rare, abnormal, or adversarial events in data. Covers anomaly-detection techniques, fraud and risk modeling, handling extreme class imbalance, and the precision/recall and latency tradeoffs of real-time detection systems. Focuses on the modeling patterns unique to needle-in-a-haystack detection problems.
You must serve fraud decisions at 100,000 transactions per second with a hard latency budget under 50 milliseconds per decision. What does that constraint rule out, what feature-serving and model-serving choices does it push you toward, and what do you give up in exchange for that speed?
Sample Answer
Direct answer
A 50-millisecond, 100,000-transactions-per-second budget rules out any synchronous external service call, any model that needs more than a few milliseconds of inference time, and any feature that requires scanning historical data at request time rather than reading a precomputed value; it pushes you toward precomputed and cached features served from an in-memory or low-latency key-value store, a compact model (a small tree ensemble or a distilled model rather than a large deep network), and asynchronous, best-effort enrichment for anything that can tolerate arriving after the decision. In exchange for that speed, you give up some detection depth: signals that require heavier computation or cross-referencing multiple sources typically move to a slower, secondary review pass rather than gating the real-time decision itself.
Structured elaboration
- What the constraint rules out: any feature computation that requires a live database scan or join at request time is too slow at this rate; any model whose inference time is a meaningful fraction of the latency budget (a large neural network, an ensemble of many heavy models) is also ruled out for the primary decision path; and any external network call in the hot path (a third-party risk-scoring API, for example) is almost never fast or reliable enough to include synchronously at this scale.
- What it pushes you toward: a feature-serving layer built on precomputed, cached values (an in-memory key-value store like Redis, with feature computation happening asynchronously in the background, well before the transaction needing them arrives); a compact, fast-inference model, often a moderately-sized gradient-boosted tree ensemble or a distilled/quantized version of a larger model; and horizontal scaling of the serving layer itself, since 100,000 transactions per second at even a few milliseconds each requires many parallel serving instances, not a single fast one.
- What you give up: the richest, most computationally expensive signals (deep graph traversals, large-context sequence models, cross-referencing external data sources) generally cannot run in this hot path at all; they move to an asynchronous secondary pass that can flag a transaction for a follow-up action after the fact, rather than gating the initial decision.
Trade-offs and pitfalls
It's tempting to try to squeeze a more sophisticated model into the budget through aggressive optimization (quantization, pruning, distillation) rather than accepting the architectural trade-off directly; that's often worth doing at the margin, but it doesn't eliminate the fundamental ceiling, a genuinely complex model with many sequential steps or external dependencies rarely fits a true sub-50-millisecond, hot-path budget no matter how it's optimized. It's usually more productive to accept the primary decision will be made by a compact model and invest the saved complexity budget into the asynchronous secondary layer instead.
A fraud model's performance has been quietly degrading over several months as fraud patterns evolve, not because of a data-pipeline break but because fraudsters are behaving differently than they did when the model was trained. How would you tell the difference between this kind of adversarial pattern drift and an ordinary data-quality problem, and what would you do about it?
Sample Answer
Direct answer
Tell the difference by looking at WHERE the degradation concentrates: adversarial pattern drift (fraudsters behaving differently on purpose) tends to show up as a widening gap specifically on the fraud class, with legitimate-transaction scoring staying stable, while an ordinary data-quality problem tends to degrade performance more broadly and often coincides with an identifiable upstream change (a schema change, a broken feature pipeline, a missing data source) rather than a gradual, fraud-specific shift.
Structured elaboration
- Signature of adversarial drift: recall on confirmed fraud cases declines gradually over weeks or months while precision and overall data quality metrics look normal; the newly-missed fraud, once confirmed, often shares a specific new pattern (a slightly different transaction structure, a newly-favored evasion technique) rather than being randomly distributed across all fraud types. This looks like the model's learned decision boundary is being specifically and gradually outmaneuvered.
- Signature of an ordinary data-quality problem: performance degradation tends to be broader (affecting both fraud and legitimate scoring, or showing up as missing/null feature values rather than a shift in the values themselves), and usually coincides with a specific, identifiable event, a schema change in an upstream table, a broken ETL job, a new data source integration, rather than a smooth trend over time.
- Diagnostic approach: check feature-level data-quality metrics (null rates, value-range sanity checks) FIRST, since ruling out an ordinary pipeline problem is cheaper and faster than a deep adversarial-pattern investigation; if data quality checks pass clean, look specifically at the confirmed-fraud subset's feature distributions over time (not the whole population) for a fraud-specific drift signature, which is the more diagnostic signal for genuine adversarial adaptation.
- What to do about genuine adversarial drift: once confirmed, retrain on the most recent confirmed fraud examples specifically (rather than diluting them across a much larger historical dataset), consider whether a new feature is needed to capture the newly-observed evasion pattern directly, and treat this as an ongoing, recurring need rather than a one-time fix, since an adaptive adversary will likely shift again.
Trade-offs and pitfalls
It's tempting to jump straight to "the fraudsters have adapted" whenever performance degrades, since it's a more interesting explanation, but a mundane data-quality problem is often the actual, cheaper-to-fix cause; always rule out the boring explanation first. Conversely, don't dismiss a real adversarial shift as "normal model decay" and simply schedule a routine retrain without first understanding what specifically changed, since a routine retrain on stale features won't fix a gap the model's feature set fundamentally can't see.
Signups spike suddenly and you suspect a wave of bot or spam accounts rather than organic growth. Write SQL to flag likely-fake accounts using heuristics such as shared device or IP across many accounts, implausible signup velocity, or reused payment instruments, and describe how you would test the rule safely before it starts blocking real users.
Sample Answer
Direct answer
Flag likely-fake accounts with SQL heuristics that look for shared identifiers across many "different" accounts, shared device or IP address, reused payment instruments, or implausibly fast signup timing, and validate the rule in shadow mode (logging what it WOULD flag without actually blocking anyone) before letting it affect real signups, so you can measure its false-positive rate against real traffic before it can cause customer harm.
Structured elaboration and worked example (executed)
-- Heuristic 1: same device/IP reused across many accounts
SELECT device_id, ip_address, COUNT(*) AS account_count
FROM signups
GROUP BY device_id, ip_address
HAVING COUNT(*) >= 3;
-- Heuristic 2: implausible signup velocity (multiple signups from the same device within 30 seconds)
SELECT s1.user_id AS user_a, s2.user_id AS user_b, s1.device_id,
(julianday(s2.signed_up_at) - julianday(s1.signed_up_at)) * 86400 AS seconds_apart
FROM signups s1
JOIN signups s2
ON s1.device_id = s2.device_id AND s1.user_id < s2.user_id
WHERE (julianday(s2.signed_up_at) - julianday(s1.signed_up_at)) * 86400 <= 30;
-- Heuristic 3: the same payment instrument reused across different accounts
SELECT payment_instrument, COUNT(DISTINCT user_id) AS distinct_users
FROM signups
GROUP BY payment_instrument
HAVING COUNT(DISTINCT user_id) > 1;
Executed against a small synthetic set of 6 signups, one legitimate cluster of 4 accounts sharing the same device and IP, one legitimate-looking single account, and one account reusing a payment instrument from the first cluster:
device/IP reuse flag: ('devA', '1.2.3.4', 4)
implausible signup velocity flag: all 6 pairs among the 4 devA accounts, 4 to 15 seconds apart
payment-instrument reuse flag: ('card_111', 2)
The device/IP heuristic correctly isolated the 4-account cluster and no one else; the velocity heuristic correctly flagged every pair within that same cluster (all well under the 30-second bar); and the payment-instrument heuristic separately caught the account outside that cluster that reused a card from within it, a distinct signal the first two heuristics alone would have missed.
Testing safely before it blocks real users
Deploy this as a SHADOW rule first, logging what it would flag against live signup traffic without taking any blocking action, and manually review a sample of what it flags for a representative period (at least long enough to see normal weekly traffic patterns, including any legitimate bursts like a marketing-driven signup wave). Only promote it to an active block, or route matched signups to a lightweight additional verification step rather than an outright rejection, once the shadow period confirms an acceptably low false-positive rate against genuinely legitimate signups sharing a device (a family, a shared work computer, a public kiosk).
Trade-offs and pitfalls
Every one of these three heuristics has an honest false-positive story: shared devices/IPs can be a household or an office; fast signup timing can be a legitimate viral referral link driving several friends to sign up in quick succession; and shared payment instruments can be a shared family card. This is exactly why shadow-mode validation before enforcement matters, and why a first response to a flagged cluster should usually be a lighter-touch action (additional verification) rather than an immediate account block.
A fraud model reports 99.5 percent accuracy, but the fraud operations team is unhappy with it. Explain why accuracy is a poor headline metric here, which evaluation metrics you would report instead, and how your choice would change if the fraud rate dropped from 1 percent to 0.05 percent.
Sample Answer
Direct answer
Accuracy is misleading here because the fraud class is a tiny fraction of all transactions, so a model that predicts "not fraud" for everything can score extremely high on accuracy while catching zero fraud. Report precision, recall, and precision-recall area under the curve (PR-AUC) instead, and expect all three to look worse as the fraud rate drops further, purely because the detection problem gets statistically harder, not because the model got worse.
Structured elaboration
- Why accuracy fails: if a base rate is 0.5%, always predicting "legitimate" already achieves 99.5% accuracy while catching 0% of fraud. Accuracy rewards the model for getting the overwhelming majority class right and says nothing about performance on the class that actually matters.
- What to report instead: precision (of the cases you flagged, what fraction were really fraud) and recall (of all the real fraud, what fraction did you catch) directly answer the two questions the fraud team actually cares about. PR-AUC summarizes the precision/recall trade-off across every possible threshold, which matters because the "right" threshold is a business decision, not a fixed property of the model.
- Why receiver operating characteristic area under the curve (ROC-AUC) is a weaker choice here: ROC-AUC's false-positive-rate axis is computed against the (very large) legitimate-transaction population, so it can look deceptively good even when precision at any usable operating point is poor. PR-AUC's precision axis is directly sensitive to the class imbalance, which is exactly the property you want a metric to be sensitive to on this problem.
- What changes as the base rate drops: for a FIXED classifier, a lower base rate mechanically lowers precision at any given recall level, because the same absolute number of false positives is now being compared against fewer true positives. A 0.05% fraud rate needs a noticeably more discriminative model (or a stricter operating threshold) than a 1% fraud rate just to hold precision constant.
Worked example
At a 0.5% fraud rate, predicting "not fraud" for every transaction is accurate 99.5% of the time and catches exactly 0% of fraud (recall 0, precision undefined). That single fact, stated plainly, is usually enough to convince a stakeholder that accuracy is the wrong headline number for this problem; it is worth leading a conversation with it before moving on to precision, recall, and PR-AUC.
Trade-offs and pitfalls
Reporting PR-AUC alone is not sufficient either. A single summary number hides where on the curve you actually plan to operate, and a fraud team cares specifically about precision and recall AT the threshold they will actually use, which is a business decision tied to review capacity (see the separate question on choosing an operating threshold under a review-queue capacity constraint). Always pair a summary metric like PR-AUC with the concrete precision/recall pair at the threshold you intend to ship.
Tell me about a time a fraud or anomaly-detection model you owned started causing a noticeable increase in false positives after it was already in production. Using the STAR format, describe how you noticed it, how you diagnosed the cause, and what you changed.
Sample Answer
Direct answer
Situation: I owned a fraud model that had been stable in production for months, and within a couple of weeks the false-positive rate visibly rose, based on a spike in support tickets from customers whose legitimate transactions were being blocked.
Structured elaboration
Task: my responsibility was to figure out whether this was a model problem, a data problem, or a genuine shift in legitimate customer behavior, and to fix it without simply loosening the threshold and quietly letting more fraud through.
Action: I first checked whether the SCORE DISTRIBUTION had shifted overall, or whether the false positives were concentrated in one segment; they were concentrated in one segment, tied to a recent legitimate change in customer behavior (a promotional campaign had temporarily driven a spike in transaction volume from a subset of customers who looked, to the model, like a burst-pattern anomaly). I confirmed this by pulling a sample of the newly-flagged cases and manually reviewing them against the promotional campaign's start date, rather than assuming any single cause upfront.
Result: rather than lowering the threshold globally, which would have also let more real fraud through, I added the campaign-participation signal as a feature so the model could distinguish "burst from a known promotion" from "burst that looks like fraud," which brought the false-positive rate for that segment back down within about a week of the fix shipping, without loosening detection elsewhere.
Trade-offs and pitfalls
The tempting shortcut in a situation like this is to just move the global threshold until complaints stop; that treats the symptom and quietly increases the recall you're giving up, often without anyone noticing until a later fraud-loss review. Diagnosing the specific, segmented cause first, even though it takes longer, avoids trading real detection capability for a quick fix to customer friction.
Unlock Full Question Bank
Get access to all 41 Anomaly and Fraud Detection interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.