Debugging and Testing ML Systems Questions
Finding, diagnosing, and fixing problems in ML code, data, and models, and building tests that catch these problems before they reach users. Covers common ML pitfalls (data leakage, shape mismatches, silent training bugs, mis-specified loss or metrics), root-cause analysis of model regressions and production incidents (accuracy drops, calibration drift, intermittent or hard-to-reproduce failures), distributed-training-specific failures (multi-GPU divergence, intermittent OOM, precision-related instability), and the diagnostic tooling that supports it (reproducibility artifacts, structured logging, instrumentation). Also covers testing ML systems directly: unit tests for data and feature pipelines, validation checks for datasets and features, test oracles and acceptance criteria for probabilistic or non-deterministic model outputs, and integration and regression tests that catch model or pipeline regressions before deployment. Emphasizes the engineering rigor that keeps ML systems correct and maintainable.
Training loss decreases overall but oscillates violently even with a small learning rate, or validation metrics behave erratically (sometimes improving, sometimes degrading run to run) even though the training loss trend looks fine. Provide a debugging checklist focused on the optimizer and training configuration, and describe at least one small, fast reproducible experiment you would run to distinguish an optimizer bug from a genuine data or label problem, and how the diagnosis changes if the instability appears only in production, not in local dev runs.
Sample Answer
Direct answer. Loss that oscillates even at a small learning rate, or validation metrics that swing run to run despite a stable training-loss trend, points at the optimizer configuration or a data/label problem interacting badly with training dynamics, and the fastest way to tell them apart is a small, fast, targeted experiment rather than more staring at the loss curve.
Optimizer/configuration checklist. Check momentum, weight decay, and gradient-accumulation settings for an accidental mismatch. The specific accumulation bug worth knowing the arithmetic for: if you accumulate 4 micro-batches and SUM their gradients instead of averaging them, the gradient handed to the optimizer is 4 times too large. Under SGD (with or without momentum) the update is linear in the gradient, so that is exactly a 4x learning rate: a configured 3e-4 behaves like 1.2e-3, and the factor is the accumulation-step count, nothing subtler. This is why the symptom can appear the day someone changes accum_steps and nothing else.
Be careful applying that arithmetic to Adam or AdamW, because it does NOT carry over, and 3e-4 is exactly the learning rate people usually run Adam at. Adam divides each gradient by a running estimate of its own magnitude, so multiplying every gradient by a constant cancels in the numerator and denominator together and the parameter update is unchanged. Summed accumulation under Adam therefore looks like NO learning-rate change at all, not a 4x one, which flips the diagnosis: if you suspect this bug and your optimizer is Adam, the absence of a step-size change is not evidence that accumulation is correct. What the bug does still break under Adam is everything that reads the raw gradient magnitude, so look there instead: a clip_grad_norm_ threshold now bites at a quarter of the intended norm, mixed-precision loss scaling sees inflated gradients and backs its scale off, and non-decoupled weight decay (plain Adam(weight_decay=) rather than AdamW) rides inside the gradient and so does get rescaled by 4. The fastest confirmation either way is to log the global gradient norm the optimizer actually receives and check whether it moves by the accumulation-step count when you change accum_steps.
The next two items apply only when training is spread across multiple GPUs, so skip them on a single-device job. (A replica here is one full copy of the model living on one GPU; each replica processes a different slice of the batch and the replicas' gradients are averaged together each step.) For models with BatchNorm, check whether normalization statistics are being computed consistently across replicas, inconsistent per-replica batch statistics is a specific, common source of run-to-run instability that looks like an optimizer bug but isn't. Then, whether single-device or distributed, try a learning-rate range test (a short run sweeping LR from very small to very large while watching where loss starts diverging) to get an empirical, rather than guessed, sense of what LR range is actually stable for this model and data.
Two small, fast experiments, and what each result means.
Test 1, the label-permutation test. Shuffle the labels randomly so no true relationship between features and labels survives, then train the same model with the same configuration. Read TWO numbers from it. First, validation performance: with the signal destroyed, it must land at chance (AUC around 0.5, accuracy around the majority-class rate). If validation comes out meaningfully ABOVE chance on shuffled labels, that is an unambiguous data bug, because there is nothing left to learn, so anything above chance means information is crossing from the training rows into the validation rows: duplicate or near-duplicate records split across both sides, a leaked identifier, or an encoding computed over the whole dataset before splitting. No optimizer change will fix that. Second, training loss: if it still falls to the same low value it reaches with real labels, the model has the capacity to memorize the training set outright, which is not itself a bug but does mean the training-loss curve carries no evidence about whether anything real was learned, so from that point on you reason from validation only. If training loss instead stalls near the entropy of the label distribution, labels are reaching the loss correctly and the optimizer is not fabricating signal.
Test 2, the tiny-sample overfit test. With the real (unshuffled) labels, try to drive training loss to near zero on a handful of examples. It should get there in a few dozen steps if the optimizer and loss function are wired correctly.
The two together. The pair is more informative than either alone:
tiny-overfit FAILS tiny-overfit PASSES
permutation val the training loop itself is pipeline is wired correctly and
at chance broken (loss detached from the the split is clean; the instability
graph, optimizer stepping the lives in the full dataset or the
wrong parameter group, an full-scale config: label noise, a
effective LR of zero). Fix this rare corrupted batch, batch
before reading anything else. composition, LR range.
permutation val two independent bugs. Fix the split contamination or a leaked
above chance loop first, because a permutation feature. A data problem; no
result from a loop that cannot optimizer or LR change will
fit 10 examples is not evidence. resolve it.
If the instability appears only in production, not local dev runs. This shifts the most likely cause away from the optimizer config itself (which is presumably identical) and toward something that differs between the two environments: the ACTUAL data (is a production data slice statistically different from the local dev sample, in a way that interacts badly with the current batch-norm or learning-rate settings), or a subtle environment difference (different hardware, different library version, different degree of numerical determinism) that changes floating-point behavior enough to matter for an already-marginally-stable configuration. The label-permutation and tiny-overfit tests should be re-run specifically on a sample of the PRODUCTION data pulled locally, not just the original dev sample, since the whole point is to test whether the data itself is the differentiator.
A large model-training run fails with a GPU out-of-memory error when you increase the batch size. List the practical mitigation strategies available to you, and describe the trade-off each one makes (extra compute time, implementation complexity, or a change in effective batch-size semantics).
Sample Answer
Direct answer. When a larger batch size pushes training past available GPU memory, the available fixes trade off differently between compute time, implementation complexity, and whether they change the actual math of what "batch size" means, so the right choice depends on which of those you can afford to spend.
The mitigation menu.
- Gradient accumulation: run several smaller micro-batches, summing their gradients before a single optimizer step, simulating a larger effective batch size without ever holding the full batch in memory at once. Trade-off: nearly free in implementation complexity and mathematically equivalent to true large-batch training for most loss functions, but costs extra wall-clock time since you're doing more forward/backward passes for the same effective batch. The exception is the batch-size-semantics one, and it is worth knowing by name: any layer that computes statistics over the batch, batch normalization above all, sees only the micro-batch, not the full effective batch, so a network containing batch normalization is NOT equivalent under accumulation and needs synchronized batch normalization across micro-batches or a normalization layer that does not depend on the batch at all.
- Mixed precision: store activations (the intermediate outputs each layer produces during the forward pass, which are kept in memory because the backward pass needs them, and which usually dominate memory during training) and weights in fp16 or bf16 (bf16 being a 16-bit format with fp32's exponent range but fewer mantissa bits, so it trades precision for far less overflow risk) instead of fp32, roughly halving memory footprint. Trade-off: often a "free" win in both memory and speed on modern hardware, but introduces its own numerical-stability considerations (dynamic loss scaling, meaning the loss is multiplied by a large factor before backward and the gradients divided by it afterwards, so that small gradient values do not vanish to zero in fp16, with the factor adjusted automatically when it causes an overflow; plus watching for NaNs from fp16 underflow and overflow), so it is not entirely free of engineering cost.
- Model or pipeline parallelism: split the model itself (not just the data) across multiple devices, so no single device needs to hold the full model or full activations. Trade-off: the most memory-scalable option but the most complex to implement correctly, and introduces communication overhead between devices.
- Activation checkpointing: don't store intermediate activations from the forward pass, recompute them during the backward pass instead. Trade-off: significantly reduces memory (often the single biggest lever for very deep networks) at the direct cost of extra compute, since the forward pass effectively runs twice for the checkpointed segments.
- Reducing sequence length: for sequence models, truncating or bucketing to a shorter maximum length directly reduces the memory footprint of attention and activations. Trade-off: free in implementation cost, but changes what the model actually sees, potentially losing information for the fraction of examples whose true length exceeds the new cap.
- Reducing the batch size itself, with the learning rate re-tuned to match: the option that is easy to forget precisely because it looks like giving up. Trade-off: free in both compute and implementation complexity, and the one option on this list whose cost lands squarely on effective batch-size semantics. A smaller batch means noisier gradient estimates and a different effective noise scale, so the learning rate and schedule usually need re-tuning (often downward), and the run's convergence path, and sometimes its final quality, genuinely changes rather than merely being reorganized.
- Parameter sharding (splitting optimizer state and/or parameters themselves across devices, as in ZeRO-style approaches, ZeRO being a family of techniques that partitions optimizer state, gradients and eventually parameters across devices instead of replicating them on every one): substantially reduces per-device memory for the optimizer state specifically (which is often larger than the model weights themselves for adaptive optimizers, that is, optimizers such as Adam that keep one or two running statistics per parameter and therefore carry several times the model's own size in state). Trade-off: adds communication overhead to gather sharded parameters when needed, and is more complex to set up than gradient accumulation or mixed precision.
How to choose. The two to know cold are mixed precision and gradient accumulation: both are close to free in engineering cost and, batch-normalization layers aside, neither changes what the model learns. Activation checkpointing is the third to reach for, since it is a pure compute-for-memory trade with no semantic change at all. Only then reach for model or pipeline parallelism, sharding, reducing sequence length, or simply shrinking the batch, since each of those trades away something the first three do not: implementation complexity and communication overhead for parallelism and sharding, information in the input for a shorter sequence, and effective batch-size semantics for a smaller batch with a re-tuned learning rate.
After adding a new feature or component, validation accuracy dropped. Design an ablation study to determine which change caused the regression: what controlled experiments you would run, how you would log and compare results, how you would control for run-to-run variance so you can assess statistical significance, and how you would reason about interactions between features rather than testing each one in complete isolation.
Sample Answer
Direct answer
Treat this as a controlled experiment, not a single before/after diff. Retrain with the new feature or component present and again with it absent, holding the code, data split, and hyperparameters fixed, and repeat each configuration across several random seeds so one lucky or unlucky run cannot masquerade as a real effect. Compare configurations with a paired statistical test on validation accuracy across matched seeds, and test the suspect changes together as well as individually, because two individually harmless changes can interact and produce a regression that neither one causes alone.
Structured elaboration
Controlled experiments to run. Define a small factorial design rather than a single before/after pair. If the change under investigation bundles more than one addition (say, a new feature A and a new preprocessing step B that shipped together), the minimum useful set of conditions is:
| condition | A present | B present |
|---|---|---|
| baseline | no | no |
| +A only | yes | no |
| +B only | no | yes |
| +A+B (the shipped change) | yes | yes |
Every condition uses the identical training code, the identical data split per seed, and the identical hyperparameters; the only thing that varies between conditions is which of the suspect additions is active. A full factorial design costs 2N configurations for N suspect changes, which is fine for 2 to 3 changes but explodes past that. For a larger bundle, run leave-one-out first (start from the full shipped candidate and remove one change at a time) to get a first-order attribution, then escalate to a small factorial only among the handful of changes leave-one-out flags as suspicious, rather than paying for the full combinatorial grid up front.
Logging and comparing results. Every run writes to a shared experiment-tracking table keyed by: exact code commit hash, data-version identifier, the active-changes bitmask for that condition, the random seed, and the resulting validation accuracy (plus secondary metrics such as precision, recall, and calibration, so a regression that only shows up in one slice is not invisible to a single scalar). Comparison happens condition-by-condition as a table of mean and standard deviation across seeds, never as a single run versus a single run.
Controlling run-to-run variance so significance is assessable. The variance that matters here has two sources: data-split variance (which examples land in train versus validation) and initialization/shuffling variance (weight init, batch order, dropout masks). Pin the SAME seed to produce the SAME data split and initialization across every condition, so seed 3 in the baseline and seed 3 in the +A+B condition differ only in whether the suspect change is active, not in which examples they happened to see. This pairing is what lets you use a paired test (paired by seed) instead of an unpaired one: a paired test looks only at the within-seed differences, so shared noise across conditions (an unusually easy or hard split for seed 3) cancels out of the comparison instead of inflating the variance estimate. Run enough seeds, typically 5 to 10 for this kind of accuracy comparison, to get a real distribution of the paired difference rather than a single point estimate, and report the paired difference's mean, standard deviation, and a paired t-test p-value.
Reasoning about interactions instead of testing in isolation. Define the interaction term as what the combined condition's drop fails to explain from the sum of the individual conditions' drops:
Δinteraction=ΔA+B−(ΔA+ΔB)where each Δ is the baseline-minus-condition accuracy drop. Fix the sign convention explicitly before reading the result, because every branch of the interpretation flips with it. A drop here is baseline minus condition, so a POSITIVE Δ means the condition is WORSE than baseline and a bigger positive number is a bigger regression. Under that convention Δinteraction has three branches:
- Δinteraction>0, super-additive: the two changes interact and make each other worse. The combined drop is larger than the sum of the individual drops, so damage is appearing that neither change produces on its own. Common mechanisms are the new feature and the new component encoding overlapping or collinear signal that confuses the model when both are present, a shared upstream preprocessing bug that only triggers when both paths run, or the second change altering what the model learns to rely on from the first (feature crowding).
- Δinteraction≈0, independent: the changes act independently and additively, and either one's individual ablation result explains its share of the regression.
- Δinteraction<0, sub-additive: the pair together costs less than the two individual drops predicted, so the changes partly mask each other (often because they degrade the same examples, and those examples can only be got wrong once). Attributing each change its full individual drop here double-counts the damage.
Only the middle branch licenses reasoning about the two changes separately. This is exactly why testing each change in strict isolation and never testing the actual shipped combination can miss the real cause: neither isolated ablation shows the full regression, only the combined one does.
Worked example
Five seeded training runs per condition, same four conditions as above, same seed producing the same data split and initialization across all four conditions for a given seed index:
| seed | baseline | +A only | +B only | +A+B (shipped) |
|---|---|---|---|---|
| 0 | 0.8552 | 0.8427 | 0.8483 | 0.8285 |
| 1 | 0.8531 | 0.8399 | 0.8444 | 0.8315 |
| 2 | 0.8533 | 0.8438 | 0.8545 | 0.8313 |
| 3 | 0.8471 | 0.8476 | 0.8446 | 0.8285 |
| 4 | 0.8482 | 0.8456 | 0.8524 | 0.8289 |
| mean | 0.8514 | 0.8439 | 0.8488 | 0.8297 |
Paired difference per seed, shipped minus baseline: -0.0267, -0.0216, -0.0220, -0.0186, -0.0193. Mean paired difference dˉ=−0.02164, sample standard deviation of the differences sd=0.00318. The paired t-statistic with n=5 runs (4 degrees of freedom):
t=sd/ndˉ=0.00318/5−0.02164≈−15.2That gives a two-sided p-value of about 0.0001, far below any reasonable significance threshold (for example alpha = 0.05), so the shipped change's regression is real and not run-to-run noise.
Now the interaction check. Individual drops from baseline: ΔA=0.8514−0.8439=0.0075, ΔB=0.8514−0.8488=0.0025. (Both deltas come from the unrounded column means, 0.85138, 0.84392, 0.84884 and 0.82974; subtracting the four-decimal means printed in the table gives 0.0026 for ΔB instead of 0.0025, which is display rounding rather than a disagreement.) If the two changes acted independently, the combined drop should be about their sum, 0.0075+0.0025=0.0100. The actually observed combined drop is ΔA+B=0.8514−0.8297=0.0216. The interaction term is:
Δinteraction=0.0216−0.0100=0.0116Run that value back through the rule: Δinteraction=+0.0116, which is positive, which is the super-additive branch, so A and B interact and the combined condition is where the regression actually lives.
Stated in one consistent unit (one point here means one percentage point of accuracy, so 0.0075 is 0.75 points): A alone costs 0.75 points, B alone costs 0.25 points, so independence predicts 1.00 point for the pair, and the shipped pair actually costs 2.16 points. The extra 1.16 points is damage that belongs to the combination, not to either change. An investigation that only ran the two individual ablations and stopped there would have seen a three-quarter-point drop and a quarter-point drop, correctly concluded that neither alone explains a two-point regression, and then wrongly concluded "neither change is the problem" and shipped both, missing the actual cause, which only appears when both are active together.
Trade-offs and pitfalls
- Combinatorial cost. The full factorial design is cheap for 2 changes (4 configurations) and painful for 5 (32 configurations); use leave-one-out to triage first and reserve the full grid for a small suspect subset.
- Pairing discipline is what makes the significance test valid. If each condition draws its own independent data split or its own independent seed instead of a shared one, the paired test's variance-reduction assumption breaks and you need substantially more runs to reach the same statistical power.
- Statistical significance is not practical significance, and mixing units hides the difference. The p-value of about 0.0001 above says the 2.16-percentage-point drop is real, not that it is big; the same five-seed paired design would return an equally tiny p-value on a 0.05-point drop nobody would block a release over. Quote the effect size in the same unit every time (percentage points throughout this answer) so the magnitude is comparable across write-ups, and weigh it against the cost of reverting or delaying the change rather than letting the p-value stand in for size.
- Underpowered ablations look like "no effect." Five seeds is often enough for a clear regression like this one, but a small true effect can hide inside run-to-run noise with too few seeds; before concluding a change is innocent, check whether the study had enough seeds to detect a meaningfully sized effect in the first place, not just whether this particular p-value crossed a threshold.
- The common wrong turn is stopping at isolated ablations. Testing A alone and B alone and declaring victory because neither individually explains the full regression, without ever testing the shipped A+B combination, is the single most common way this kind of investigation goes wrong; interactions are common enough in real feature sets (shared preprocessing, collinear signal, feature crowding) that the combined condition belongs in the design from the start, not as an afterthought.
- A change that is genuinely a schema change to an existing feature rather than a brand-new feature follows the identical design: the schema-changed version and the pre-change version are simply the two arms of one of the paired conditions above.
Training checkpoints sometimes fail to restore correctly, causing longer retraining times or forcing a restart from scratch. What practical checks and strategies would you implement to make checkpointing reliable, and what would you consider when choosing a file format and storage location for checkpoints at scale?
Sample Answer
Direct answer. Checkpoint restore failures are almost always one of three things: a write that got interrupted partway (a torn/partial file), a version or schema mismatch between the checkpoint and the code trying to load it, or ambiguity about which checkpoint is actually the latest valid one, and each has a specific, cheap prevention.
Practical checks and strategies.
- Atomic writes. Write the checkpoint to a temporary file, then rename it to the final path only once the write completes successfully; a rename is atomic on virtually all filesystems, so a process that crashes mid-write leaves the temporary file corrupted but the "real" checkpoint path untouched, rather than a half-written file masquerading as a valid checkpoint.
- Versioning. Store checkpoints with an explicit version or step number in the filename or a manifest, and keep more than just the single latest one (for example, the last 3-5), so a corrupted or bad latest checkpoint doesn't force a restart from scratch, you fall back to the second-most-recent instead.
- Schema/format compatibility checks. Store a small metadata record alongside the checkpoint (library version, model architecture hash, key hyperparameters) and validate it against the current code BEFORE attempting the full weight-loading, so a genuine incompatibility fails fast with a clear message rather than partially loading and producing a subtly broken model.
File formats and storage considerations at scale. Prefer a format with built-in integrity checking or at least a way to verify completeness (a stored checksum, or a format that fails loudly on truncation rather than silently loading partial data). For cloud storage specifically, be aware that some object stores don't guarantee strong read-after-write consistency in every configuration, a checkpoint write that "succeeds" from the writer's perspective might not be immediately, reliably readable everywhere, worth confirming for your specific storage backend rather than assuming standard filesystem semantics apply. At scale (frequent large checkpoints), balance retention (keeping enough history to recover from a bad checkpoint) against storage cost, and consider whether you actually need full-precision checkpoints for every intermediate save or whether a compressed or lower-precision format is sufficient for all but the final/best checkpoint.
Write a Python script that parses a simple line-based training log (lines like epoch=1 step=100 loss=0.345 val_loss=0.32) and detects anomalies: a sudden loss spike (more than 3x the rolling mean), any NaN occurrence, or a stalled validation improvement over a window of steps. The script should log each alert with surrounding context and produce a short summary report at the end.
Sample Answer
Direct answer. The parser needs to track a small rolling window of recent loss values (to define "spike" relative to recent behavior, not a fixed threshold that would need retuning per run), explicitly check for NaN before comparing numbers (a NaN comparison silently evaluates to False in Python and would let a NaN slip past a naive spike check), and separately track validation loss to detect a stalled (flat) trend over a window of steps. Each alert carries the lines around it so it can be read without opening the raw log, and a summary function rolls the alert list up at the end.
import re
from collections import Counter, deque
def context(lines, lineno, radius=1):
lo, hi = max(0, lineno - 1 - radius), min(len(lines), lineno + radius)
return [l.strip() for l in lines[lo:hi]]
def parse_log_and_detect_anomalies(lines, window=5, spike_factor=3.0, stall_window=10):
pattern = re.compile(r"epoch=(\d+)\s+step=(\d+)\s+loss=([\w.]+)\s+val_loss=([\w.]+)")
recent_losses = deque(maxlen=window)
val_history = []
alerts = []
for lineno, line in enumerate(lines, start=1):
m = pattern.search(line)
if not m:
continue
epoch, step, loss_s, val_loss_s = m.groups()
try:
loss = float(loss_s)
except ValueError:
loss = float('nan')
try:
val_loss = float(val_loss_s)
except ValueError:
val_loss = float('nan')
if loss != loss: # a NaN is the only float that isn't equal to itself
alerts.append({"line": lineno, "type": "nan_loss", "context": context(lines, lineno)})
elif recent_losses:
rolling_mean = sum(recent_losses) / len(recent_losses)
if rolling_mean > 0 and loss > spike_factor * rolling_mean:
alerts.append({"line": lineno, "type": "loss_spike", "loss": loss,
"rolling_mean": round(rolling_mean, 4), "context": context(lines, lineno)})
if loss == loss:
recent_losses.append(loss)
if val_loss != val_loss: # "any NaN occurrence" covers the validation field too
alerts.append({"line": lineno, "type": "nan_val_loss", "context": context(lines, lineno)})
else:
val_history.append(val_loss)
if len(val_history) >= stall_window:
recent_window = val_history[-stall_window:]
if max(recent_window) - min(recent_window) < 1e-4:
alerts.append({"line": lineno, "type": "stalled_validation",
"window": stall_window, "context": context(lines, lineno)})
return {"total_lines_parsed": len(lines), "alerts": alerts, "alert_count": len(alerts)}
def summary_report(result):
out = [f"parsed {result['total_lines_parsed']} lines, {result['alert_count']} alerts"]
counts = Counter(a["type"] for a in result["alerts"])
for kind, n in counts.most_common():
hit_lines = [a["line"] for a in result["alerts"] if a["type"] == kind]
out.append(f" {kind}: {n} (first line {hit_lines[0]}, last line {hit_lines[-1]})")
return "\n".join(out)
Verified against a synthetic log with four known injected anomalies: a loss spike to 15.0 at step 50 (against a rolling mean well under 1.0), a literal val_loss=nan at step 60, a literal loss=nan at step 80, and a validation loss that flatlines at 0.32 for the last 12 steps. The two NaN injections are on different fields on purpose, because the question asks for ANY NaN occurrence and a parser that only guards the training loss will happily swallow a NaN validation loss, which is the more likely of the two to appear when an evaluation pass fails rather than the training step. The generator is worth showing, because the alert line numbers below only mean something if the fixture is reproducible:
def make_log(n=100, flat_tail=12, nan_loss_step=80, nan_val_step=60):
lines = []
for step in range(1, n + 1):
loss = round(0.985 ** step, 4)
val_loss = 0.32 if step > n - flat_tail else round(0.985 ** step + 0.02, 4)
loss_field = "nan" if step == nan_loss_step else (15.0 if step == 50 else loss)
val_field = "nan" if step == nan_val_step else val_loss
lines.append(f"epoch={(step - 1) // 25 + 1} step={step} loss={loss_field} val_loss={val_field}")
return lines
print(summary_report(parse_log_and_detect_anomalies(make_log()))) gives:
parsed 100 lines, 6 alerts
stalled_validation: 3 (first line 98, last line 100)
loss_spike: 1 (first line 50, last line 50)
nan_val_loss: 1 (first line 60, last line 60)
nan_loss: 1 (first line 80, last line 80)
All four injected anomaly types are detected at the correct line numbers.
Why the stall fires three times and not twelve, which is the part worth checking by hand. The stall condition only becomes true once the entire trailing window is flat, so with stall_window=10 a flat run of L steps alerts on its last L - 9 lines. The flat tail here occupies lines 89 to 100 (12 steps), so the first window that lies wholly inside it ends at line 98, and the alerts land on 98, 99 and 100. Line 97 does not alert, because its window still reaches back to line 88, which has val_loss=0.2845 and is not yet flat. Two useful consequences: the repetition is expected behavior rather than a bug (the condition stays true for every subsequent line until the trend changes), and the arithmetic doubles as a regression test on the detector itself, since a 12-step flat tail that starts alerting at 97 means the window boundary is off by one (that is what a 13-step tail produces, and it also pushes alert_count from 6 to 7).
Each alert carries its neighbours, which is what "surrounding context" has to mean for an on-call engineer reading the alert instead of the log:
{
"line": 50,
"type": "loss_spike",
"loss": 15.0,
"rolling_mean": 0.4916,
"context": [
"epoch=2 step=49 loss=0.4768 val_loss=0.4968",
"epoch=2 step=50 loss=15.0 val_loss=0.4897",
"epoch=3 step=51 loss=0.4626 val_loss=0.4826"
]
}
The neighbouring lines are what make this diagnosable at a glance: a loss of 15.0 against a rolling mean of 0.4916 is a 30x jump out of a step that looked entirely healthy, and the step after it is back to normal, which points at one bad batch rather than a diverging run.
The loss != loss idiom is deliberate rather than a typo: in IEEE floating point, NaN is the only value that is not equal to itself, so this check catches a NaN without needing to import math.isnan, and it's what makes the parser robust to a NaN slipping silently past the spike-detection branch (a plain loss > threshold comparison against NaN would evaluate to False, hiding the failure entirely). The same idiom is applied to val_loss as val_loss != val_loss, and the consequence there is different and worth stating: a NaN validation loss is not only alerted on, it is also kept OUT of val_history, so a run of NaNs can never be mistaken for a flat, stalled trend.
The summary report is summary_report above: a Counter over alert types plus the first and last line each type occurred on, which gives an on-call engineer a five-second read on what happened without scrolling through the raw alert list. It stays useful as the log grows, since a 40,000-line run with a long stall would otherwise bury the two single-line alerts under thousands of repeated stall entries.
Unlock Full Question Bank
Get access to all 39 Debugging and Testing ML Systems interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.