Debugging and Testing ML Systems Questions
Finding, diagnosing, and fixing problems in ML code, data, and models, and building tests that catch these problems before they reach users. Covers common ML pitfalls (data leakage, shape mismatches, silent training bugs, mis-specified loss or metrics), root-cause analysis of model regressions and production incidents (accuracy drops, calibration drift, intermittent or hard-to-reproduce failures), distributed-training-specific failures (multi-GPU divergence, intermittent OOM, precision-related instability), and the diagnostic tooling that supports it (reproducibility artifacts, structured logging, instrumentation). Also covers testing ML systems directly: unit tests for data and feature pipelines, validation checks for datasets and features, test oracles and acceptance criteria for probabilistic or non-deterministic model outputs, and integration and regression tests that catch model or pipeline regressions before deployment. Emphasizes the engineering rigor that keeps ML systems correct and maintainable.
Write a detailed bug report for an intermittent ML training failure that surfaces as a non-deterministic CUDA kernel launch happening only on specific GPU types. Include what timestamps and environment matrix (drivers, CUDA, cuDNN, framework versions) you would capture, minimal steps to reproduce, sample seeds, which logs to attach, an impact assessment, and your recommended next experiment or workaround.
Sample Answer
Direct answer
A bug report for an intermittent, GPU-type-specific training failure has to do the work a live debugging conversation would otherwise do informally, because the person who eventually picks it up will not have access to the machine, the moment, or your memory of what you tried. That means an exact environment matrix (down to driver, CUDA, the NVIDIA GPU computing toolkit, and cuDNN, NVIDIA's deep-learning kernel library, versions, since this exact bug class is version-sensitive by nature), a minimal reproduction that isolates the GPU type as the variable, every seed pinned so any attempted repro is comparing apples to apples, the specific logs to attach and why each one matters, an honest impact assessment, and a concrete next experiment, not just a description of the symptom.
Structured elaboration
Why this bug class specifically demands more than a symptom description. "Training sometimes crashes" is not actionable; a bug report is only as useful as the next person's ability to reproduce, or rule out, the same failure without your involvement. For a bug that is intermittent (does not happen every run) and hardware-specific (only certain GPU types), the report has to actively narrow the search space: which axis is the variable, hardware, driver version, a specific kernel, and provide enough detail that the report itself becomes evidence, not just a complaint.
Timestamps, and why they belong in the report as their own field, not folded into "environment." Record the exact wall-clock time of every failing run, alongside the exact wall-clock time of every successful run in the same investigation, in a consistent timezone, UTC (Coordinated Universal Time, the timezone-independent standard most infrastructure logs already use). This is what lets a later reader correlate a failure against something that changed at a specific moment but is not itself part of the environment matrix: a driver auto-update that rolled out to part of the fleet, a concurrent job that briefly contended for the same GPU, or a thermal-throttling event visible in a separate infrastructure dashboard. A bug report that only states "this happens sometimes" gives a later investigator nothing to line up against those external timelines; a report with exact timestamps lets them check in minutes whether the failures cluster around a specific external event instead of being uniformly distributed, which is itself a strong clue about the root cause.
What the environment matrix needs to capture, and why each field matters. GPU driver version and CUDA (Compute Unified Device Architecture, NVIDIA's parallel-computing platform) toolkit version together determine which low-level kernel implementations get selected for a given operation; a driver/CUDA combination that works on one GPU architecture can expose a genuine race condition or an unsupported code path on another. cuDNN (CUDA Deep Neural Network library, NVIDIA's library of GPU-accelerated primitives for deep learning) version matters separately because it ships its own kernel selection heuristics, independent of the CUDA toolkit version. Framework version (the training library itself) matters because kernel dispatch logic lives partly in the framework, not only in CUDA/cuDNN. Recording all four together, rather than just one, is what lets someone later correlate this exact failure against a known upstream bug report in any of the four projects.
Minimal reproduction steps. State the smallest model, batch size, and number of training steps that still reproduces the failure, and the specific GPU type(s) it reproduces on versus the GPU type(s) it was also tried on and did not reproduce on, since the negative result (GPU type X: 0 failures in N attempts) narrows the search space just as much as the positive one does.
Seeds. Pin and report every seed used in the reproduction attempts, since without them a colleague's "I couldn't reproduce it" is uninformative, they might simply have drawn a different random path through the same non-deterministic kernel. Also report whether the failure reproduces on the SAME seed across multiple attempts (which would suggest the non-determinism lives somewhere other than the obvious PyTorch/NumPy random state, likely the CUDA kernel scheduling itself) or only appears on some seeds and not others.
Which logs to attach. The full stdout/stderr from a failing run (the actual CUDA error string and code, not a paraphrase of it), the nvidia-smi output captured at or near the time of failure (GPU memory usage and utilization can distinguish an out-of-memory-adjacent race from a pure kernel-launch bug), and, if available, output from CUDA_LAUNCH_BLOCKING=1 (an environment variable that forces synchronous kernel launches, which converts an asynchronous, hard-to-localize CUDA error into one that points at the actual failing kernel call, at the cost of much slower execution, which is exactly why it is a debugging-only flag rather than something left on by default).
Impact assessment. State concretely which training runs and which GPU types are affected, and the failure rate observed so far in this specific investigation (framed honestly as "N failures in M attempts on this GPU type, this seed set," not extrapolated into a general production failure rate unless that rate has actually been measured at production scale).
Recommended next experiment or workaround. Propose the next diagnostic step that would most efficiently narrow the remaining hypothesis space (for example, running with CUDA_LAUNCH_BLOCKING=1 to localize the failing kernel, or bisecting across driver versions on the affected GPU type specifically), and a concrete interim workaround if one exists (routing training for the affected GPU type to a known-good driver version, or avoiding the specific operation suspected of triggering it), so the team is not blocked on training while the root cause is still under investigation.
Worked example
A filled bug report for exactly the scenario in the question:
Title: Intermittent CUDA kernel launch failure during training, reproduces only on GPU type A100, not on V100 or H100 (A100, V100, and H100 are successive generations of NVIDIA's data-center training GPUs)
Environment matrix
Field Affected runs Unaffected runs GPU type NVIDIA A100 (80GB) NVIDIA V100, NVIDIA H100 GPU driver 535.104.05 535.104.05 (same driver, different GPU type) CUDA toolkit 12.2 12.2 cuDNN 8.9.4 8.9.4 Framework PyTorch 2.3.0 PyTorch 2.3.0 Observed: 6 failures in 40 training attempts on A100 (15% failure rate in this investigation's sample; not yet measured at full production scale). 0 failures in 25 attempts on V100 and 0 failures in 15 attempts on H100, same code, same driver/CUDA/cuDNN/framework versions.
Timestamps of each A100 failure (UTC): 2026-07-14 03:12, 2026-07-14 09:47, 2026-07-15 03:20, 2026-07-16 03:15, 2026-07-18 03:09, 2026-07-19 22:58.
Tally of that list (a clustering claim that has not actually been counted is not evidence, so the count belongs in the report next to the raw timestamps):
classification timestamps count inside the 03:00-03:20 UTC window 07-14 03:12, 07-15 03:20, 07-16 03:15, 07-18 03:09 4 of 6, on 4 distinct days outside it 07-14 09:47, 07-19 22:58 2 of 6 Four of six landing inside a 20-minute window is the detail worth flagging first, ahead of the raw failure count. Twenty minutes is 20/1440 = 1.39% of a day, so if failures were spread uniformly through the clock you would expect 6 x 0.0139 = 0.08 of them to land in that window, not 4; the probability of seeing 4 or more of 6 there by chance is about 5.5e-07, roughly 1 in 1.8 million. That is far too strong to write off as coincidence, so it suggests correlation with a recurring scheduled job on the shared cluster (a nightly maintenance window is the working hypothesis, not yet confirmed) rather than a purely code-triggered race condition. The two failures outside the window are exactly why it stays a working hypothesis: a maintenance-window cause accounts for 4 of the 6 observations and something else, a second trigger or plain chance, still has to account for 07-14 09:47 and 07-19 22:58.
Minimal reproduction: a 2-layer transformer block, batch size 32, sequence length 128, fails intermittently within the first 50 training steps on A100. Does not require the full model or dataset; the minimal repro script isolates a single forward/backward pass repeated in a loop.
Seeds: reproduced with seeds 0, 17, and 42 on A100 (not seed-specific: same seed 0 both failed and succeeded across different attempts, which points toward non-determinism in kernel scheduling itself rather than in the framework-level random state).
Logs attached: (1) full stdout/stderr from a failing run, including the raw CUDA error string and code, not a summary; (2)
nvidia-smioutput captured within 5 seconds of the failure; (3) a run withCUDA_LAUNCH_BLOCKING=1set, which converts the normally-asynchronous CUDA error into one that names the specific failing kernel call, at the cost of roughly 3 to 5x slower execution for that run, which is why it is not run by default.Impact: blocks reliable overnight training runs specifically on the team's A100 nodes; V100 and H100 nodes are unaffected and available as an immediate workaround.
Recommended next experiment: (1) immediate workaround: route affected training jobs to V100 or H100 nodes while this is investigated; (2) next diagnostic step: check the cluster's job scheduler and infrastructure logs for anything recurring in the 03:00-03:20 UTC window to confirm or rule out the maintenance-window hypothesis the 4-of-6 clustering points at, and separately check whether the two out-of-window failures (07-14 09:47, 07-19 22:58) share any property with each other that the four in-window ones do not, since a hypothesis that only ever explains 4 of 6 is not yet a complete account; (3) in parallel, re-run the minimal repro with
CUDA_LAUNCH_BLOCKING=1on A100 across 20 attempts to identify the specific failing kernel by name, then check that kernel against NVIDIA's published A100 driver 535.x errata before assuming this is a novel bug rather than a known one.
Trade-offs and pitfalls
- Reporting failure rate without stating the sample size invites the reader to misjudge severity in either direction. "15% failure rate" sounds alarming or mild depending on whether it is 6 of 40 or 150 of 1000; always report both numbers together, and be explicit about whether the rate has been measured at production scale or only within this specific investigation's sample, per the reproducibility discipline of never presenting an estimate as a broader measurement than it actually is.
CUDA_LAUNCH_BLOCKING=1is a debugging tool, not a permanent fix. It serializes kernel launches, which can slow training by several times; attaching a run captured with it on is valuable evidence, but the recommendation should never be to leave it on in production training as a workaround for the underlying bug.- A same-seed-both-outcomes observation is a stronger clue than it looks, and is easy to bury in a wall of text. It specifically points toward non-determinism at the CUDA kernel-scheduling level rather than the framework's own random-number state, which changes where the next investigator should look first; surface it explicitly near the top of the report, not just in a logs appendix.
- A stated pattern that has not been counted is the same mistake as a failure rate quoted without its sample size. "The failures cluster overnight" is an assertion; "4 of 6, on 4 distinct days, inside a 20-minute window where 0.08 would be expected by chance" is evidence the next investigator can check and either build on or refute. Put the tally in the report next to the raw list, and let the prose summarise the tally rather than the other way round.
- A minimal reproduction that is not actually minimal wastes the next investigator's time. If the smallest failing case still requires the full model and dataset, keep reducing it (smaller batch size, fewer layers, synthetic data instead of the real dataset) before filing, since every extra unnecessary component in the repro is one more variable a colleague has to also rule out.
Design a comprehensive test-suite strategy for an ML codebase intended to prevent regressions: unit tests (data transforms, loss functions), integration tests (short training runs), dataset tests (schema and distribution checks), and model-behavior tests (smoke inputs, invariants). Describe which tests you would run at pull-request time versus nightly versus pre-deploy, why that split makes sense given each test's cost and signal, and give one concrete example test per category with an approximate runtime budget so the whole suite stays usable in CI.
Sample Answer
Direct answer. A test suite for an ML codebase needs four genuinely different kinds of tests, because each catches a different class of regression that the others miss: unit tests for the deterministic pieces of code, integration tests for whether components still work together, dataset tests for whether the data itself is still sane, and model-behavior tests for whether the model's actual outputs still make sense. Running all of them on every commit would make the pipeline too slow to be usable, so they're split across PR-time, nightly, and pre-deploy based on cost and how often each one is expected to catch something real.
The four categories, with one concrete example and a runtime budget each.
| Category | What it checks | Example | Runtime budget |
|---|---|---|---|
| Unit | Deterministic code: data transforms, loss functions | test_normalize_returns_zero_mean_unit_variance() on a fixed small array | under 1 second |
| Integration | Components work together | A training run for 5 steps on 50 synthetic examples completes without error and loss decreases | 30-90 seconds |
| Dataset | The data is still sane | Schema check (column types, required fields) plus a distributional check (feature means within N standard deviations of a stored baseline) on the latest data snapshot | 1-5 minutes (depends on data volume) |
| Model-behavior | The model's outputs still make sense | Smoke-input invariance test (a small, fixed set of inputs with known expected qualitative behavior, e.g. a spam classifier should not flag an empty string as spam) and a metric-regression test against a frozen golden dataset | 2-10 minutes |
When each runs, and why. Unit and a fast subset of integration tests run at pull-request time, they're cheap enough that a slow PR check would actively discourage people from running tests at all, and they catch the majority of code-level bugs before they ever reach a shared branch. Dataset tests and the full model-behavior suite run nightly against the latest production-like data snapshot, since they depend on data that changes daily and can be too slow or too dependent on external state to gate every PR. Model-behavior tests against the golden dataset also run again specifically at pre-deploy time, as a final gate right before a new model artifact goes to production, because that's the last point where catching a regression is cheap; catching the same regression after deploy means an incident instead of a blocked release.
Two closing practical notes. First, a release-validation checklist belongs at the pre-deploy stage specifically and should include not just model-quality thresholds but also privacy checks and an explicit rollback plan, automated into the deploy pipeline as hard gates rather than a manual sign-off step that people learn to skip under time pressure. Second, keep the nightly and pre-deploy dataset/model-behavior tests deterministic and reasonably fast even though they're not PR-blocking: a nightly suite that takes six hours to fail stops being useful days before anyone notices, because the feedback loop is too slow to actually change behavior.
A production model's performance drops sharply right after a change to the upstream data-ingestion pipeline. Outline a systematic debugging approach: validating raw inputs, comparing feature distributions before and after the pipeline change, verifying schema and null-handling behavior, replaying historical data through the new pipeline to check for silent differences, and using a shadow deployment to isolate whether the regression is in the data or the model. Describe the preventative tests you would add so a future pipeline change can't cause the same regression silently.
Sample Answer
Direct answer
Work from cheapest to most expensive: validate the raw inputs first, then compare feature distributions and verify schema and null-handling on the same data, then replay historical data through the new pipeline to see whether the pipeline code itself behaves differently on identical inputs, and only then reach for a shadow deployment to separate a data-side cause from a model-side one, since a shadow deployment is the most expensive diagnostic and the earlier steps usually already answer the question. The single most common failure this sequence protects against is treating an aggregate, whole-population distribution check as sufficient when the real regression is concentrated in one segment the aggregate view dilutes into invisibility.
Structured elaboration
Validating raw inputs. Before touching any statistics, confirm the pipeline change did not simply break ingestion: row counts per upstream source in the expected range, required fields present, types conforming to the expected schema. This is the cheapest check and catches gross breakage (a source that silently stopped sending a field, a partial ingestion failure) before spending effort on subtler distributional analysis that a gross failure would make meaningless anyway.
Comparing feature distributions before and after the pipeline change. For each feature, compare its distribution from before the change against its distribution from after, using a distribution-free test such as the two-sample Kolmogorov-Smirnov (KS) test, appropriate here because feature distributions are typically continuous and not reliably normal, so a test that does not assume a particular shape is the safer default. That test returns two numbers and the comparison below turns on telling them apart. The KS statistic is the single largest vertical gap between the two samples' cumulative distribution curves (for each value on the horizontal axis, the curve gives the share of that sample falling at or below it), so it runs from 0 when the two distributions are identical to 1 when they do not overlap at all, and it reads directly as "at their widest disagreement, these two samples differ by this much accumulated share." The p-value says only how unlikely a gap that large would be if both samples really did come from the same distribution; it says nothing about how large the gap is. Those come apart at scale: with enough rows, a gap far too small to matter still produces a tiny p-value, so the statistic is the effect size and has to be read alongside the p-value rather than replaced by it. Critically, do this segmented by whatever cohort dimensions are available (region, customer type, input source), not only on the whole population: a change concentrated in one segment can be small enough relative to the whole population that an aggregate-only comparison misses it, while the same comparison restricted to the affected segment shows it clearly.
Verifying schema and null-handling behavior. Confirm explicitly, not by inference, that types, allowed value ranges, and the specific handling of missing values match the contract the current model was trained against. A silent change in null-handling, for example a field that used to arrive as an explicit null and now silently gets coerced to zero by an upstream default, produces a systematic, hard-to-spot shift that a generic distribution comparison can sometimes miss if the coerced value happens to fall within an otherwise plausible range.
Replaying historical data through the new pipeline. Take a fixed batch of raw data from before the change, whose resulting features were already recorded by the old pipeline, and run that identical raw batch through the new pipeline code. Diff the newly computed features against the originally recorded ones for the exact same input rows. This is the cleanest possible causal test available: because the raw input is literally identical, any difference in the output features can only come from the pipeline code itself, not from the real world having changed, which is exactly the distinction needed to separate "the code changed behavior" from "the underlying data genuinely shifted."
Using a shadow deployment to isolate data versus model. Once the cheaper checks above have narrowed things down, run the current production model against the new pipeline's live features in shadow mode, scored but never served to users, and compare its offline performance on a labeled sample of that shadow traffic against its known historical performance. If the same, unchanged model degrades when fed the new pipeline's features, the fault sits in the features, not the model. If a retrain is also under consideration, comparing the old model and a newly retrained model against the identical new feature set is what isolates whether any remaining gap is model-side rather than data-side.
Preventative tests for the future. Turn the historical-replay diff from an ad hoc investigation step into an automated regression test: a fixed historical batch with its expected feature output checked into the test suite, run automatically whenever the pipeline code changes, so a future silent behavior change fails a test instead of reaching production. Add an explicit schema and null-handling contract test that asserts the exact type, range, and null-treatment behavior the model depends on. Add a scheduled, segmented distribution-drift monitor comparing live feature distributions against the training-time reference on an ongoing basis, not only around known pipeline changes, since not every silent regression will coincide with a deploy someone remembers to check against.
Worked example
Comparing an order_value feature before and after a pipeline change that silently broke currency normalization for international orders only, on synthetic data (4,000 rows before, 4,000 after, about 25% flagged international in each), first at the whole-population level, then segmented. The seed is pinned so every number below is reproducible rather than a one-off draw:
import numpy as np
from scipy import stats
N, INTL_SHARE, BUG_FACTOR = 4000, 0.25, 1.35
MU, SIGMA = np.log(40.36) - 0.125, 0.5 # order_value is lognormal, mean about 40
rng = np.random.default_rng(5268) # pinned, so this table reproduces
before = rng.lognormal(MU, SIGMA, N)
before_intl = rng.random(N) < INTL_SHARE
after = rng.lognormal(MU, SIGMA, N)
after_intl = rng.random(N) < INTL_SHARE
# The bug: currency normalization silently mis-scales international orders only.
after = np.where(after_intl, after * BUG_FACTOR, after)
comparisons = [
("whole population", before, after),
("domestic segment only", before[~before_intl], after[~after_intl]),
("international segment only", before[before_intl], after[after_intl]),
]
print(f"{'comparison':<28}{'n before':>9}{'n after':>9}{'KS stat':>10}{'p-value':>12}")
for label, a, b in comparisons:
res = stats.ks_2samp(a, b)
print(f"{label:<28}{len(a):>9}{len(b):>9}{res.statistic:>10.4f}{res.pvalue:>12.2g}")
m_before, m_after = before[before_intl].mean(), after[after_intl].mean()
print(f"\ninternational mean order_value: {m_before:.2f} -> {m_after:.2f} "
f"(ratio {m_after/m_before:.4f}, bug applied {BUG_FACTOR})")
Output:
comparison n before n after KS stat p-value
whole population 4000 4000 0.0610 6.8e-07
domestic segment only 2994 3000 0.0188 0.65
international segment only 1006 1000 0.2441 1e-26
international mean order_value: 40.08 -> 54.11 (ratio 1.3499, bug applied 1.35)
| comparison | KS statistic | p-value |
|---|---|---|
| whole population | 0.0610 | 6.8e-07 |
| domestic segment only | 0.0188 | 0.65 |
| international segment only | 0.2441 | 1e-26 |
The whole-population test does detect something (p=6.8e-07), but its KS statistic of 0.061 looks like a mild, easy-to-dismiss shift. Read literally, 0.061 says that at the point where the before and after cumulative curves are furthest apart they differ by about 6 percentage points of accumulated mass, which on a 0-to-1 scale is close to the identical end. With 4,000 rows on each side, even a gap that small is comfortably significant, and that combination, a tiny statistic with a convincing p-value, is exactly the trap: read the p-value alone and it looks like a confirmed problem, read the statistic alone and it looks like nothing, and neither reading tells you where to go next. Segmenting shows what actually happened: the domestic segment shows no significant difference at all (p=0.65, indistinguishable from noise), while the international segment alone shows a far larger and far more significant shift (KS statistic 0.244, p effectively zero). On the same 0-to-1 scale, 0.244 means those two curves separate by over 24 percentage points of accumulated mass at their widest, exactly four times the whole-population gap, and it does that on only about a quarter of the rows, which is precisely why averaging it in with the unaffected three quarters shrank it to 0.061. The international segment's mean order value moved from 40.08 to 54.11, a ratio of 1.3499, recovering almost exactly the 1.35x mis-scaling the underlying bug actually applied. An investigation that stopped at the whole-population number would have seen a modest, ambiguous signal; segmenting turned it into an unambiguous, localized, and nearly root-cause-identifying result.
Trade-offs and pitfalls
- Reaching for the shadow deployment before the cheaper checks wastes the most expensive tool on a question the earlier steps usually already answer. Sequencing this from cheapest to most expensive is not just tidiness, it avoids spending shadow-deployment effort re-discovering what a distribution comparison would have shown directly.
- An aggregate-only distribution comparison can genuinely miss a real, severe, segment-concentrated regression, exactly as the worked example shows; always segment by every cohort dimension available before concluding a feature is unaffected.
- The historical-replay diff is the step most often skipped, and it is the one that actually distinguishes a code bug from a genuine real-world shift. Without it, a team can spend real effort investigating "why did the world change" when the honest answer is "the pipeline code changed and the world did not."
- A replay test only covers the inputs it was built from. It will not catch a bug that only manifests on an input pattern the historical batch never contained, so it complements, rather than replaces, the ongoing distribution-drift monitor.
Draft the artifact checklist an ML-specific production-outage postmortem needs beyond a generic incident postmortem template: which model and dataset versions, experiment IDs, feature-store snapshots, and reproduction steps should be captured so the incident can actually be reproduced and understood later, not just narrated. Explain why each item matters specifically for an ML system rather than a generic service outage.
Sample Answer
Direct answer
A generic service postmortem's artifact list (code commit, config, deploy timestamp, request logs) captures everything needed to reproduce a stateless service, because its behavior is fully determined by code and config. A machine learning (ML) system's behavior additionally depends on a large, opaque model artifact and the data that produced it, neither of which is recoverable from source code alone, so an ML-specific postmortem needs five items a generic template does not ask for: the exact model artifact version, the exact training dataset version, the training-run's experiment identifier, a feature-store snapshot of what the model actually saw at serving time, and reproduction steps that tie all of those together into one runnable recipe. Without these five, the incident can be narrated but not actually reconstructed.
Structured elaboration
Model version or artifact identifier. Record the exact model artifact identifier (a registry version, a content hash of the weights file), not a human description like "the model deployed on the 12th." This matters specifically for ML because a model's decision logic lives in a large binary weights artifact that is not derivable from source code the way a compiled service is; two "deploys of the same code" at different times can carry genuinely different model weights if a retrain happened in between, so the artifact identifier, not the deploy timestamp, is the only thing that actually pins down what was making decisions during the incident.
Training dataset version or snapshot identifier. Record the exact identifier of the dataset snapshot the model was trained on, including whatever preprocessing or feature computation was applied as of that snapshot. This matters specifically for ML because datasets are not static the way source code is: they get backfilled, corrected, or silently regenerated over time, so "the training data" as it exists today can already differ from what actually trained the incident-era model. Without a pinned snapshot identifier, a later attempt to inspect "the training data" may be looking at data that no longer matches what produced the deployed model.
Experiment or training-run identifier. Record the identifier linking to the full experiment-tracking record for the training run that produced the deployed model: hyperparameters, random seed, training code commit, and training environment. This matters specifically for ML because reproducing an ML incident sometimes requires reproducing the training process itself, for example to determine whether the incident is a one-off artifact of that specific training run or something that would recur from any training run using the same code and data. A generic service never needs this distinction, since its code is inherently, deterministically reproducible from source; a model trained again from the identical code and data can still land on meaningfully different weights depending on the random seed and other run-specific factors, so the experiment record is what tells you whether the run itself was unusual.
Feature-store snapshot at serving time. Record the actual feature values the model received for a representative sample of the affected requests, not just the raw request payload. This matters specifically for ML because a model's real input is computed by a separate feature pipeline, often running asynchronously and sometimes diverging from what the training pipeline computed for the same logical input, a failure mode commonly called training-serving skew. A generic postmortem's request log tells you what was asked; an ML postmortem additionally needs to know what the model actually saw, since those two things are not the same system and can silently disagree.
Reproduction steps. A runnable recipe that ties the four items above together: fetch this exact model artifact, feed it this exact set of feature vectors (either replayed from the feature-store snapshot or recomputed from the pinned dataset version with the pinned feature-computation code), using this exact serving code version, and confirm the same failing output reproduces. This matters specifically for ML because a generic incident's reproduction usually reduces to "revert to commit X and replay the request," a single axis to pin. An ML incident's reproduction has to pin the model artifact, the feature computation, and the serving code as three separate axes, any one of which can independently explain a discrepancy if left unpinned, so the checklist has to specify exactly which artifacts to fetch and in what order, not just point at a commit hash.
Worked example
A filled-in artifact block for a hypothetical incident, showing the shape these five items take in practice, distinct from the narrative sections (timeline, impact, corrective actions) a generic postmortem template already provides:
| field | example value | why a generic template would not ask for this |
|---|---|---|
| model artifact version | fraud-model registry v482, sha256:9f2a... | a deploy timestamp alone cannot tell you which of several retrains was actually live |
| training dataset snapshot | transactions_train snapshot 2026-06-14T02:00Z | the live dataset today may already differ from what trained this model |
| experiment / training-run ID | mlflow run a1b2c3, seed=17, training code commit e4f5g6 | needed to tell whether the incident is reproducible from any training run or specific to this one |
| feature-store snapshot (sample) | user_id=88213, snapshot_ts=2026-06-20T09:14Z, features={...} | the request payload alone does not show what the model's feature pipeline actually computed |
| reproduction recipe | "load artifact v482, replay feature snapshot above through serving code at commit h7i8j9, confirm output matches the incident's logged prediction" | ties the other four together into something an engineer can actually run, not just read |
Each row exists because a generic postmortem's usual artifact (a code commit and a deploy timestamp) genuinely does not capture it: none of the middle three rows have any equivalent in a stateless service's incident record.
Trade-offs and pitfalls
- Capturing the model version but not the dataset version is a common half-measure. It tells you what weights were live but not what produced them, which blocks any attempt to understand whether a training-data problem caused the incident.
- A feature-store snapshot captured too late, after the incident window, is close to useless if the online feature pipeline is itself mutable or time-decaying. Snapshot capture has to happen close to the actual serving time of the affected requests, not whenever someone gets around to writing the postmortem.
- Recording artifact identifiers without a working reproduction recipe still leaves the incident unreproducible in practice. The identifiers are necessary but not sufficient; someone still has to know, and document, the exact sequence of steps to fetch and combine them.
- This checklist is deliberately narrow. It covers only the artifacts an ML system needs beyond a generic postmortem, not the postmortem's document structure, timeline narrative, or corrective-action process, which a generic incident-postmortem template already handles perfectly well for an ML incident just as it would for any other.
List five quick sanity checks or 'toy model' experiments you could run to determine, within minutes, whether a large-model production problem originates from input data, model code, or infrastructure. For each, state the expected command or action and what result would implicate that category.
Sample Answer
Direct answer. The goal of a five-minute sanity pass is to categorize the problem (data, code, or infrastructure) cheaply, before committing to a deep investigation in any one direction, using checks that are individually fast and each rule a whole category in or out.
Five checks.
- Single-batch inference with a fixed seed. Action: run one small, known batch through the model with all randomness pinned. Expected result if this passes: the model and its basic inference path are functioning; a failure here (a crash, or a NaN, not-a-number, the value a floating-point operation returns when it has no valid answer, which then silently propagates through every computation downstream of it until something finally checks) implicates model code or model artifact corruption directly, before you've spent any time on data or infra.
- Run on CPU instead of GPU. Action: force the same input through on CPU. If the result differs meaningfully from the GPU result, that's a strong signal the problem is GPU/driver/kernel-specific (numeric precision differences, a nondeterministic op, a hardware fault), not a logic bug that would reproduce identically anywhere.
- Verify container/environment checksum or version. Action: confirm the running container image hash (or key library versions) matches what was actually tested and approved. If it doesn't match, stop investigating the model entirely, you're debugging the wrong artifact.
- Feed a known, previously-correct input and compare to a stored expected output. Action: run a fixed regression example through the current pipeline and diff against a golden result computed earlier. If this input, which used to produce a known-correct result, now produces something different, that isolates the regression to something that changed SINCE the golden result was captured, narrowing the search window immediately. That narrows TIME rather than category, so this check needs one more move to land where the other four do: re-run the same golden input against the PREVIOUS container image. If the old image reproduces the golden result, the regression is in code or environment, specifically in whatever changed between the two images. If the old image also fails on the golden input, the model artifact or the environment beneath both images is what went bad, not the newly deployed code. And if the golden input passes on the current image while live traffic is still wrong, nothing on the model side changed at all and the problem is in the input data, since the only thing that differs between the passing run and the failing one is the input.
- Check upstream data freshness/availability with a simple timestamp query. Action: query the most recent timestamp available in the primary upstream data source. If it's stale (older than expected), that implicates a data-pipeline problem upstream of the model entirely, and further model-level debugging is premature until the data pipeline itself is fixed.
Why five minutes, and why these five specifically. Each check is fast (seconds to a couple of minutes) and, critically, each one's PASS or FAIL result rules out or strongly implicates one of the three broad categories (input data, model code, infrastructure). Check 4 is the one that needs two steps rather than one to get there, since its first result narrows the time window and only the re-run against the previous image converts that into a category, which is still well inside the five-minute budget. Running all five in sequence gives a rough triage classification before anyone commits real time to a specific hypothesis, cheap insurance against spending an hour debugging model code for a problem that turns out to be an environment mismatch found in 30 seconds by check 3.
Unlock Full Question Bank
Get access to all Debugging and Testing ML Systems interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.