Debugging and Testing ML Systems Questions

Finding, diagnosing, and fixing problems in ML code, data, and models, and building tests that catch these problems before they reach users. Covers common ML pitfalls (data leakage, shape mismatches, silent training bugs, mis-specified loss or metrics), root-cause analysis of model regressions and production incidents (accuracy drops, calibration drift, intermittent or hard-to-reproduce failures), distributed-training-specific failures (multi-GPU divergence, intermittent OOM, precision-related instability), and the diagnostic tooling that supports it (reproducibility artifacts, structured logging, instrumentation). Also covers testing ML systems directly: unit tests for data and feature pipelines, validation checks for datasets and features, test oracles and acceptance criteria for probabilistic or non-deterministic model outputs, and integration and regression tests that catch model or pipeline regressions before deployment. Emphasizes the engineering rigor that keeps ML systems correct and maintainable.

HardTechnical
46 practiced

Write a detailed bug report for an intermittent ML training failure that surfaces as a non-deterministic CUDA kernel launch happening only on specific GPU types. Include what timestamps and environment matrix (drivers, CUDA, cuDNN, framework versions) you would capture, minimal steps to reproduce, sample seeds, which logs to attach, an impact assessment, and your recommended next experiment or workaround.

HardTechnical
53 practiced

Design a comprehensive test-suite strategy for an ML codebase intended to prevent regressions: unit tests (data transforms, loss functions), integration tests (short training runs), dataset tests (schema and distribution checks), and model-behavior tests (smoke inputs, invariants). Describe which tests you would run at pull-request time versus nightly versus pre-deploy, why that split makes sense given each test's cost and signal, and give one concrete example test per category with an approximate runtime budget so the whole suite stays usable in CI.

HardTechnical
52 practiced

A production model's performance drops sharply right after a change to the upstream data-ingestion pipeline. Outline a systematic debugging approach: validating raw inputs, comparing feature distributions before and after the pipeline change, verifying schema and null-handling behavior, replaying historical data through the new pipeline to check for silent differences, and using a shadow deployment to isolate whether the regression is in the data or the model. Describe the preventative tests you would add so a future pipeline change can't cause the same regression silently.

MediumTechnical
72 practiced

Draft the artifact checklist an ML-specific production-outage postmortem needs beyond a generic incident postmortem template: which model and dataset versions, experiment IDs, feature-store snapshots, and reproduction steps should be captured so the incident can actually be reproduced and understood later, not just narrated. Explain why each item matters specifically for an ML system rather than a generic service outage.

EasyTechnical
56 practiced

List five quick sanity checks or 'toy model' experiments you could run to determine, within minutes, whether a large-model production problem originates from input data, model code, or infrastructure. For each, state the expected command or action and what result would implicate that category.

Unlock Full Question Bank

Get access to all Debugging and Testing ML Systems interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.