MLOps: Monitoring, Retraining, and Lifecycle Management Questions

Operating machine learning systems reliably over time. Covers model and data monitoring, drift and degradation detection, feedback loops, retraining and model-freshness strategy, versioning and model registries, and pipeline and workflow orchestration. Focuses on keeping deployed models healthy and reproducible across their lifecycle.

MediumTechnical
71 practiced

What is model calibration, why does it matter for probabilistic decision-making, and how would you measure and monitor it in production? Describe one technique to improve calibration and when you'd apply it, and how your online/offline monitoring approach changes for a multi-class classifier with delayed labels.

HardTechnical
59 practiced

A production model fails only intermittently, and sample replay reproduces the failure inconsistently. Describe a forensic playbook: correlation with upstream pipeline metrics, per-feature distribution comparisons, hidden confounders, time-of-day patterns, and an algorithmic approach to RANK likely root causes using distribution divergence and explainability signals (for example SHAP). Include what you'd document for a postmortem.

MediumTechnical
65 practiced

Sketch a PyTorch training loop (pseudocode is fine) that supports incremental training from an existing checkpoint, and logs model-version metadata (dataset snapshot id, hyperparameters, training start/end timestamps) to a model registry. What safety checks would you run before writing the new model version?

HardTechnical
50 practiced

A monitoring system flags multivariate drift but doesn't say which of hundreds of upstream features or data sources is responsible. Design an approach to attribute the drift to specific features: what statistical and algorithmic tools you'd reach for and why, how you'd combine multiple signals, how you'd prioritize which features to investigate first, and how you'd quantify your confidence in an attribution.

MediumTechnical
56 practiced

What is catastrophic forgetting in continual learning? Give two mitigation strategies when incrementally retraining a model, one replay-based and one regularization-based, and describe a scenario where each is preferable.

Unlock Full Question Bank

Get access to all MLOps: Monitoring, Retraining, and Lifecycle Management interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.