MLOps: Monitoring, Retraining, and Lifecycle Management Questions

Operating machine learning systems reliably over time. Covers model and data monitoring, drift and degradation detection, feedback loops, retraining and model-freshness strategy, versioning and model registries, and pipeline and workflow orchestration. Focuses on keeping deployed models healthy and reproducible across their lifecycle.

HardTechnical
71 practiced

You want to detect multivariate drift robustly. Compare three approaches: a multivariate statistical test (e.g. Hotelling's T-squared), dimensionality reduction followed by univariate tests, and a trained two-sample classifier. Discuss computational cost and interpretability trade-offs for a production system.

MediumTechnical
64 practiced

Compare the Kolmogorov-Smirnov (KS) test and the Population Stability Index (PSI) for detecting numeric feature drift. Cover their assumptions, sensitivity to sample size, and when you'd prefer one over the other in production monitoring. Then walk through a worked PSI calculation by hand for a small binned example.

EasyTechnical
64 practiced

Define model monitoring for a production ML system. List the key categories of signals you'd track (data/feature drift, model performance, latency, resource usage, and business KPIs), explain why each matters operationally, and clarify the difference between monitoring and observability with a short example of when better observability (not just monitoring) speeds up root-cause identification.

MediumTechnical
55 practiced

Compare Kolmogorov-Smirnov (KS), Population Stability Index (PSI), Kullback-Leibler (KL) divergence, and Maximum Mean Discrepancy (MMD) as drift-detection tools. For each, discuss sensitivity to sample size, applicability to multivariate or categorical data, and numerical stability. Then explain how a trained two-sample classifier (domain classifier) can serve as an alternative to all four, and what its practical failure modes are.

MediumTechnical
56 practiced

What is catastrophic forgetting in continual learning? Give two mitigation strategies when incrementally retraining a model, one replay-based and one regularization-based, and describe a scenario where each is preferable.

Unlock Full Question Bank

Get access to all 16 MLOps: Monitoring, Retraining, and Lifecycle Management interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.