MLOps: Monitoring, Retraining, and Lifecycle Management Questions

Operating machine learning systems reliably over time. Covers model and data monitoring, drift and degradation detection, feedback loops, retraining and model-freshness strategy, versioning and model registries, and pipeline and workflow orchestration. Focuses on keeping deployed models healthy and reproducible across their lifecycle.

MediumSystem Design
65 practiced

Design an ML observability dashboard for a deployed classification model (for example, a loan-approval or fraud model). List at least 8 metrics grouped into performance, data/drift, infrastructure, and business categories, specify SLIs/SLOs and one proactive alert rule per category, and name concrete tooling choices (for example Prometheus/Grafana, Evidently, or a managed equivalent) you'd wire it into.

MediumTechnical
48 practiced

From a Site Reliability Engineer's perspective, how do the operational requirements of an ML model service differ from a traditional stateless microservice? Discuss determinism, dependency on training data, model versioning, rollback complexity, and reproducibility, and explain how incidents manifest differently.

MediumTechnical
64 practiced

Design an alerting strategy for model and data drift that balances sensitivity with alert fatigue. Define severity tiers, aggregation windows, and suppression/decorrelation rules; explain how you'd select thresholds using historical production data (for example via control charts) rather than guessing; and give three concrete example alerting rules with metric, threshold, and immediate on-call action. Include how you'd flag inputs that look adversarial or malicious rather than merely drifted.

EasyTechnical
55 practiced

Define SLI, SLO, and SLA in the context of ML systems. Propose a set of SLIs and SLO thresholds for an online recommendation model, covering inference latency (p95), availability, a prediction-quality KPI (for example CTR uplift), and feature freshness, and describe how you'd implement alerting and escalation on SLO breach.

MediumSystem Design
61 practiced

Design an alerting taxonomy that clearly differentiates a job-FAILURE alert from a data-quality-REGRESSION alert on the same pipeline. Propose example SLIs and thresholds for each category, who gets notified for each (on-call engineer, data owner, or downstream consumer team), and how you would keep this distinction from collapsing into one generic 'something is wrong' page.

Unlock Full Question Bank

Get access to all 9 MLOps: Monitoring, Retraining, and Lifecycle Management interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.