MLOps: Monitoring, Retraining, and Lifecycle Management Questions
Operating machine learning systems reliably over time. Covers model and data monitoring, drift and degradation detection, feedback loops, retraining and model-freshness strategy, versioning and model registries, and pipeline and workflow orchestration. Focuses on keeping deployed models healthy and reproducible across their lifecycle.
Design an ML observability dashboard for a deployed classification model (for example, a loan-approval or fraud model). List at least 8 metrics grouped into performance, data/drift, infrastructure, and business categories, specify SLIs/SLOs and one proactive alert rule per category, and name concrete tooling choices (for example Prometheus/Grafana, Evidently, or a managed equivalent) you'd wire it into.
From a Site Reliability Engineer's perspective, how do the operational requirements of an ML model service differ from a traditional stateless microservice? Discuss determinism, dependency on training data, model versioning, rollback complexity, and reproducibility, and explain how incidents manifest differently.
Design an alerting strategy for model and data drift that balances sensitivity with alert fatigue. Define severity tiers, aggregation windows, and suppression/decorrelation rules; explain how you'd select thresholds using historical production data (for example via control charts) rather than guessing; and give three concrete example alerting rules with metric, threshold, and immediate on-call action. Include how you'd flag inputs that look adversarial or malicious rather than merely drifted.
Define SLI, SLO, and SLA in the context of ML systems. Propose a set of SLIs and SLO thresholds for an online recommendation model, covering inference latency (p95), availability, a prediction-quality KPI (for example CTR uplift), and feature freshness, and describe how you'd implement alerting and escalation on SLO breach.
Design an alerting taxonomy that clearly differentiates a job-FAILURE alert from a data-quality-REGRESSION alert on the same pipeline. Propose example SLIs and thresholds for each category, who gets notified for each (on-call engineer, data owner, or downstream consumer team), and how you would keep this distinction from collapsing into one generic 'something is wrong' page.
Unlock Full Question Bank
Get access to all 9 MLOps: Monitoring, Retraining, and Lifecycle Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.