Model Evaluation and Validation Questions

Measuring whether a model is good enough to trust and ship. Covers metric selection for classification, regression, and ranking (precision/recall, ROC-AUC, calibration, RMSE), offline validation design, evaluation-metric-to-business-objective alignment, and production safety guardrails. Emphasizes choosing metrics that reflect real objectives and avoiding misleading evaluations.

EasyTechnical
82 practiced

Define overfitting and underfitting, and explain how learning curves (training versus validation performance as a function of training-set size) let you tell them apart. Given a curve where training error stays low while validation error stays high and roughly flat, what's going on and what would you change? Then describe how the curves would look instead if the model were underfitting, and what you would do in that case.

MediumSystem Design
78 practiced

Design how you would detect data and concept drift in a production system handling roughly a thousand requests a second. Compare statistical tests such as Kolmogorov-Smirnov, Population Stability Index, and ADWIN, discuss detection with limited labeled data (for example unsupervised proxies like autoencoder reconstruction error), and describe how you would set thresholds, window sizes, and alerting so the system stays sensitive to real drift without false-alarming on seasonality.

EasyTechnical
93 practiced

Explain the precision-recall trade-off. Using two concrete business examples where the right call goes in opposite directions (for instance an email spam filter versus a medical diagnostic screen), walk through which metric you would prioritize in each and how you would set the operating threshold given the different costs and class prevalences involved.

HardTechnical
87 practiced

Design an approximate streaming ROC-AUC calculator in Python that ingests an incoming stream of (y_true, score) pairs under limited memory, supports incremental updates, and can be queried for an approximate AUC at any time. Discuss the algorithmic choices (fixed binning, t-digest, quantile sketches), the memory-versus-accuracy trade-off, and how you would merge sketches computed on different shards.

MediumTechnical
88 practiced

Explain uplift (heterogeneous treatment effect) modeling and the metrics used to evaluate it, such as the Qini coefficient and uplift@k. Describe a business use case, for example a marketing campaign, where uplift modeling is clearly preferable to simply predicting conversion probability directly, and how you would run an experiment to validate that targeting by uplift actually increases ROI.

Unlock Full Question Bank

Get access to all Model Evaluation and Validation interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.