Model Evaluation and Validation Questions

Measuring whether a model is good enough to trust and ship. Covers metric selection for classification, regression, and ranking (precision/recall, ROC-AUC, calibration, RMSE), offline validation design, evaluation-metric-to-business-objective alignment, and production safety guardrails. Emphasizes choosing metrics that reflect real objectives and avoiding misleading evaluations.

HardTechnical
123 practiced

Explain cluster-randomized experiments, where you randomize at the level of a user, household, or region rather than an individual event, and why clustering is necessary when there is spillover or correlated behavior within a cluster. Define the intra-cluster correlation coefficient and describe how it affects the required sample size and variance estimation.

HardSystem Design
86 practiced

Design a comprehensive evaluation framework for a large-scale search or recommendation product serving tens of millions of users monthly. Cover offline metrics (NDCG@k, recall@k, MAP), how you would correct for position and exposure bias, the online metrics you would track (CTR, revenue, retention), the logging schema needed for counterfactual evaluation, and how offline evaluation, online A/B tests, and champion-challenger deployment fit together.

MediumTechnical
83 practiced

Implement stratified k-fold cross-validation from scratch (no scikit-learn) for a multi-class dataset, preserving class proportions across folds as closely as possible, including for classes with fewer examples than folds. Then explain how you would adapt the same idea for multi-label data, where samples can carry more than one class, to preserve label co-occurrence patterns across folds.

HardTechnical
77 practiced

Prove, or sketch a proof, that ROC-AUC equals the probability that a randomly chosen positive example is scored higher than a randomly chosen negative example. State the assumptions you use, and explain how this relates to the Mann-Whitney U statistic.

EasyTechnical
81 practiced

For a multi-class classification problem, explain micro versus macro averaging of precision, recall, and F1. Walk through a concrete example where label frequencies are skewed (for instance a customer-support intent classifier with 10 unbalanced intents), showing how the two averages diverge, and advise which one you would present to stakeholders and why.

Unlock Full Question Bank

Get access to all Model Evaluation and Validation interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.