Model Evaluation and Validation Questions
Measuring whether a model is good enough to trust and ship. Covers metric selection for classification, regression, and ranking (precision/recall, ROC-AUC, calibration, RMSE), offline validation design, evaluation-metric-to-business-objective alignment, and production safety guardrails. Emphasizes choosing metrics that reflect real objectives and avoiding misleading evaluations.
Explain cluster-randomized experiments, where you randomize at the level of a user, household, or region rather than an individual event, and why clustering is necessary when there is spillover or correlated behavior within a cluster. Define the intra-cluster correlation coefficient and describe how it affects the required sample size and variance estimation.
Sample Answer
Cluster-randomized experiments randomize treatment at the group level (users, households, schools, regions) rather than individuals. You use them whenever interference or correlated behavior makes individual randomization invalid: e.g., within-household spillover, network effects, or shared environments where one person’s treatment affects others’ outcomes. Randomizing clusters preserves the causal contrast and avoids contamination.
Intra-cluster correlation coefficient (ICC, ρ) measures the similarity of outcomes within clusters: ρ = σ_b² / (σ_b² + σ_w²), where σ_b² is between-cluster variance and σ_w² is within-cluster variance. ICC ranges 0–1; higher ρ means outcomes within the same cluster are more alike.
Impact on sample size and variance:
- Design effect (DE) = 1 + (m − 1)·ρ, with m = average cluster size. DE inflates variance relative to independent individuals.
- Effective sample size Neff ≈ N / DE (N = total individuals). So required total N must be multiplied by DE to retain power.
- Variance of treatment effect estimates must account for clustering: Var_clust = DE · Var_ind. Ignoring ICC underestimates standard errors, inflating Type I error.
Analysis recommendations:
- Power calculations should use estimated ICC and cluster sizes; increasing number of clusters is more effective than increasing cluster size when ρ>0.
- Use cluster-robust standard errors, mixed-effects models (random intercepts), or GEE with exchangeable correlation to correctly estimate SEs.
- Ensure sufficient degrees of freedom (enough clusters) because inference depends on number of clusters, not individuals.
Example: m=20, ρ=0.05 → DE = 1 + 19·0.05 = 1.95, so nearly double the sample needed compared with individual randomization.
Design a comprehensive evaluation framework for a large-scale search or recommendation product serving tens of millions of users monthly. Cover offline metrics (NDCG@k, recall@k, MAP), how you would correct for position and exposure bias, the online metrics you would track (CTR, revenue, retention), the logging schema needed for counterfactual evaluation, and how offline evaluation, online A/B tests, and champion-challenger deployment fit together.
Sample Answer
Requirements & goals:
- Evaluate ranking quality (relevance), business outcomes (CTR, revenue, retention), and long-term user satisfaction at 100M DAU with low risk.
Offline metrics & protocol:
- Relevance: NDCG@k (Normalized Discounted Cumulative Gain: rewards a ranked list for placing relevant items near the top, discounts relevant items that appear further down, and normalizes against the best possible ordering so scores are comparable across queries), Recall@k (of all the relevant items that exist, the fraction that showed up anywhere in the top k), MAP@k (Mean Average Precision: the average precision computed at each rank position where a relevant item appears in the top k, then averaged across queries) computed on holdout sessions; use session-level aggregation and per-user temporal splits (train on t, test on t+delta).
- Calibration & confidence: compute confidence intervals via bootstrapping by user.
- Diversity & novelty: catalog-based measures (intra-list diversity, coverage).
Correcting position/exposure bias:
- Propensity scoring via logged exposure probabilities (from serving logs): use IPS (inverse propensity scoring) and SNIPS (Self-Normalized IPS: the same reweighting idea as IPS, but rescaled by the sum of the propensity weights rather than by the raw sample count, which reduces the wild variance plain IPS can have when some items had very low exposure probability) to unbiasedly estimate CTR and NDCG.
- Train position-bias models (e.g., an examination model / PBM) to estimate propensities when they aren't logged directly.
- Use doubly robust estimators combining IPS with outcome models to reduce variance.
Online metrics:
- Immediate: raw CTR, conversion rate, revenue per thousand impressions.
- Short-term engagement: session length, day-over-day retention.
- Long-term value: 7/30/90-day retention, LTV, churn rate, downstream purchases.
Logging schema (must be complete & immutable):
- Event id, user_id (hashed), timestamp, session_id, request_id, placement_id, rank_list (item_ids + positions), served_probabilities (model score, softmax prob), exposure_flag per item, click/engagement events with timestamps, item metadata (owner, category), context (device, region), policy_version, experiment_id, traffic_bucket, reward signals (purchase, watch_time), prior user-state feature snapshot. Ensure deterministic replay keys and a sampling indicator for subsampling.
Counterfactual eval & offline simulator:
- Offline simulator: replay logged requests, simulate alternative policies using logged propensity or importance weights. Include synthetic user-response models learned from logs for stress tests (e.g., adversarial content).
- Use IPS/SNIPS/doubly robust for policy evaluation. Validate simulators by backtesting on historical A/B tests.
A/B testing & system components:
- Experiment platform: traffic allocation, randomization (user-level), exposure logging, kill switch.
- Metrics pipeline: near-real-time aggregator for guardrail metrics, weekly cohort analyses for long-term metrics.
- Policy rollout: staged (canary to ramp), automatic risk checks (statistical significance and business bounds).
- Analysis tools: automated uplift estimation, sequential testing with alpha-spending (a rule for how much of your total false-positive budget you're allowed to spend by peeking at a running experiment's results early, so repeated interim looks don't silently inflate the overall false-positive rate the way naive repeated significance testing would), and variance reduction via stratification/ANCOVA (Analysis of Covariance: adjusting an experiment's outcome metric for a pre-experiment covariate, like each user's prior engagement level, to strip out predictable noise before comparing groups, which tightens confidence intervals without needing more traffic).
How offline evaluation, online A/B tests, and champion-challenger deployment fit together:
- Offline metrics are the fast, cheap FILTER: any new policy must beat the current champion on the offline holdout and the offline simulator (using IPS/SNIPS/DR) before it is allowed anywhere near real traffic. This is the stage that screens out most bad candidates for near-zero cost.
- A/B testing is the causal VALIDATION step: a policy that clears the offline bar gets a randomized, low-traffic online test against the current champion, because even debiased offline estimators can miss position-bias or feedback-loop effects that only appear with real exposure.
- Champion-challenger is the ONGOING PRODUCTION pattern once a challenger has won its A/B test: instead of a full one-shot replacement, the challenger is promoted to serve a small, sustained slice of live traffic (e.g. 5-10%) permanently alongside the incumbent champion, with the same online metrics tracked continuously rather than for a fixed test window. This catches slow drift, seasonality, and small regressions a short A/B window would miss, and gives an instant, no-redeploy rollback (shift traffic back to the champion) if the challenger degrades later. Only after a sustained period of the challenger matching or beating the champion does it get promoted to be the new champion, at which point a fresh challenger can be tested against it.
- Together the three form a funnel of increasing cost and decreasing risk: offline (cheap, many candidates screened, imperfect signal) -> A/B (moderate cost, causal, time-boxed) -> champion-challenger (small ongoing cost, the steady-state safety net that a time-boxed test can't provide).
Long-term impact tracking:
- Cohort-based LTV and retention dashboards, causal impact analyses (difference-in-differences, synthetic controls), monitor content-provider effects and feedback loops (popularity bias).
- Periodic offline retraining with debiased labels and causal features to prevent feedback loops.
Trade-offs & operational notes:
- Logging volume: sample some heavy fields but keep deterministic keys for replay.
- Bias-variance: IPS is unbiased but high variance; prefer doubly robust estimators in production.
- Privacy: hash/anonymize PII; consider differential privacy for aggregate dashboards.
This framework provides unbiased offline evaluation, safe online experimentation, and a champion-challenger steady state to iterate recommendation policies at scale with continuous, low-cost safety monitoring.
Implement stratified k-fold cross-validation from scratch (no scikit-learn) for a multi-class dataset, preserving class proportions across folds as closely as possible, including for classes with fewer examples than folds. Then explain how you would adapt the same idea for multi-label data, where samples can carry more than one class, to preserve label co-occurrence patterns across folds.
Sample Answer
Direct answer. Assign each class's examples to folds independently, round-robin, so every class's proportions land as evenly as possible across folds, including classes with fewer members than folds (some folds simply get zero examples of that ultra-rare class, which is unavoidable, not a bug).
Code (executed and verified: exact coverage, proportions checked per fold).
import numpy as np
def stratified_kfold_indices(y, k, seed=0):
y = np.asarray(y)
rng = np.random.default_rng(seed)
folds = [[] for _ in range(k)]
for c in np.unique(y):
idx = np.where(y == c)[0]
rng.shuffle(idx)
for i, ix in enumerate(idx):
folds[i % k].append(ix) # round-robin within each class
return [np.array(sorted(f)) for f in folds]
Worked example (recomputed). On a synthetic 83-row, 3-class dataset (50/30/3 split, so the smallest class has fewer members than the 5 requested folds), the resulting folds had sizes 17/17/17/16/16, with the size-3 class distributed as 1/1/1/0/0 across folds, i.e. every fold gets either exactly the class's fair round-robin share or the class's fair share minus one when it's too small to reach every fold, and every index across all 5 folds was verified to appear in exactly one fold, matching the original 83 rows exactly.
Structured elaboration: adapting to multi-label data. With multi-label data, a sample carries several class memberships at once, so simple per-class round-robin (which assumes one label per row) breaks down: assigning by one label ignores the others. The standard fix is iterative stratification (Sechidis et al.), which repeatedly assigns whichever remaining sample has the rarest, most under-represented label combination to whichever fold currently needs it most, rather than processing one class at a time; scikit-learn's IterativeStratification (in the scikit-multilearn or iterstrat packages) implements this.
Trade-offs and pitfalls. Round-robin assignment within each class is simple and exactly preserves the overall class ratio in the limit, but it is NOT immune to combining classes in a way that skews a joint (multi-label) distribution, which is precisely why multi-label stratification needs a different algorithm rather than just running the single-label version once per label. A class with fewer members than folds guarantees at least one fold sees zero examples of it; if that class is the one you care most about, consider fewer folds, or a repeated/nested design that gives every fold a chance to see it across repeats.
Prove, or sketch a proof, that ROC-AUC equals the probability that a randomly chosen positive example is scored higher than a randomly chosen negative example. State the assumptions you use, and explain how this relates to the Mann-Whitney U statistic.
Sample Answer
Assumptions: scores s(x) are real-valued; positive and negative instances are drawn independently from distributions F_pos, F_neg. If ties occur, treat them by splitting probability (half credit).
Sketch of proof:
-
Define AUC as area under ROC: AUC = ∫{0}^{1} TPR(FPR^{-1}(u)) du. Equivalently, for a threshold t, FPR(t)=P(s_neg ≥ t), TPR(t)=P(s_pos ≥ t). Then
AUC = ∫{-∞}^{∞} TPR(t) dFPR(t)
= ∫ P(s_pos ≥ t) d[P(s_neg ≥ t)]. -
Convert the integral to a double probability over independent samples s_pos, s_neg using Fubini/measure-change:
AUC = P(s_pos > s_neg) + 0.5·P(s_pos = s_neg).
Intuition: For fixed pair (s_pos, s_neg), consider thresholds t between them; the pair contributes 1 to the integral if s_pos > s_neg, 0 if s_pos < s_neg, and 0.5 if equal.Tiny numeric check: let positives = [0.9, 0.6] and negatives = [0.5, 0.2]. List all n_pos × n_neg = 4 (pos, neg) pairs and score each: (0.9, 0.5): pos>neg; (0.9, 0.2): pos>neg; (0.6, 0.5): pos>neg; (0.6, 0.2): pos>neg. All 4 of 4 pairs have pos>neg, so P(s_pos > s_neg) = 4/4 = 1.0 and the formula above gives AUC = 1.0. This matches the empirical AUC computed directly from the ROC curve: sorting all four scores descending gives [0.9(pos), 0.6(pos), 0.5(neg), 0.2(neg)], a perfect separation of positives above negatives, which is exactly what AUC = 1.0 means. This is the concrete case the abstract Fubini/measure-change step is asserting in general: AUC literally counts (and normalizes) how many positive-negative pairs are correctly ordered.
More formal line: AUC = E_{s_neg}[P(s_pos ≥ s_neg)] = P(s_pos > s_neg) + P(s_pos = s_neg)·P(counting tie as half) hence the stated identity.
Relation to Mann–Whitney U:
Mann–Whitney U tests whether values from one sample tend to be larger than another. For samples of sizes n_pos, n_neg, U equals the number of (pos, neg) pairs where s_pos > s_neg plus 0.5 times ties. Normalizing U by n_pos·n_neg gives exactly the empirical AUC. Thus AUC is the probability interpretation of the Mann–Whitney U statistic.
For a multi-class classification problem, explain micro versus macro averaging of precision, recall, and F1. Walk through a concrete example where label frequencies are skewed (for instance a customer-support intent classifier with 10 unbalanced intents), showing how the two averages diverge, and advise which one you would present to stakeholders and why.
Sample Answer
Approach & notation
Given a 3×3 confusion matrix C where C[i][j] = true class i predicted as j (rows = true, cols = pred). For class k (k=1..3):
- True Positives (TP_k) = C[k][k]
- False Positives (FP_k) = sum over i != k of C[i][k]
- False Negatives (FN_k) = sum over j != k of C[k][j]
Per-class formulas
Precision_k:
precision_k = TP_k / (TP_k + FP_k)
Recall_k:
recall_k = TP_k / (TP_k + FN_k)
F1_k (per-class):
f1_k = 2 * precision_k * recall_k / (precision_k + recall_k)
Macro vs Micro F1
- Macro-F1: average of per-class F1s
macro_f1 = (f1_1 + f1_2 + f1_3) / 3
- Micro-F1: compute global TP, FP, FN (sum over classes) then F1 from aggregated precision/recall
micro_precision = sum_k TP_k / (sum_k TP_k + sum_k FP_k)
micro_recall = sum_k TP_k / (sum_k TP_k + sum_k FN_k)
micro_f1 = 2 * micro_precision * micro_recall / (micro_precision + micro_recall)
Worked example (skewed customer-support intents)
Take a simplified 3-intent slice of a support classifier (billing_question, cancel_subscription, technical_issue) where billing_question dominates traffic, the kind of skew a real 10-intent classifier shows:
Confusion matrix C (rows = true, cols = predicted):
pred_billing pred_cancel pred_technical row total
true_billing 760 25 15 800
true_cancel 8 10 2 20
true_technical 5 2 8 15
column total 773 37 25 835
Per-class precision/recall/F1:
billing: TP=760, FP=13, FN=40 -> precision=760/773=0.983, recall=760/800=0.950, f1=0.966
cancel: TP=10, FP=27, FN=10 -> precision=10/37=0.270, recall=10/20=0.500, f1=0.351
technical: TP=8, FP=17, FN=7 -> precision=8/25=0.320, recall=8/15=0.533, f1=0.400
macro_f1 = (0.966 + 0.351 + 0.400) / 3 = 0.572
micro: sum TP=778, sum FP=57, sum FN=57, so micro_precision = micro_recall = 778/835 = 0.932, micro_f1 = 0.932
The two averages diverge by 0.36: micro-F1 (0.932) is dragged almost entirely by the large billing_question class the model already handles well, while macro-F1 (0.572) exposes that the model is mediocre on the two rare, business-important intents (cancel_subscription and technical_issue). For a stakeholder report I would present macro-F1 here, since a headline 0.93 would hide that the model is failing on the rare intents a support team most needs correctly routed, and I would name the two weak per-class F1 scores explicitly rather than only the macro average.
When prefer Macro-F1
Use macro-F1 when class balance matters and you want equal weight per class: e.g., detection of rare but critical classes (fraud, disease). Macro-F1 penalizes poor performance on minority classes; micro-F1 can be dominated by large classes and hide failures on rare but important classes.
Unlock Full Question Bank
Get access to all Model Evaluation and Validation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.