Model Evaluation and Validation Questions
Measuring whether a model is good enough to trust and ship. Covers metric selection for classification, regression, and ranking (precision/recall, ROC-AUC, calibration, RMSE), offline validation design, evaluation-metric-to-business-objective alignment, and production safety guardrails. Emphasizes choosing metrics that reflect real objectives and avoiding misleading evaluations.
Define overfitting and underfitting, and explain how learning curves (training versus validation performance as a function of training-set size) let you tell them apart. Given a curve where training error stays low while validation error stays high and roughly flat, what's going on and what would you change? Then describe how the curves would look instead if the model were underfitting, and what you would do in that case.
Sample Answer
Overfitting: a model captures noise or idiosyncrasies in the training data and performs much better on training than on unseen data. Underfitting: a model is too simple to capture the underlying signal and performs poorly on both training and validation.
Learning curves: plot model performance (e.g., accuracy or RMSE) on training and validation sets versus training set size.
- Overfitting signature: training performance is high, validation performance is much lower; gap persists as data grows. Validation may improve slowly with more data. Worked example (train/validation accuracy at increasing training-set size n): n=100: train=99%, val=58%; n=500: train=98%, val=61%; n=1000: train=97%, val=62%; n=2000: train=97%, val=63%. The training curve sits near-flat and high the whole way, the validation curve creeps up only slowly, and the gap between them (roughly 35-40 points) barely narrows even as n quadruples: that persistent, wide, flat gap is exactly the "training error stays low while validation error stays high and roughly flat" pattern the question describes.
- Underfitting signature: both training and validation performance are poor and close to each other; adding more data doesn't help much. Worked example: n=100: train=64%, val=61%; n=500: train=66%, val=64%; n=1000: train=67%, val=65%; n=2000: train=67%, val=66%. Both curves sit low from the start, stay within a couple of points of each other throughout, and neither one climbs meaningfully as n grows: the model has hit a capacity ceiling that more data can't fix, unlike the overfitting case where the gap (not the level) was the problem.
Practical BI pipeline changes
- If I detect overfitting: add regularization or simplify the model used in automated reports (e.g., switch from a high-cardinality decision tree to a regularized logistic regression or apply L1 feature selection). Also enforce cross-validation in model training and reduce dimensionality (aggregate categorical levels, drop redundant features) before publishing predictions to dashboards. Applied to the worked example above, adding L2 regularization and dropping redundant features would be expected to pull training accuracy down from the high-90s toward the mid-80s while lifting validation accuracy from the low 60s toward the low-to-mid 70s, narrowing the gap rather than closing it in one step.
- If I detect underfitting: enrich features and increase model capacity: add derived features (time-based aggregations, interaction terms), loosen regularization, or use a more flexible model (e.g., gradient boosted trees) in the modeling stage. Re-run validation and update ETL to include the new features so dashboards reflect improved predictions. Applied to the worked underfitting example above, swapping in gradient boosted trees and adding derived features would be expected to lift both curves together, e.g. training accuracy from ~67% toward the mid-80s and validation from ~66% toward the high-70s, since the ceiling here was capacity, not overfitting.
These changes balance predictive accuracy with interpretability and operational constraints typical in BI deliverables.
Design how you would detect data and concept drift in a production system handling roughly a thousand requests a second. Compare statistical tests such as Kolmogorov-Smirnov, Population Stability Index, and ADWIN, discuss detection with limited labeled data (for example unsupervised proxies like autoencoder reconstruction error), and describe how you would set thresholds, window sizes, and alerting so the system stays sensitive to real drift without false-alarming on seasonality.
Sample Answer
Start by stating the identifying assumptions:
- Covariate shift: p_train(y|x) ≈ p_prod(y|x) but p_train(x) ≠ p_prod(x).
- Label shift: p_train(x|y) ≈ p_prod(x|y) but p_train(y) ≠ p_prod(y).
Covariate shift
- Detection test:
- Train a domain classifier to distinguish train vs. prod x (using unlabeled prod features). If classifier AUC ≫ 0.5, covariate shift exists. Alternatively use two-sample tests (KS per feature, multivariate MMD).
- Correction technique:
- Importance weighting: weight each training example by w(x)=p_prod(x)/p_train(x).
- Estimate w(x) via density ratio methods (KLIEP: Kullback-Leibler Importance Estimation Procedure, which directly fits a ratio model by minimizing KL divergence to the true density ratio; uLSIF: unconstrained Least-Squares Importance Fitting, which fits the same ratio via a squared-error objective that has a closed-form solution and is faster to compute), or via classifier-based density ratio: fit logistic regression D(x)=P(prod|x) then w(x) ≈ D(x)/(1−D(x)) × (N_train/N_prod). In practice, the classifier-based approach is the default first choice: it reuses infrastructure you already have for training classifiers and is easy to debug by inspecting D(x) directly. KLIEP and uLSIF are more specialized density-ratio estimators worth reaching for only if the classifier-based weights prove unstable (D(x) saturating near 0 or 1) or the feature space is very high-dimensional.
- Practicalities: clip/regularize weights, calibrate D(x), or use Kernel Mean Matching (KMM: solves directly for weights that match the mean feature embedding of train and prod in a kernel space, without ever estimating the ratio pointwise) to avoid extreme weights. KMM is a rarer fallback for when even KLIEP/uLSIF produce badly unstable weights.
- Adapt offline evaluation:
- Compute importance-weighted metrics: E_prod[L] ≈ sum_i w(x_i) L(y_i, f(x_i)) / sum_i w(x_i).
- Use weighted cross-validation and report variance; if weights high-variance, use stratified or bootstrap confidence intervals.
Label shift
- Detection test:
- If you can obtain a small labeled sample from production, compare class prior distributions directly.
- Without labels, use Black Box Shift Estimation (BBSE): train a classifier on training data, compute its confusion matrix on a validation set, apply it to unlabeled prod X to get predicted label frequencies; large mismatch indicates label shift. BBSE is the standard default for label shift specifically (as opposed to the covariate-shift methods above): it needs only the classifier's confusion matrix and unlabeled production predictions, never new labeled production data.
- Correction technique:
- Confusion-matrix inversion / BBSE: estimate target class priors π_prod by solving C · π_prod = q_prod, where C is confusion matrix P(pred|true) and q_prod is vector of P(pred) on prod data. Use EM (Saerens et al.) to jointly refine priors and posterior estimates.
- Reweight training examples by ratio π_prod(y)/π_train(y) when computing loss.
- Account for classifier calibration and regularize inversion if C is ill-conditioned.
- Adapt offline evaluation:
- Recompute metrics by reweighting per-class: metric_prod ≈ sum_y (π_prod(y)/π_train(y)) · E_train[ L | y ].
- If you only have predicted labels on prod, use corrected confusion-matrix estimates to adjust precision/recall values.
Comparing the statistical tests: KS, PSI, and ADWIN
- Kolmogorov-Smirnov (KS): a two-sample test comparing the empirical CDFs of one feature in train vs. a fixed production window; returns a test statistic and a p-value. Best suited to a scheduled batch comparison between two fixed samples (e.g., "this week vs. training").
- Population Stability Index (PSI): buckets a feature's values into bins (typically deciles derived from the training distribution), then computes PSI = Σ_bins (pct_prod − pct_train) · ln(pct_prod / pct_train) across bins. Unlike KS it is not a formal hypothesis test: it is an industry-standard magnitude score with widely used rule-of-thumb thresholds (PSI < 0.1: no significant shift; 0.1-0.25: moderate shift worth watching; > 0.25: major shift requiring action). It is simpler to compute and to explain to non-statistician stakeholders than KS, which is why it is the common choice for a dashboard-facing drift metric.
- ADWIN (Adaptive Windowing): unlike KS and PSI, which compare two fixed windows chosen in advance, ADWIN is a streaming algorithm that maintains a variable-length window of recent values and automatically shrinks the window from the old end whenever it detects the older and newer sub-windows have significantly different means. This suits continuous, real-time monitoring where you do not want to hand-pick a window size, at the cost of being harder to reason about and tune than a KS/PSI check run on a fixed schedule.
Detection with limited labeled data
- Unsupervised proxy: train an autoencoder on training-time features, then compute the reconstruction error (mean squared error between input and reconstruction) for each incoming production example. A rising rolling-average reconstruction error means production inputs increasingly do not look like anything the autoencoder learned to reconstruct well: a drift signal needing zero production labels, complementing the domain-classifier and BBSE approaches above (which also need no production labels, but only tell you a distribution changed, not how unusual any individual example is).
Worked example: thresholds, window size, and seasonality
- Track PSI weekly on a feature (e.g., average session length) against a 30-day trailing training baseline. Week 1 PSI = 0.04 (normal noise). By week 3, PSI has climbed to 0.30. With the rule-of-thumb threshold above (PSI > 0.25 triggers action), this fires an alert at week 3, not week 1.
- Window size: use a rolling window spanning at least one full seasonal cycle (e.g., 7 days for a feature with weekday/weekend structure) rather than a single day, and compare same-phase windows (this week's Mon-Sun vs. last week's Mon-Sun, or vs. the same week 4 weeks back) rather than arbitrary day-to-day windows. This alignment is what actually prevents ordinary weekly seasonality from false-alarming the detector.
- Alerting: fire a warning at PSI in [0.1, 0.25) or KS p-value < 0.05 sustained for 2 consecutive windows (filtering one-off noise), and page on-call at PSI > 0.25 or 3 consecutive breaching windows.
Practical notes
- Choose test based on available resources: domain classifier and MMD require only unlabeled prod X; BBSE needs a reliable classifier and a validation confusion matrix.
- Monitor continually: track feature drift, domain-classifier AUC, estimated weights and π_prod; alert when corrections large.
- Uncertainty: propagate estimation error (bootstrap) and guard against over-correction (clipping, L2 on weights).
- When assumptions break, prefer collecting labels in prod and retraining or using domain-adaptive models (importance-weighted training, domain-adversarial networks, or joint-covariate/label-shift estimators).
Explain the precision-recall trade-off. Using two concrete business examples where the right call goes in opposite directions (for instance an email spam filter versus a medical diagnostic screen), walk through which metric you would prioritize in each and how you would set the operating threshold given the different costs and class prevalences involved.
Sample Answer
Precision measures the fraction of positive predictions that are correct; recall (sensitivity) measures the fraction of true positives that are found. They trade off because raising the decision threshold usually increases precision (fewer false positives) but lowers recall (more false negatives), and lowering the threshold does the opposite.
Example 1: Email spam filtering:
- Business goal: minimize user annoyance from false positives (important emails marked spam) while keeping inboxes reasonably clean.
- Priority: precision > recall (but not extreme). Missing some spam (lower recall) is acceptable; misclassifying legitimate mail (false positive) is costly.
- Thresholding: choose a threshold that achieves high precision on validation data (e.g., 95%+) while measuring user impact. Use precision-recall curve and set threshold where marginal gain in precision outweighs lost recall. Consider whitelist/soft quarantine for borderline cases.
- Worked cost-based threshold: using expected_cost = C_FN·P_FN + C_FP·P_FP with C_FP=$5 (annoyance of a legitimate email wrongly blocked) and C_FN=$0.10 (nuisance of a spam email getting through), spam prevalence π=5%, and these validation-set rates at three candidate thresholds:
θ=0.3: TPR=0.98, FPR=0.15 -> P_FN=π(1−TPR)=0.001, P_FP=(1−π)·FPR=0.1425 -> expected_cost = 0.10×0.001 + 5×0.1425 = $0.7126 per email
θ=0.5: TPR=0.95, FPR=0.05 -> P_FN=0.0025, P_FP=0.0475 -> expected_cost = 0.10×0.0025 + 5×0.0475 = $0.2378 per email
θ=0.7: TPR=0.88, FPR=0.01 -> P_FN=0.006, P_FP=0.0095 -> expected_cost = 0.10×0.006 + 5×0.0095 = $0.0481 per email
θ=0.7 has the lowest expected cost of the three and still keeps sensitivity above an 85% recall floor. Pushing the threshold even higher would lower the pure expected cost further, since C_FP so heavily outweighs C_FN, but that recall floor, not the unconstrained cost minimum, is what keeps the chosen threshold practical for a spam filter that still needs to catch most spam.
Example 2: Medical diagnosis for a serious but treatable disease:
- Business goal: catch as many true cases as possible; downstream confirmatory tests exist.
- Priority: recall > precision. Missing a sick patient (false negative) has high cost; false positives are tolerable if they lead to further testing.
- Thresholding: choose a low threshold to maximize sensitivity subject to an acceptable false-positive rate. Optimize expected utility using a cost matrix: expected_cost = C_FN * P_FN + C_FP * P_FP, where P_FN/P_FP come from validation. Calibrate model probabilities and pick threshold that minimizes expected cost (or maximizes net benefit), adjusting for disease prevalence (use Bayes’ theorem) so that predicted risk maps to real-world positive rates. If prevalence is low, precision will naturally be low at high recall: communicate PPV to clinicians and use triage (e.g., risk stratification) rather than a single binary cut.
Practical steps for both:
- Calibrate probabilities (Platt/Isotonic), plot precision-recall and ROC, and compute expected cost using stakeholder-specified C_FP/C_FN and prevalence. If costs asymmetric, directly optimize threshold on expected cost or use cost-sensitive loss during training (class weights or focal loss). Validate threshold on holdout data reflecting real prevalence and monitor post-deployment to adjust for drift.
Design an approximate streaming ROC-AUC calculator in Python that ingests an incoming stream of (y_true, score) pairs under limited memory, supports incremental updates, and can be queried for an approximate AUC at any time. Discuss the algorithmic choices (fixed binning, t-digest, quantile sketches), the memory-versus-accuracy trade-off, and how you would merge sketches computed on different shards.
Sample Answer
Direct answer. Bin incoming scores into a fixed set of buckets and keep two running counts per bucket (positives and negatives seen so far); AUC can then be recovered from those bucket counts alone via the same rank-based formula as the exact computation, using each bucket's midpoint rank as an approximation for the true rank of every point that landed in it.
Code (executed and verified against the exact scikit-learn AUC).
import numpy as np
class StreamingApproxAUC:
def __init__(self, n_bins=1000, score_min=0.0, score_max=1.0):
self.n_bins, self.lo, self.hi = n_bins, score_min, score_max
self.pos_counts = np.zeros(n_bins)
self.neg_counts = np.zeros(n_bins)
def _bin(self, score):
b = int((score - self.lo) / (self.hi - self.lo) * self.n_bins)
return min(max(b, 0), self.n_bins - 1)
def add(self, y_true, score):
b = self._bin(score)
(self.pos_counts if y_true == 1 else self.neg_counts)[b] += 1
def auc(self):
P, N = self.pos_counts.sum(), self.neg_counts.sum()
if P == 0 or N == 0:
return float("nan")
cum_neg_below = np.cumsum(self.neg_counts) - self.neg_counts
rank_score = cum_neg_below + 0.5 * self.neg_counts # half-credit for same-bucket ties
return float(np.sum(self.pos_counts * rank_score) / (P * N))
def merge(self, other):
merged = StreamingApproxAUC(self.n_bins, self.lo, self.hi)
merged.pos_counts = self.pos_counts + other.pos_counts
merged.neg_counts = self.neg_counts + other.neg_counts
return merged
Worked example (recomputed on 20,000 points against sklearn's exact roc_auc_score). With 1,000 bins, exact AUC = 0.89629 versus approximate AUC = 0.89628, an absolute error of 0.000004. With a much coarser 20 bins, the approximation degrades to 0.89514, an error of 0.001144, over 280 times larger, quantifying exactly what "coarser binning trades accuracy for memory" costs in practice. Merging two independently-maintained shards (each covering half the stream) reproduced the single-pass approximate AUC exactly, confirming the sketch can be computed in parallel and combined afterward.
Structured elaboration: algorithmic choices and the memory/accuracy trade-off. Fixed binning (used above) needs only O(n_bins) memory total, independent of stream length, and merges trivially (element-wise sum of two count arrays), which is why it's the natural first choice; its weakness is that scores concentrated in a narrow range get coarse resolution unless bins are chosen adaptively for that range. A t-digest or quantile sketch instead adapts bin (centroid) density to where the data actually is, giving much better resolution in the score distribution's tails at the cost of a more complex merge operation and a data structure that isn't just a flat array. Choosing between them is really a question of whether your score distribution is roughly uniform over its range (fixed binning is fine) or heavily skewed/concentrated (a t-digest earns its complexity).
Trade-offs and pitfalls. More bins costs proportionally more memory but the accuracy gain is not linear, going from 20 to 1,000 bins here cut the error by roughly 280x for a 50x increase in bin count, showing sharply diminishing but still real returns; picking a bin count is a genuine memory-versus-accuracy budget decision, not a default to leave unexamined. The half-credit convention for same-bucket ties (0.5 * self.neg_counts) matters more as bins get coarser, since more true ties get artificially created by bucketing that wouldn't have been ties at full precision; this is the mechanism behind most of the accuracy loss at 20 bins.
Explain uplift (heterogeneous treatment effect) modeling and the metrics used to evaluate it, such as the Qini coefficient and uplift@k. Describe a business use case, for example a marketing campaign, where uplift modeling is clearly preferable to simply predicting conversion probability directly, and how you would run an experiment to validate that targeting by uplift actually increases ROI.
Sample Answer
Uplift (treatment-effect) modeling predicts the causal incremental effect of applying a treatment (e.g. a marketing action) on an individual's outcome versus not treating them. Unlike standard outcome prediction P(y|x), uplift estimates tau(x) = E[Y|X=x, T=1] - E[Y|X=x, T=0]. Common model families: two-model (separate models for treated/control), S-/T-/X-learners, and causal forests or meta-learners that directly target treatment heterogeneity.
Evaluation metrics:
- Uplift curve: plot cumulative incremental response when targeting top-ranked individuals by predicted tau(x). X-axis = proportion targeted; Y-axis = cumulative incremental gains.
- Qini curve & Qini coefficient: the uplift analogue of the ROC/AUC; the Qini curve plots incremental responders vs targeted population; the Qini coefficient is the area between the Qini curve and the random-targeting baseline, higher is better.
- uplift@k (area under the uplift curve, AUUC): incremental gain when treating the top k% (practical for budgeted campaigns).
- These metrics require randomized treatment assignment (or careful causal adjustment) to get an unbiased ground-truth increment; use cross-validation across trials and confidence intervals via bootstrap.
Why prefer uplift over outcome prediction:
Scenario: promotional marketing with a cost to contact. A response model predicts who will buy, but many high-propensity buyers would purchase anyway, so contacting them wastes budget. Uplift modeling identifies persuadable customers (positive tau) and avoids 'do-not-disturb' customers (negative tau, who are harmed by contact). Business impact: higher ROI, lower cost-per-incremental-conversion, and a better customer experience for people who would have converted regardless.
Validating that targeting by uplift actually increases ROI
The naive design (compare a contacted, uplift-targeted group against an untreated control) measures the value of the TREATMENT, not the value of the TARGETING STRATEGY, and will look good even if the uplift model is no better than random targeting, because contacting anyone at all usually beats contacting no one. To isolate the targeting strategy's value, run a three-arm randomized experiment instead:
- Arm 1 (uplift-targeted): contact the top-k% of customers ranked by predicted tau(x).
- Arm 2 (random-targeted, same budget): contact a random k% of customers, holding contact volume and cost identical to Arm 1.
- Arm 3 (no-contact holdout): a small held-out slice used only to estimate the untreated baseline conversion rate needed to validate the tau(x) estimates themselves, not for the ROI comparison.
The ROI comparison that actually answers the question is Arm 1 vs Arm 2: since both arms incur the same contact cost, any difference in conversion or revenue between them isolates the value of picking the right k% via uplift, from the value of contacting people at all. Compute incremental profit per contacted customer for each arm as (conversion_rate * avg_order_value) - cost_per_contact, and test the difference between arms with a two-sample test or bootstrap CI, sized via a power analysis using the historical variance in per-customer revenue. Pre-register the primary metric (incremental profit per contact, Arm 1 minus Arm 2) before launch, and require both statistical significance and a minimum practical ROI delta (covering rollout and maintenance cost) before rolling the uplift model out to full targeting.
Key practical notes:
- Need randomized or well-controlled observational data plus propensity adjustment if the historical treatment assignment was not randomized.
- Common pitfalls: selection bias in who was historically treated, and lack of overlap (some segments never treated historically, so tau(x) is unidentified there without extrapolation).
Unlock Full Question Bank
Get access to all Model Evaluation and Validation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.