Data Preparation and Class Imbalance for ML Questions
Preparing training data and handling skewed or shifting distributions. Covers preprocessing and cleaning for model input, data augmentation, distribution shift, class imbalance techniques (resampling, reweighting, threshold tuning), and cold-start scenarios where labeled data is scarce. Focuses on the data-side decisions that determine whether a model can learn at all.
Implement a function that yields stratified train/test index arrays for a binary label vector across n folds, ensuring every test fold contains at least one positive instance. Explain how your implementation handles the case where the number of positives is smaller than the number of folds.
You are deciding between an unsupervised anomaly-detection method (for example Isolation Forest) and a supervised classifier trained on imbalanced labels for a rare-event detection problem. What criteria (label availability and quality, novelty of the pattern you are trying to catch, operational constraints) would lead you to favor one approach over the other, and how would you combine both in production?
A stakeholder tells you to 'catch every case' of a rare, costly outcome (fraud, churn, a safety incident) regardless of false positives. How do you explain the precision/recall trade-off to a non-technical audience, translate it into business terms (operational cost, reviewer time), and propose a plan that balances detection with the team's capacity to act on alerts?
You receive inconsistent categorical labels for the same underlying value from different sources (for example 'NY', 'New York', 'N.Y.', 'newyork' all meaning the same place). Propose a scalable standardization approach combining rule-based normalization, fuzzy matching, and human-in-the-loop verification for ambiguous cases, and describe how you would store and version the canonical mapping so it stays consistent across pipelines over time.
Behavioral: describe a time a preprocessing or data-quality decision you made changed the outcome of a model, an experiment, or a business decision. Use the STAR method (Situation, Task, Action, Result): what was your reasoning for the chosen approach, how did you validate its impact, and what did you learn?
Unlock Full Question Bank
Get access to all Data Preparation and Class Imbalance for ML interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.