Exploratory Data Analysis and Data Quality Questions

Understanding and preparing an unfamiliar dataset before analysis or modeling. Covers systematic profiling through summary statistics, distribution and outlier inspection, and relationships between variables to form initial hypotheses, alongside turning raw data into a trustworthy base: handling missing values, deduplication, outlier treatment, type and consistency checks, and validation. Includes critical thinking about sampling and measurement bias and what a dataset can and cannot support.

EasyTechnical
70 practiced

You show a dashboard with a sudden jump in conversion rate. A stakeholder immediately assumes a specific recent feature caused it. In a five-minute conversation, what caveats and quick diagnostic checks would you offer, and how would you explain what the data can and can't prove yet?

MediumTechnical
119 practiced

You need to explore a table with hundreds of millions of rows that doesn't fit comfortably in memory or return interactively. What sampling strategy would you use to get representative statistics -- random vs stratified vs reservoir sampling -- and what do you give up with each?

EasyTechnical
60 practiced

A numeric column holds the same value for 95% of rows, with rare non-null values in the remaining 5%. How would you investigate whether to keep, transform, or drop this column, and what would change your answer?

HardTechnical
66 practiced

In practice, how do you actually tell MCAR from MAR rather than just defining them? Walk through a concrete diagnostic (for example, Little's MCAR test, or regressing a missingness indicator on the observed covariates), and be honest about its limitations on a real operational dataset.

MediumTechnical
60 practiced

You notice your dataset systematically undercounts a subset of the population it claims to represent -- events on weekends because a logging job only runs weekdays, or users in rural areas because of how the data was collected. During EDA, what tests and visualizations would reveal this kind of sampling bias, and what would you flag before anyone builds on the data?

Unlock Full Question Bank

Get access to all Exploratory Data Analysis and Data Quality interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.