Causal Inference Questions
Establishing cause-and-effect from observational and experimental data. Covers correlation versus causation, confounding, treatment-effect estimation, and quasi-experimental methods such as difference-in-differences, matching, and instrumental variables. Includes incrementality reasoning when true randomization is not possible.
Explain uplift (heterogeneous treatment effect) modeling: when it is preferable to relying on the average effect from a randomized experiment, how it differs from a standard predictive model, the common families of modeling approaches and how they differ, at least one evaluation metric appropriate for ranking users by predicted uplift, and the practical deployment considerations, including why it needs randomized training data.
Randomization is not available for a marketing or product change you need to evaluate causally. Compare instrumental variables, regression discontinuity, difference-in-differences, and synthetic control as identification strategies: for each, give a concrete scenario where it is the right tool, the key assumption you would need to validate, and one diagnostic or falsification check you would run before trusting the result.
A report shows that users who enable personalization have 30% higher retention. List at least five plausible confounders that could explain this correlation on their own, and briefly explain how each would bias a naive interpretation that personalization causes the retention lift.
In Python, simulate a dataset where X and Y are correlated only because of a shared confounder Z: generate Z ~ N(0,1), X = 0.8Z + noise, Y = 1.2Z + noise (independent normal noise terms). Compute (a) the naive Pearson correlation between X and Y, and (b) the partial correlation between X and Y controlling for Z (by regressing each out of Z and correlating the residuals). Show that (b) is close to zero while (a) is not, and state one limitation of partial correlation as a method for controlling confounders in general.
You're prioritizing analytics investment: should you spend engineering effort enabling randomized rollouts across regions, or spend it collecting additional covariates to support better observational analyses? Lay out a brief decision framework considering cost, time-to-insight, risk, and your actual ability to get a defensible causal answer either way.
Unlock Full Question Bank
Get access to all 41 Causal Inference interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.