InterviewStack.io LogoInterviewStack.io

Causal Inference Questions

Establishing cause-and-effect from observational and experimental data. Covers correlation versus causation, confounding, treatment-effect estimation, and quasi-experimental methods such as difference-in-differences, matching, and instrumental variables. Includes incrementality reasoning when true randomization is not possible.

MediumTechnical
63 practiced

What is a natural experiment? Give an example in an e-commerce or advertising context, explain why the assignment mechanism approximates random assignment, and describe the statistical checks you would run to validate that it actually behaves like one.

HardTechnical
64 practiced

Explain the role of influence functions and the sandwich (robust) variance estimator for semiparametric ATE estimators such as IPW and doubly robust estimation. Give an intuitive explanation of how an influence function captures each observation's contribution to the estimator, and why this produces valid standard errors even under fairly weak assumptions about the nuisance models.

HardTechnical
61 practiced

Discuss causal machine learning approaches for heterogeneous treatment effect estimation, such as causal forests and double/debiased machine learning (DML). Explain the identification assumptions, how cross-fitting and orthogonalization (the Neyman-orthogonal score) reduce bias from flexibly estimated nuisance functions, and what to watch for when tuning and interpreting the results.

HardTechnical
60 practiced

Describe an end-to-end pipeline to estimate heterogeneous treatment effects for personalization (who benefits most from a feature or campaign): experiment design, modeling approach (causal forests, uplift meta-learners), feature engineering, how you would validate candidate high-uplift subgroups to guard against overfitting and spurious findings, and how you would translate the output into a prioritized targeting policy with an expected incremental-revenue estimate.

HardTechnical
63 practiced

Explain counterfactual (off-policy) evaluation for a new ranking or recommendation policy using only logged data collected under the old policy: name and describe at least three distinct families of estimators for this problem, and state the main assumption each requires and where it can fail (for example, when the logging policy has no support for an action the new policy would take, or when non-stationarity means the logged data no longer reflects current behavior).

Unlock Full Question Bank

Get access to all Causal Inference interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.