InterviewStack.io LogoInterviewStack.io

A/B Test Design & Statistical Rigor Questions

Designing and statistically defending a controlled online experiment: framing a testable hypothesis, defining control and treatment variants, choosing the randomization unit, setting the primary success metric, and computing sample size, power, and minimum detectable effect. Covers the statistical foundations that make a readout trustworthy, including hypothesis testing, p-values, confidence intervals, statistical vs practical significance, and Type I/II error. Emphasizes avoiding the common pitfalls that invalidate a test, such as peeking, multiple-comparison inflation, underpowered designs, and how test duration and stopping rules affect the validity of conclusions.

MediumTechnical
44 practiced

Beyond the initial launch experiment, why would you keep a long-run holdout group even after a feature or a pricing algorithm change has fully shipped? Explain how you would decide the size of the holdout, how long to maintain it, and what you are trying to learn from it that the original launch experiment could not tell you. How would you communicate the cost of maintaining a holdout to stakeholders who want the new experience rolled out to everyone?

MediumTechnical
52 practiced

What is the Stable Unit Treatment Value Assumption (SUTVA) in online experimentation? Explain its two components, and give two concrete examples from real online products where SUTVA is violated (for example, a social feed where a treated user's action visibly changes what their connections in control see, or a shared inventory or capacity constraint that lets treatment eat into control's resources). Explain why each violation biases how you would interpret the A/B test result.

MediumTechnical
42 practiced

An experiment shows a statistically significant positive lift on the primary metric, but a guardrail metric moved in the wrong direction, for example a click-through-rate win alongside a retention or revenue-per-user regression. The team wants to ship. Walk through the analysis plan you would run before recommending rollout or rollback: additional robustness checks, whether the guardrail result itself is adequately powered, how you would weigh a short-term win against a longer-term cost, and the decision rule you would apply.

MediumTechnical
40 practiced

A key business metric has high variance and a long-tailed distribution, making it hard to detect real treatment effects without a huge sample. Propose a concrete variance-reduction strategy that combines data transformations with a covariate-based technique such as CUPED or stratification, plus any instrumentation changes needed to support it. Describe the implementation steps, the trade-offs of your approach, and how you would validate the variance reduction actually achieved using historical data.

HardTechnical
44 practiced

You manage a social or messaging product where users influence each other, for example friends can see and react to a new sticker pack or feed feature. A standard user-level A/B test can be biased here because treating one user changes what their connections experience. Propose at least two experimental designs that mitigate this network interference, such as cluster or graph-cluster randomization and ego-network (egocentric) randomization. Specify the randomization unit and exposure mapping for one of them, and describe how you would estimate both the direct effect on treated users and the indirect spillover effect on their connections.

Unlock Full Question Bank

Get access to all 23 A/B Test Design & Statistical Rigor interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.