InterviewStack.io LogoInterviewStack.io

A/B Test Design & Statistical Rigor Questions

Designing and statistically defending a controlled online experiment: framing a testable hypothesis, defining control and treatment variants, choosing the randomization unit, setting the primary success metric, and computing sample size, power, and minimum detectable effect. Covers the statistical foundations that make a readout trustworthy, including hypothesis testing, p-values, confidence intervals, statistical vs practical significance, and Type I/II error. Emphasizes avoiding the common pitfalls that invalidate a test, such as peeking, multiple-comparison inflation, underpowered designs, and how test duration and stopping rules affect the validity of conclusions.

HardTechnical
41 practiced

You run the same experiment across many countries, or across many device types and new-versus-returning users, and see a small but statistically significant uplift overall. Describe how you would assess whether the effect is genuinely heterogeneous across these segments: which interaction tests or models you would use, how you would power the per-segment analysis, and how you would correct for testing many segments at once so you don't just find noise. Compare full pooling, no pooling per segment, and partial pooling using a hierarchical model that borrows strength across segments, and recommend a rollout strategy given what you find.

MediumTechnical
45 practiced

Define the novelty effect and the primacy effect in the context of a multi-week online experiment: what causes each, and in which direction does each bias an early readout? Describe the visualizations, models, or statistical checks you would use to tell a genuine, persistent treatment effect apart from a temporary novelty spike or a fading resistance-to-change effect, and explain how you might adjust the experiment's duration or analysis to account for it.

MediumTechnical
43 practiced

An experiment launched during a holiday week, or right after a new marketing campaign, shows a large lift in week one that decays and flattens out over the following two weeks. Explain how you would distinguish a genuine novelty-effect decay from seasonality, from a selection-bias artifact of the traffic source, and from a real persistent effect. Describe how you would redesign the experiment or its analysis window to reach a trustworthy conclusion.

HardTechnical
51 practiced

You suspect an observed uplift in your A/B test is driven by a novelty effect that will fade over time rather than a persistent treatment effect. Design an experiment and analysis strategy to distinguish the two: specify the time windows you would compare, how you would model the decay, and the decision rule you would use before concluding the effect is real and durable.

HardTechnical
47 practiced

Your A/B test shows no overall lift, but a particular user segment, say mobile users, shows a statistically significant positive uplift. How would you validate whether this is a genuine heterogeneous treatment effect rather than a false positive from looking at many segments? What analyses would you run, and if you're not yet certain, what decision process would you use to decide whether to ship for that segment, run a confirmatory follow-up experiment, or abandon the finding?

Unlock Full Question Bank

Get access to all 23 A/B Test Design & Statistical Rigor interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.