Statistical Inference and Hypothesis Testing Questions
Reasoning about uncertainty in data and drawing formal conclusions from samples. Covers probability rules and common distributions, the Central Limit Theorem, sampling, standard error, confidence intervals, and Bayesian reasoning, together with the significance-testing framework: null and alternative hypotheses, p-values, statistical power, Type I and Type II errors, effect sizes, and choosing the right test (t-test, chi-square, non-parametric). Emphasizes correctly interpreting statistical results and avoiding common misreadings of significance in business and product contexts rather than memorizing formulas.
You are reviewing an internal analysis that reports a large effect but only shows results for the significant subgroup analyses. Describe how you would audit the analysis to identify potential p-hacking or selective reporting. List concrete checks you would perform, and propose a robust reanalysis plan to produce defensible inference.
Sample Answer
Direct answer
An analysis that only shows the significant subgroup results, with no mention of how many subgroups were tried, is a textbook signature of p-hacking or selective reporting. Audit it by reconstructing the full set of comparisons that were actually run (not just the ones reported), re-testing that full set with a multiplicity correction, and re-running the analysis end to end on the raw data. The reanalysis plan should pre-specify a small, justified set of primary comparisons and report the complete picture, significant and non-significant alike, not a curated slice of it.
Structured elaboration
Concrete audit checks
| Check | What it reveals |
|---|---|
| Request the original analysis plan and code | Whether the reported subgroup was pre-specified or found by trying many and keeping the winner |
| Reproduce the reported numbers by running the code on the raw data | Whether the figures are even reproducible from the stated pipeline |
| Enumerate every subgroup, covariate combination, and outcome that plausibly could have been tested | The true "family" size the multiplicity correction needs to be computed over |
| Re-test the full family with a correction (Bonferroni or Benjamini-Hochberg) | Whether the reported effect survives once the full search is accounted for |
| Look for p-values clustered just under 0.05 | A classic fingerprint of stopping or specification choices made to cross the threshold |
| Check for undisclosed exclusions or covariate adjustments | Whether cherry-picked exclusions, not a real effect, are driving significance |
Robust reanalysis plan
- Reproduce the original result from raw data and code before doing anything else.
- List every comparison that was actually run, whether or not it appeared in the report.
- Apply a multiplicity correction (Bonferroni for a strict guarantee, Benjamini-Hochberg for a more powered but still principled one) across that full list, not just the reported subset.
- Report the complete table: every tested subgroup, its unadjusted and adjusted p-value, its effect size, and its sample size, so nothing is hidden by omission.
- Clearly separate confirmatory findings (survive correction) from exploratory ones (interesting, but not proven, and worth a dedicated follow-up test).
Worked example
Suppose 12 subgroup comparisons were actually run internally, but the report only surfaces the one with p=0.031. Re-testing the full family of 12 with Holm-Bonferroni (step-down, more powerful than flat Bonferroni but still FWER-controlling, meaning it keeps the family-wise error rate, the chance that even one of the reported findings is a false positive, at or below the target rate across all 12 comparisons, the same guarantee flat Bonferroni gives, just with more power):
| Rank | p-value (sorted) | Holm threshold α/(m−rank+1) | Passes? |
|---|---|---|---|
| 1 (smallest) | 0.031 | 0.00417 | No |
| 2 | 0.090 | 0.00455 | No |
| 3 | 0.120 | 0.00500 | No |
| ... | ... | ... | ... |
(with m=12 and α=0.05; computed directly from the Holm step-down formula)
The smallest p-value in the full family, 0.031, needed to clear 0.00417 to survive Holm-Bonferroni; it doesn't. The "significant" subgroup finding that made it into the report does not survive once the other 11 comparisons that were actually run are accounted for. This is exactly the pattern an audit is designed to catch, and it doesn't require re-collecting any data, only re-testing honestly against the full family.
Trade-offs & pitfalls
- The hardest part of this audit is usually not the statistics, it's getting an honest accounting of how many comparisons were actually tried; without that, no correction can be computed correctly, so insist on the code and logs, not just a verbal assurance.
- Correcting for multiplicity can occasionally kill a genuinely real effect along with the false ones; that's the acknowledged cost of FWER or FDR control, not a reason to skip it, but it's worth flagging any borderline case as "worth a dedicated confirmatory follow-up" rather than dismissing it outright.
- Presenting this audit to the original analysts or to leadership needs care: the goal is fixing the process (pre-registration, full reporting), not accusing anyone of misconduct, since selective reporting this way is frequently an unconscious byproduct of exploratory analysis rather than deliberate manipulation.
- A reanalysis that only tightens the statistics but keeps reporting only the "winning" comparisons repeats the original mistake in a more sophisticated wrapper; the fix has to include reporting the full family, not just applying a correction to the subset that was already selected.
Revenue has increased for two quarters while retention and NPS have declined. Produce a structured analysis plan to reconcile these conflicting signals: the hypotheses you would test, the metrics and cohorts you would analyze, the statistical tests you would run, and the decisions that might follow.
Sample Answer
Direct answer
Revenue up while retention and NPS decline for two quarters is a classic sign that the business is winning short-term monetization at the expense of long-term health, most often through mix shift (different, sometimes lower-fit users or a different pricing structure) rather than the product genuinely improving. The investigation should treat "revenue growth" and "retention/NPS decline" as two outcomes potentially driven by a shared, upstream cause (pricing, acquisition mix, or a product change), not evaluate them independently.
Structured elaboration
Hypotheses to test, roughly in order of how commonly they explain this exact pattern
- Pricing or monetization change increased ARPU but also increased friction or perceived value mismatch, driving both the revenue bump and dissatisfaction.
- Acquisition mix shifted toward higher-paying but lower-fit users (e.g. a paid channel or enterprise push brought in users who convert to revenue quickly but churn or complain more).
- A specific feature or experiment increased short-term monetization while degrading the core experience for a subset of users.
- One-time or concentrated revenue events (large enterprise deals, a promotion) inflate aggregate revenue while the core, ongoing user base's underlying health is actually flat or declining.
- Measurement or attribution issue: revenue recognition change or duplicate counting inflating the topline number without a real behavioral change at all.
Metrics and cohorts to analyze
| Area | What to pull |
|---|---|
| Revenue composition | MRR/ARR vs. one-time bookings, ARPU by cohort, revenue concentration (top-N accounts share of total) |
| Retention | Day 1/7/30 retention and full cohort retention curves, split pre- vs. post-change |
| Satisfaction | NPS trend by cohort and segment, support ticket volume and category, refund rate |
| Acquisition | Channel mix over the two quarters, cost per acquisition, early engagement by channel |
| Engagement | DAU/MAU, core feature usage, session depth, for revenue-contributing vs. non-contributing users |
Statistical tests and analyses to run
- Cohort retention heatmap crossed with ARPU decile: is the revenue growth concentrated in cohorts whose retention is also declining, or in a separate, healthy cohort?
- Difference-in-differences comparing cohorts before and after the suspected driver (e.g. a pricing change), against a comparable unaffected baseline, to isolate its effect on both revenue and retention.
- Regression of retention on covariates including acquisition channel and revenue tier, to test whether channel or tier explains the retention decline once controlled for, versus it being broad-based across the whole user base.
- Trend significance test on the NPS and retention time series themselves (e.g. comparing quarter-over-quarter means with appropriate variance estimates) to confirm the decline isn't within normal seasonal noise before treating it as a real signal worth acting on.
Decisions that might follow
| Finding | Likely action |
|---|---|
| Revenue growth concentrated in a mix shift toward lower-fit users | Reconsider acquisition targeting; the revenue may not be sustainable |
| Pricing change is driving both effects | Evaluate whether the near-term revenue gain is worth the retention cost; consider a more gradual or segmented pricing approach |
| Decline broad-based, not explained by mix or pricing | Deeper product investigation needed; revenue growth may be masking a real product regression |
| Decline concentrated in one specific feature/experiment | Roll back or fix that specific change, independent of the broader revenue trend |
Worked example
Suppose a cohort retention heatmap crossed with ARPU decile shows that the top 20% of ARPU users (by revenue) have 30-day retention of 38%, versus 61% for the bottom 80%. Splitting further by acquisition channel shows the top-ARPU decile is 70% sourced from a single paid channel launched two quarters ago, versus that channel being only 15% of the broader user base. This pattern (a new, concentrated acquisition channel simultaneously driving the ARPU decile up and dragging blended retention down) points toward hypothesis 2: the revenue growth and the retention/NPS decline share a common cause in acquisition mix, rather than the product itself degrading for its existing base. That reframes the fix as a channel-quality and targeting problem, not a product-quality problem, though the team should still verify NPS specifically within that channel's cohort to rule out a genuine experience gap for those users too.
Trade-offs & pitfalls
- Don't average away the story. A blended, company-wide retention number can look like a modest decline while masking a severe drop in one segment offset by stability elsewhere; always cross revenue and retention cuts by the same cohort dimensions.
- Correlation between the revenue and retention trends is not proof of a shared cause. Test the specific hypothesized mechanism (e.g. via difference-in-differences around the suspected driver) rather than asserting a link because two lines moved in opposite directions in the same period.
- NPS is a noisy, low-response-rate metric. A shift in who responds to the survey (not just how they feel) can produce an apparent decline; check response rate and respondent composition alongside the score itself.
- A quarter or two of concentrated enterprise deals is easy to over-interpret as "the business is healthy." Report revenue concentration (e.g. share from the top 10 accounts) alongside the headline number so a one-time deal isn't mistaken for durable growth.
- Acting on hypothesis 1 (pricing) versus hypothesis 2 (acquisition mix) implies very different fixes, so resist moving to remediation before the cohort-level analysis actually distinguishes between them; treating a mix-shift problem as a product or pricing problem wastes a cycle and doesn't fix the underlying issue.
What is p-hacking, and how does it happen in practice? Give an example of how testing many metrics or slicing data into many subgroups until something looks significant can produce a false positive, and explain why this inflates the true false-positive rate above the nominal alpha even when each individual test used alpha = 0.05.
Sample Answer
Direct answer
P-hacking is any practice, deliberate or not, of searching over many possible analyses, metrics, subgroups, cutoffs, models, and selectively reporting the one that crosses a significance threshold. It breaks the guarantee that a p-value under 0.05 means "only a 5% chance of seeing this by chance," because that guarantee was only ever a promise about one pre-specified test, not about the best of many attempts.
Structured elaboration
How it happens in practice
- Testing many metrics (conversion, revenue, retention, time-on-page) and reporting only whichever one turned out significant.
- Slicing the same data into many subgroups (by country, device, cohort, day of week) until one slice looks significant, then presenting that slice as the headline finding.
- Peeking at results repeatedly and stopping the moment p < 0.05 appears, sometimes called optional stopping.
- Trying several reasonable-looking exclusion or data-cleaning rules and keeping the one that produces significance.
None of these require bad intent. A team that genuinely believes it's just being thorough by checking 15 metrics is p-hacking exactly as much as someone doing it cynically; the inflation comes from the number of looks, not the motive.
Why this inflates the true false-positive rate above the nominal alpha
Each individual test, run and interpreted in isolation, does have a 5% false-positive rate under its own null. The problem is the family-wise error rate: the probability that at least one of k independent tests falsely rejects its null, even when every single null is true, is
P(at least one false positive)=1−(1−α)kbecause the probability that all k tests correctly fail to reject is (1−α)k under independence, and "at least one rejects" is the complement of that.
Worked example
With alpha = 0.05 per test:
| Number of tests (k) | P(at least one false positive) |
|---|---|
| 1 | 1−(1−0.05)1=0.0500 |
| 5 | 1−(1−0.05)5=0.2262 |
| 10 | 1−(1−0.05)10=0.4013 |
| 20 | 1−(1−0.05)20=0.6415 |
With just 5 metrics or subgroups tested independently at alpha = 0.05 each, there's already a 22.6% chance that something looks significant purely by chance, not 5%. At 20 tests, for example one significance test per subgroup across 4 countries and 5 device types, the chance of at least one false "win" is over 64%, meaning it's more likely than not that something will look reportable even if nothing real is happening. Reporting only the metric or slice that crossed the threshold, while staying silent about the others that didn't, is p-hacking: it turns a controlled 5% risk into an uncontrolled 22 to 64% risk while still labeling the result "p < 0.05."
Trade-offs and pitfalls
- Pre-register the primary metric and the analysis plan, including which subgroups, if any, will be tested, before looking at data. Anything examined afterward is exploratory, not confirmatory, and should be labeled that way in the report.
- If multiple metrics or subgroups genuinely need testing, control for it: Bonferroni (alpha divided by k per test) is simple but conservative; Benjamini-Hochberg FDR control is the more common choice when running many tests and the goal is to bound the expected proportion of false discoveries rather than the probability of any single false discovery.
- A subtle wrong turn: applying the multiplicity correction only across the metrics that ended up being reported, rather than across everything that was actually looked at. That undercounts k and still overstates significance.
- The fix is not "never look at subgroups." Exploratory subgroup analysis is legitimate and valuable for generating hypotheses; the failure mode is presenting an exploratory finding with the same confidence and framing as a pre-registered confirmatory one.
You're presenting A/B test results to a product manager who asks: what's the difference between a p-value, a confidence interval, and effect size? Explain each concept in plain language, state what each does and does not tell you, and give an example sentence you would use to summarize results to a non-technical stakeholder.
Sample Answer
Direct answer
The p-value tells you whether the observed difference is unlikely to be pure chance under "no effect." The confidence interval (CI) tells you the range of effect sizes the data are consistent with. Effect size tells you how big the difference actually is, in units the business cares about. You need all three together: a tiny p-value with a tiny effect size is not a reason to act, and a wide confidence interval is a warning that the point estimate alone is not precise enough to bet on.
Structured elaboration
P-value
- What it is: the probability of seeing data this extreme (or more) if there were truly no difference between the groups.
- What it tells you: whether "no effect" is a poor explanation for what you observed.
- What it does NOT tell you: the probability the treatment works, or how large the effect is. A p-value of 0.001 and a p-value of 0.04 can come from effects of the same practical size, just with different sample sizes or noise.
Confidence interval
- What it is: a range of effect sizes that are plausible given the data and the model, at a chosen confidence level (typically 95%).
- What it tells you: both the size of the estimated effect and how precisely it's been measured. A narrow interval means the data pin the effect down tightly; a wide one means there's a lot of remaining uncertainty.
- What it does NOT tell you: it is not literally "a 95% probability the true value is in this specific interval." the 95% describes the long-run behavior of the procedure across repeated experiments, not a probability statement about this one realized interval.
Effect size
- What it is: the actual magnitude of the difference, absolute (percentage points) or relative (percent lift).
- What it tells you: whether the change is worth the engineering cost and rollout risk, independent of whether it's statistically significant.
- What it does NOT tell you: on its own, whether the estimate is reliable. An effect size without a confidence interval could be pure noise.
Worked example
A checkout test: baseline click-through p0=0.080, treatment p1=0.086, n=20,000 per arm.
Pooled test statistic:
z=2pˉ(1−pˉ)/np1−p0=2.1748⇒p≈0.029695% CI for the absolute difference (unpooled standard error):
(p1−p0)±1.96np0(1−p0)+np1(1−p1)=[0.0006, 0.0114](all values computed directly from these formulas with scipy.stats.norm)
Summary sentence for the PM: "Click-through rose from 8.0% to 8.6%, a 7.5% relative lift (p = 0.030). We're 95% confident the true absolute lift is somewhere between 0.06 and 1.14 percentage points. It's a real improvement, though the interval is wide enough that the low end is a modest win, not a blockbuster."
Trade-offs & pitfalls
- Presenting only the p-value invites the "significant equals big and certain" misread. Always pair it with the interval and the effect size in the business's own units.
- A p-value just under 0.05 with a confidence interval that barely excludes zero (as in this example, lower bound 0.0006) is a different story than a p-value of 0.0001 with a tight interval far from zero. Treat "significant" as a single bit of information, not the whole picture.
- Wide confidence intervals are common with realistic sample sizes and should be surfaced, not hidden. Narrowing the interval requires either more data or a less noisy metric, not a different way of describing the same data.
You are testing a change in a social feed where treatment may affect not only treated users but their friends (interference). Describe experimental designs appropriate under interference, discuss loss of power versus feasibility trade-offs, and explain how to estimate direct and spillover effects.
Sample Answer
Direct answer
When treatment can spill over to a user's friends, individual-level randomization breaks: control users are contaminated by treated friends, biasing the estimated effect toward zero (or in either direction, depending on the spillover mechanism). The standard fix is to randomize at the level of a cluster that contains most of each user's interactions (cluster or graph-partitioned randomization) so spillover happens mostly within an arm, and to explicitly model direct versus spillover effects with a two-stage design if you need both estimated separately.
Structured elaboration
Why individual randomization fails. The core assumption behind a standard A/B test is SUTVA (the Stable Unit Treatment Value Assumption): one user's outcome depends only on their own treatment assignment, not on anyone else's. Interference (a user's outcome also depends on their friends' assignments) violates SUTVA directly, and the individually-randomized estimator no longer estimates the effect you actually care about.
Design options, in order of how much interference they control:
| Design | How it works | Power impact | When to use |
|---|---|---|---|
| Cluster randomization | Randomize disjoint clusters (geography, community) to treatment/control as whole units | Large loss (effective n shrinks to number of clusters, not users) | Interference is local and clusters can be drawn to contain most interactions |
| Graph-clustered randomization | Partition the social graph (e.g. via Louvain or METIS, graph-partitioning algorithms that automatically group densely-connected nodes together) to minimize edges crossing cluster boundaries, then randomize clusters | Moderate-large loss, better contamination control than arbitrary clusters | A real social graph is available and natural clusters don't exist |
| Two-stage (hierarchical) design | Randomize clusters to a treated-fraction, then randomize individuals within each cluster at that fraction | Moderate loss, but recovers both direct and spillover effects | You need direct and spillover effects estimated separately, not just a combined average |
| Ego/exposure-mapping design | Assign individual treatment probabilities and define exposure as a function of own- and neighbor-treatment (e.g. "treated if ≥k friends treated") | Least power loss of the interference-aware designs | Cluster-level randomization is operationally infeasible at scale |
Estimating direct and spillover effects. With a two-stage design, the direct effect is estimated by comparing treated vs. untreated individuals within the same cluster (same neighborhood exposure, different own-treatment), and the spillover effect is estimated by comparing untreated individuals across clusters that received different treated fractions (same own-treatment status, different neighborhood exposure). Formally, define an exposure mapping ei=f(own treatmenti,fraction of neighbors treatedi) and estimate:
Yi=β0+β1⋅Ti+β2⋅Tˉneighbors(i)+Xi′γ+ϵiwhere β1 approximates the direct effect (holding neighbor exposure fixed) and β2 approximates the spillover effect. Because assignment is graph-dependent, use randomization-based inference (permutation tests over the actual assignment mechanism) or cluster-robust standard errors rather than assuming i.i.d. errors.
Worked example
Consider a network of 100,000 users split into 200 communities of ~500 users each via graph partitioning. A pure individual-level 50/50 randomization would give an effective sample size close to 100,000 for a naive estimator, but that estimator is biased under interference because roughly half of each user's friends are in the opposite arm, contaminating both groups toward the same blended outcome. Switching to cluster randomization on the 200 communities means the effective sample size for detecting the total effect is closer to the number of independent clusters, roughly 200, not 100,000 users. Using the standard sample-size formula with clusters as the unit and a typical intraclass correlation (ICC, ρ: how similar outcomes are for users within the same community relative to users in different communities, 0 = no similarity, 1 = identical within a community) of ρ=0.05 within a community of size m=500, the design effect is:
DEFF=1+(m−1)ρ=1+(500−1)(0.05)=25.95meaning you would need roughly 26x the individually-randomized sample size to retain the same power once clustering is accounted for, which is the concrete cost of eliminating interference bias this way. This is why a two-stage design (randomizing the treated fraction within each cluster rather than the whole cluster) is usually preferred in practice: it keeps most of the individual-level sample size while still separating direct and spillover effects, at the cost of a more complex analysis.
Trade-offs & pitfalls
Cluster and graph-clustered designs trade power for validity, and the design effect above shows that trade can be severe when the intraclass correlation is not small; a design with too few clusters (say, fewer than ~20-30) is functionally underpowered no matter how large each cluster is, because inference is fundamentally happening at the cluster level. A common wrong turn is running a standard individually-randomized A/B test on a clearly-networked feature (social feed, referral, messaging) and reporting the naive difference-in-means as "the effect," when SUTVA violation likely biases it toward zero and understates the true impact. Another pitfall is defining the exposure mapping incorrectly (e.g. assuming linear-in-neighbor-treatment-fraction when the true spillover mechanism is a threshold effect); a misspecified exposure model biases both the direct and spillover estimates even though the design itself was sound.
Unlock Full Question Bank
Get access to all Statistical Inference and Hypothesis Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.