Statistical Inference and Hypothesis Testing Questions
Reasoning about uncertainty in data and drawing formal conclusions from samples. Covers probability rules and common distributions, the Central Limit Theorem, sampling, standard error, confidence intervals, and Bayesian reasoning, together with the significance-testing framework: null and alternative hypotheses, p-values, statistical power, Type I and Type II errors, effect sizes, and choosing the right test (t-test, chi-square, non-parametric). Emphasizes correctly interpreting statistical results and avoiding common misreadings of significance in business and product contexts rather than memorizing formulas.
You need to compare mean customer satisfaction across four geographic regions. Explain why you would use one-way ANOVA instead of multiple pairwise t-tests, how to interpret a significant F-statistic, and which post-hoc methods you would use to identify which regions differ while controlling Type I error.
Sample Answer
Direct answer
Use one-way ANOVA instead of running six pairwise t-tests across four regions because ANOVA tests a single global null hypothesis (all four region means are equal) with one test, controlling the overall false-positive rate; running six separate pairwise t-tests at 0.05 each lets the false-positive rate compound well past 5%. A significant F-statistic means the variance between region means is larger than would be expected from within-region noise alone, but it only tells you some region differs, not which one, so it's followed by a post-hoc test like Tukey's HSD to identify the specific pairs.
Structured elaboration
Why not just run pairwise t-tests. With 4 regions there are (24)=6 possible pairs. Running each pairwise comparison at α=0.05 independently, the chance of at least one false positive across all six, assuming the tests were independent, is:
1−(1−0.05)6≈0.265So even with genuinely equal region means, there's roughly a 26.5% chance of at least one "significant" pairwise result purely by chance, more than five times the nominal 5% rate. ANOVA avoids this by testing one combined hypothesis:
H0:μ1=μ2=μ3=μ4Ha:at least one μi differsInterpreting the F-statistic. F is the ratio of between-group variance to within-group variance:
F=average variance of observations within each groupvariance of the group means around the grand meanA large F means the region means are more spread out relative to each other than the natural noise within each region would predict, and a significant p-value (compared to the F-distribution with the appropriate numerator/denominator degrees of freedom) means it's unlikely this spread arose by chance alone. Critically, rejecting H0 only supports "at least one region differs from at least one other," not a specific pairwise claim, which is exactly why a post-hoc step is needed.
Assumptions to check. Independent observations; approximately normal residuals within each group (or large enough samples per group for the CLT); homogeneity of variance across groups, checkable with Levene's or Bartlett's test. Report an effect size like η2 alongside the F-test, since a significant F with a tiny effect size may not be practically meaningful.
Post-hoc methods. Lead with Tukey's HSD: it's the standard default for all-pairwise comparisons after a significant ANOVA, controlling the family-wise error rate across all pairs simultaneously while providing simultaneous confidence intervals (its Tukey-Kramer variant handles unbalanced group sizes). That's the one answer most interviews are actually testing for. Beyond that baseline, a few named alternatives are worth knowing exist but are depth rather than the first answer: Bonferroni correction (dividing α by the number of comparisons) is simpler but more conservative, costing power; Holm-Bonferroni is uniformly more powerful than plain Bonferroni while still controlling the family-wise error rate; and Benjamini-Hochberg trades some family-wise protection for more power when the concern is false discovery rate across many comparisons rather than strict family-wise control.
Worked example
Satisfaction scores for 5 customers in each of 4 regions:
| Region | Scores | Mean |
|---|---|---|
| 1 | 72, 75, 78, 74, 71 | 74.0 |
| 2 | 80, 83, 79, 85, 81 | 81.6 |
| 3 | 69, 72, 68, 70, 73 | 70.4 |
| 4 | 77, 76, 79, 75, 78 | 77.0 |
One-way ANOVA gives F≈22.38, p≈0.000006 (verified with scipy.stats.f_oneway), strongly rejecting H0: the region means are not all equal. The group means (74.0, 81.6, 70.4, 77.0) suggest Region 2 is notably higher and Region 3 notably lower, but confirming exactly which pairs differ, and by how much, requires the post-hoc step: running Tukey's HSD on this data (rather than six uncorrected pairwise t-tests, which at n=5 per group and this much separation would likely also individually reach significance, but without the family-wise error control that makes the conclusion trustworthy).
Trade-offs & pitfalls
- A significant F only earns you "something differs," not "which pairs differ." Reporting a significant ANOVA as if it identifies a specific region without running the post-hoc step is a common overreach.
- Running naive uncorrected pairwise t-tests after a significant ANOVA "just to see" reintroduces exactly the inflated false-positive problem ANOVA was meant to solve; the whole point of the post-hoc method is that it controls for the multiple comparisons being made.
- If the equal-variance assumption clearly fails (e.g. Levene's test is significant), switch to Welch's ANOVA, which doesn't assume equal variances, rather than proceeding with standard ANOVA and hoping it's robust enough; if normality is a serious concern instead, the Kruskal-Wallis test (the rank-based, nonparametric analogue of one-way ANOVA) is the fallback.
Design a Bayesian A/B testing approach for binary conversion outcomes. Specify suitable priors and likelihood, explain how you would compute posterior probabilities that variant beats control, recommend stopping rules and decision thresholds, and describe how you would present posterior summaries and expected financial impact to stakeholders. Discuss sensitivity to prior choices.
Sample Answer
Direct answer
Model each variant's conversion rate with a Binomial likelihood and a Beta prior, which gives a closed-form Beta posterior for each arm; from there, sample or integrate to get the posterior probability that the variant beats control, the full uplift distribution, and an expected financial impact, rather than a single p-value. Pre-specify stopping rules on those posterior quantities (not on gut feel about when the numbers "look good") so the flexibility of Bayesian monitoring doesn't become an accidental form of optional stopping.
Structured elaboration
Model specification
For arm i∈{C,V} with ni trials and xi successes:
xi∼Binomial(ni,pi),pi∼Beta(αi,βi)Conjugacy gives a closed-form posterior:
pi∣data∼Beta(αi+xi, βi+ni−xi)Prior choice
| Prior | When to use |
|---|---|
| Beta(1,1) (uniform) | No trustworthy historical baseline |
| Beta(2,2) (weakly informative) | Mild shrinkage away from extreme rates with minimal domain input |
| Skeptical, centered on baseline p0 with small effective sample size (α+β≈10) | Historical baseline conversion is trusted and should anchor early reads |
Always run a sensitivity check across at least two of these; report how much the conclusion depends on the choice.
Decision rules and stopping
Pre-specify before looking at live data:
- Efficacy: stop and ship if P(pV>pC)≥0.975 and the expected uplift clears a minimum business-relevant threshold δ.
- Futility: stop with no rollout if P(pV>pC)≤0.05, or P(uplift>δ)≤0.10.
- Maximum horizon: a hard cap on sample size or calendar time regardless of posterior state, to bound operational risk.
Pre-specifying these thresholds (rather than picking a plausible-sounding number after seeing the data) is what keeps continuous Bayesian monitoring honest - the posterior probability itself does not automatically protect against a false-discovery-inflating "peek until it looks good" pattern.
Presenting to stakeholders
Report: posterior mean rate per arm, 95% credible interval on the uplift, P(variant beats control), P(uplift exceeds the business-minimum threshold), and expected financial impact with its own interval. Pair a posterior density plot (control vs. variant) with a single uplift-distribution histogram, since stakeholders generally find "here's the full range of plausible uplift" more actionable than a bare probability.
Worked example
Control: 100 conversions in 1,000 trials. Variant: 128 conversions in 1,000 trials. Prior: Beta(1,1) for both arms.
pC∣data∼Beta(1+100, 1+900)=Beta(101,901) pV∣data∼Beta(1+128, 1+872)=Beta(129,873)import numpy as np
rng = np.random.default_rng(42)
a_c, b_c = 1 + 100, 1 + 1000 - 100 # Beta(101, 901)
a_v, b_v = 1 + 128, 1 + 1000 - 128 # Beta(129, 873)
samples = 200_000
pc = rng.beta(a_c, b_c, size=samples)
pv = rng.beta(a_v, b_v, size=samples)
p_beats = (pv > pc).mean() # 0.97559
delta = pv - pc
delta_mean = delta.mean() # 0.02790 (2.79 points)
ci_lo, ci_hi = np.percentile(delta, [2.5, 97.5]) # (0.00014, 0.05585)
Result: P(pV>pC)=0.976, mean uplift 2.79 percentage points, 95% credible interval [0.01, 5.58] points. To stakeholders: "There's a 97.6% chance the variant converts better than control. The most likely uplift is about 2.8 points, with a plausible range of essentially flat to nearly 5.6 points better; the interval only barely brushes zero, so this comfortably clears a 0.975 efficacy threshold."
Translating that uplift into expected financial impact needs one more input: a value-per-conversion and a traffic volume. Suppose this funnel step sees 50,000 visitors a month and the average order value is $40:
monthly_visitors = 50_000
aov = 40
# reusing the posterior results computed above
delta_mean, ci_lo, ci_hi = 0.02790, 0.00014, 0.05585
extra_conversions_mean = delta_mean * monthly_visitors # 1395.0
extra_conversions_lo = ci_lo * monthly_visitors # 7.0
extra_conversions_hi = ci_hi * monthly_visitors # 2792.5
impact_mean = extra_conversions_mean * aov # 55800.0
impact_lo = extra_conversions_lo * aov # 280.0
impact_hi = extra_conversions_hi * aov # 111700.0
To stakeholders: "The expected financial impact is about $55,800 per month in incremental revenue, with a 95% credible range of roughly $280 to $111,700 per month, mirroring how wide or narrow the uplift interval itself is."
Sensitivity check
Re-running with a skeptical prior centered at the observed control rate (Beta(10,90), effective sample size 100) pulls both posteriors toward 0.10 and changes the conclusion:
import numpy as np
rng = np.random.default_rng(42)
samples = 5_000_000
ac, bc = 10 + 100, 90 + 1000 - 100 # Beta(110, 990)
av, bv = 10 + 128, 90 + 1000 - 128 # Beta(138, 962)
pc = rng.beta(ac, bc, samples)
pv = rng.beta(av, bv, samples)
p_beats = (pv > pc).mean() # 0.9708
Result: P(pV>pC)=0.971 under the skeptical prior, versus 0.976 under the uniform prior. That is below the 0.975 efficacy bar this answer's own decision rule requires before shipping, so the pre-specified stopping rule would not trigger under this prior, even though it does under the uniform one. This is exactly the kind of prior-sensitivity flip the check exists to catch: 1,000 observations per arm is not enough to fully overwhelm a skeptical prior with effective sample size 100 when the true effect sits this close to the threshold, so the honest read is "not yet robust to prior choice," not "unchanged." Flag results like this to stakeholders explicitly, since it means the prior is still doing real work in the conclusion, and either more data or an agreed-upon prior needs to be settled before shipping.
Trade-offs & pitfalls
- Continuous posterior monitoring without pre-specified thresholds reintroduces optional stopping. The posterior probability itself is always "valid" at any instant, but repeatedly checking it and stopping the first time it crosses an ad hoc threshold inflates the effective false-positive rate just as surely as peeking at frequentist p-values does; the fix is pre-committing to the stopping rule, not the Bayesian framework alone.
- Prior sensitivity is the main line of scrutiny a Bayesian result will get. Always show at least an uninformative-vs-skeptical comparison; a conclusion that flips between them is not yet strong enough to act on.
- Point-probability thresholds (P>0.975) ignore magnitude. A tiny, business-irrelevant uplift can still cross a probability threshold with enough traffic; always pair the probability with the uplift-magnitude threshold (P(uplift>δ)), not the probability alone.
- Beta-Binomial assumes independent, stable conversion probability per arm. Delayed conversions, non-stationary traffic mix, or per-user repeat exposure violate the i.i.d. Binomial assumption and require a more elaborate model (e.g., time-to-event or hierarchical structure).
What is p-hacking, and how does it happen in practice? Give an example of how testing many metrics or slicing data into many subgroups until something looks significant can produce a false positive, and explain why this inflates the true false-positive rate above the nominal alpha even when each individual test used alpha = 0.05.
Sample Answer
Direct answer
P-hacking is any practice, deliberate or not, of searching over many possible analyses, metrics, subgroups, cutoffs, models, and selectively reporting the one that crosses a significance threshold. It breaks the guarantee that a p-value under 0.05 means "only a 5% chance of seeing this by chance," because that guarantee was only ever a promise about one pre-specified test, not about the best of many attempts.
Structured elaboration
How it happens in practice
- Testing many metrics (conversion, revenue, retention, time-on-page) and reporting only whichever one turned out significant.
- Slicing the same data into many subgroups (by country, device, cohort, day of week) until one slice looks significant, then presenting that slice as the headline finding.
- Peeking at results repeatedly and stopping the moment p < 0.05 appears, sometimes called optional stopping.
- Trying several reasonable-looking exclusion or data-cleaning rules and keeping the one that produces significance.
None of these require bad intent. A team that genuinely believes it's just being thorough by checking 15 metrics is p-hacking exactly as much as someone doing it cynically; the inflation comes from the number of looks, not the motive.
Why this inflates the true false-positive rate above the nominal alpha
Each individual test, run and interpreted in isolation, does have a 5% false-positive rate under its own null. The problem is the family-wise error rate: the probability that at least one of k independent tests falsely rejects its null, even when every single null is true, is
P(at least one false positive)=1−(1−α)kbecause the probability that all k tests correctly fail to reject is (1−α)k under independence, and "at least one rejects" is the complement of that.
Worked example
With alpha = 0.05 per test:
| Number of tests (k) | P(at least one false positive) |
|---|---|
| 1 | 1−(1−0.05)1=0.0500 |
| 5 | 1−(1−0.05)5=0.2262 |
| 10 | 1−(1−0.05)10=0.4013 |
| 20 | 1−(1−0.05)20=0.6415 |
With just 5 metrics or subgroups tested independently at alpha = 0.05 each, there's already a 22.6% chance that something looks significant purely by chance, not 5%. At 20 tests, for example one significance test per subgroup across 4 countries and 5 device types, the chance of at least one false "win" is over 64%, meaning it's more likely than not that something will look reportable even if nothing real is happening. Reporting only the metric or slice that crossed the threshold, while staying silent about the others that didn't, is p-hacking: it turns a controlled 5% risk into an uncontrolled 22 to 64% risk while still labeling the result "p < 0.05."
Trade-offs and pitfalls
- Pre-register the primary metric and the analysis plan, including which subgroups, if any, will be tested, before looking at data. Anything examined afterward is exploratory, not confirmatory, and should be labeled that way in the report.
- If multiple metrics or subgroups genuinely need testing, control for it: Bonferroni (alpha divided by k per test) is simple but conservative; Benjamini-Hochberg FDR control is the more common choice when running many tests and the goal is to bound the expected proportion of false discoveries rather than the probability of any single false discovery.
- A subtle wrong turn: applying the multiplicity correction only across the metrics that ended up being reported, rather than across everything that was actually looked at. That undercounts k and still overstates significance.
- The fix is not "never look at subgroups." Exploratory subgroup analysis is legitimate and valuable for generating hypotheses; the failure mode is presenting an exploratory finding with the same confidence and framing as a pre-registered confirmatory one.
Design a battery of statistical tests and a workflow to detect disparate impact across protected groups and intersectional subgroups for a binary classifier. Include handling unequal sample sizes, effect size reporting, confidence intervals, and FDR control across many subgroup tests.
Sample Answer
Direct answer
You need a battery, not a single test, because "disparate impact" spans multiple metrics (selection rate, TPR, FPR, PPV) and multiple subgroups of very different sizes, and testing many subgroups at once inflates false positives. The core building blocks are: per-subgroup proportion tests (or Fisher's exact for tiny cells) against a reference group, effect-size reporting (risk difference and Cohen's h) so statistical significance never gets confused with practical significance, confidence intervals sized correctly for unequal N, and Benjamini-Hochberg FDR control across the whole battery of subgroup-metric tests.
Structured elaboration
Start here: for most cells, a two-proportion z-test against the reference group plus Benjamini-Hochberg correction across the battery is enough. The rest below, Fisher's exact, Wilson intervals, Cohen's h, and Benjamini-Yekutieli, are what you reach for once specific cells get small or metrics turn out to be correlated; treat them as depth to mention, not the first tools you pick up.
1. Define the metrics and the comparison structure
- Primary metrics: selection rate / positive prediction rate (demographic parity), true positive rate (equal opportunity), false positive rate, positive predictive value.
- Build one 2x2 table per (subgroup, metric) pair: subgroup vs. reference group.
- Include intersectional cells (e.g. race x gender), but pre-register a minimum cell size N_min (commonly 50) below which a cell is reported descriptively only, not tested. Testing tiny cells either has no power or produces unstable estimates.
2. Pick the test per cell based on size
| Situation | Test | Why |
|---|---|---|
| Both groups large (expected count >= 5 on both sides) | Two-proportion z-test | CLT-justified, standard, fast |
| Either group small or outcome rare | Fisher's exact (2x2) | Exact, no asymptotic assumption |
| CI for a single small-group proportion | Wilson interval, not Wald | A Wilson interval builds its bounds around the estimate in a way that accounts for the estimate's own uncertainty, instead of just adding and subtracting a fixed margin the way a Wald interval does, which keeps it from ever producing an impossible interval (like a negative proportion) for small or extreme-rate groups; Wald undercoverage is severe below n around 30 |
3. Effect size, not just p-value
Cohen's h is the standard effect size for comparing two proportions:
h=2arcsin(psub)−2arcsin(pref)Conventionally, |h| below 0.2 is a small effect, 0.2 to 0.5 medium, above 0.5 large. Report h alongside the risk difference
RD=psub−prefbecause with a large enough N a test can flag a half-percentage-point RD as "significant" even though it is operationally meaningless. Flag on effect size AND significance, never significance alone.
4. Confidence intervals under unequal N
For the risk difference between subgroup and reference, use the unpooled standard error (the pooled SE is only for the significance test's null variance, not for the CI):
RD±z0.975nsubpsub(1−psub)+nrefpref(1−pref)Unequal N shows up directly here: a small subgroup gets a much wider interval for the same point estimate, which is exactly the signal a reviewer needs before escalating a finding from a 50-person cell.
5. Multiplicity: Benjamini-Hochberg across the whole battery
With G subgroups times M metrics you run G x M tests. Controlling each at alpha = 0.05 individually lets the family-wise chance of at least one false flag climb far above 5%. Sort p-values ascending and flag every rank k up to the largest k such that
p(k)≤G×Mkqfor a target FDR q (commonly 0.05). BH controls the expected proportion of false discoveries among flagged subgroups, which is the right target here: you would rather occasionally re-check a false flag than let a real disparity through because of an overly strict Bonferroni bar.
6. Reporting
One row per (subgroup, metric): N, RD, 95% CI, Cohen's h, raw p-value, BH q-value, flagged yes/no. Pair the statistical layer with a predefined practical-significance threshold (for example RD > 0.03 or |h| > 0.2) so a statistically significant but practically trivial row doesn't consume review bandwidth.
Worked example
Reference group: n0 = 8000, x0 = 2400, so p0 = 2400/8000 = 0.3000. Testing 5 subgroups against that reference on the same metric (selection rate):
| Group | n | x | p-hat | z | p-value | RD [95% CI] | Cohen's h |
|---|---|---|---|---|---|---|---|
| G1 | 1200 | 300 | 0.2500 | -3.5470 | 0.00039 | -0.0500 [-0.0765, -0.0235] | -0.1121 |
| G2 | 600 | 210 | 0.3500 | 2.5692 | 0.01019 | 0.0500 [0.0105, 0.0895] | 0.1068 |
| G3 (intersectional) | 150 | 30 | 0.2000 | -2.6526 | 0.00799 | -0.1000 [-0.1648, -0.0352] | -0.2320 |
| G4 | 900 | 234 | 0.2600 | -2.4924 | 0.01269 | -0.0400 [-0.0704, -0.0096] | -0.0891 |
| G5 (intersectional, sparse) | 50 | 10 | 0.2000 | -1.5391 | 0.12377 | -0.1000 [-0.2113, 0.0113] | -0.2320 |
(z-test uses pooled SE for the test statistic and normal approximation for the p-value; RD confidence intervals use the unpooled SE formula above.)
Applying BH at q = 0.05 across these 5 tests: sorted p-values are 0.00039, 0.00799, 0.01019, 0.01269, 0.12377, compared against thresholds (k/5) x 0.05 = 0.01, 0.02, 0.03, 0.04, 0.05. Ranks 1 through 4 clear their thresholds, so G1, G2, G3, and G4 are flagged, with BH q-values 0.00195, 0.01586, 0.01586, and 0.01586 respectively; G5's q-value is 0.12377, so it is not flagged.
The instructive part: G3 and G5 have the identical point effect (RD = -0.10, h = -0.232), but only G3 survives correction. G5's 50-person sample is too small to distinguish that effect from noise; its CI, [-0.2113, 0.0113], still straddles zero. This is why sample-size-aware confidence intervals matter more than the raw effect size when deciding whether to escalate a subgroup finding or flag it as underpowered and needing more data.
Trade-offs and pitfalls
- BH assumes tests are independent or positively dependent (PRDS, positive regression dependency: roughly, the tests are never negatively correlated with each other). If subgroup metrics are heavily correlated (testing PPR and TPR for the same subgroup are not independent), consider Benjamini-Yekutieli, which is more conservative but valid under arbitrary dependence.
- The z-test and Wald-style CI are unreliable below roughly n = 30 or when p is near 0 or 1. Defaulting to Fisher's exact and Wilson intervals for every small cell avoids a class of "false disparity" flags caused purely by asymptotic breakdown.
- A common wrong turn: running every test unadjusted and hand-picking the "interesting" flagged subgroups for a report. That is p-hacking with extra steps. Pre-register the full battery, all metrics times all subgroups down to N_min, before looking at results.
- Unadjusted and covariate-adjusted parity can disagree: a raw PPR gap may shrink or reverse after adjusting for a legitimate business covariate. Report both, and be explicit that "adjusted for X" is itself a judgment call about which covariates are legitimate business factors versus proxies for the protected attribute.
- Intersectional cells lose power fastest as cell counts shrink multiplicatively. Prefer a hierarchical, partial-pooling model across intersectional cells over running dozens of underpowered pairwise tests once N_min filtering starts discarding most of the intersectional signal.
You need to determine a sample size to estimate average customer lifetime value within a margin of error of 0.5 units at 95% confidence. Population standard deviation is unknown but a pilot sample of 40 customers gives sd ≈ 4. Describe the steps to compute a recommended sample size and show the calculation using the pilot sd. Discuss any iterative steps you would take in practice.
Sample Answer
Direct answer
Solve the margin-of-error formula backward for n: given a target margin E, a confidence level (95%, so z∗=1.96), and an estimate of the standard deviation from a pilot sample, the required sample size is n=(z∗s/E)2, rounded up. With the pilot's s≈4 and a target margin of 0.5, that comes out to about 246 customers, treated as a starting plan to be revisited once real data comes in.
Structured elaboration
Setting up the formula. A 95% CI for a mean has half-width (margin of error) E=z∗⋅SE=z∗⋅ns. Solving for n:
n=(Ez∗s)2Using z∗=1.96 for planning (not t) is a standard simplification: at the sample sizes this formula tends to produce, t and z are close enough that using z up front and refining with t afterward is fine.
Why the pilot is only a starting point. The formula needs a standard deviation, but the true population σ for customer lifetime value is unknown; that's exactly what the 40-customer pilot supplies as an estimate. Because it's an estimate from a small sample, it's noisy, and CLV in particular is often right-skewed with a long tail of high-value customers, which can make a 40-person pilot understate the true variance if none of the top-tail customers happened to land in it.
Worked example
Target margin E=0.5, 95% confidence (z∗=1.96), pilot standard deviation s=4 from npilot=40:
n=(0.51.96×4)2=(15.68)2≈245.9⇒n=246(Verified by direct computation.) If historical response/completion rates for this kind of data pull run around 80%, plan to sample or invite more than 246 to end up with 246 completed observations: 246/0.8=307.5⇒308 invites.
Iterative refinement in practice.
- Collect the initial planned batch (or an early tranche of it).
- Recompute the sample standard deviation from the larger, more reliable batch. If it's meaningfully different from the pilot's 4, recompute n with the updated s; a larger true s means the original plan undershoots the target margin.
- Once close to the target n, switch the final margin-of-error check to the t-distribution with df=n−1 for the precise value; at n≈246, t0.975,245≈1.970 versus z=1.96, a difference of about 0.01 in the critical value, small enough that it rarely changes the plan.
- Most product sample-size calculations stop at step 3. Two further refinements apply only in specific cases: if the population of eligible customers is small relative to n (e.g. a niche segment with only a few thousand total customers), apply a finite-population correction (a downward adjustment to n that accounts for sampling a large fraction of a small, finite population rather than an effectively infinite one), which shrinks the required sample size.
- If sampling is clustered (e.g. by region or cohort) rather than a simple random sample, inflate n by a design effect, or DEFF (a multiplier that accounts for the extra correlation between units sampled from the same cluster, which reduces how much independent information the same number of clustered units carries compared to a true simple random sample), to account for the extra correlation clustering introduces.
Trade-offs & pitfalls
- Treating the pilot-based n=246 as final rather than a planning estimate is the main risk: if the true variance is higher than the pilot suggested (common with skewed CLV data and a small pilot), the study will land with a wider-than-intended interval unless the sample size gets revisited.
- Non-response and attrition are easy to forget until data collection is already underway; inflating for expected response rate up front avoids a late scramble to recruit more.
- The margin-of-error formula assumes simple random sampling; ignoring clustering or stratification in the actual collection design while using the unadjusted formula will understate the sample size actually needed.
Unlock Full Question Bank
Get access to all Statistical Inference and Hypothesis Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.