Statistical Inference and Hypothesis Testing Questions
Reasoning about uncertainty in data and drawing formal conclusions from samples. Covers probability rules and common distributions, the Central Limit Theorem, sampling, standard error, confidence intervals, and Bayesian reasoning, together with the significance-testing framework: null and alternative hypotheses, p-values, statistical power, Type I and Type II errors, effect sizes, and choosing the right test (t-test, chi-square, non-parametric). Emphasizes correctly interpreting statistical results and avoiding common misreadings of significance in business and product contexts rather than memorizing formulas.
List five common probability distributions used in analytics and ML. For each, give one concrete use case, state the distribution's parameters, and explain one simple method to estimate those parameters from data. Also explain the relationship between the Poisson distribution for counts and the exponential distribution for waiting times.
Sample Answer
Quick answer
Five distributions that come up constantly in analytics and ML: Normal, Binomial, Poisson, Exponential, and Uniform. Each models a different shape of randomness (continuous symmetric, count of successes, count of rare events, waiting time, equally-likely range), and each has a simple, direct estimator for its parameters from data. Poisson and Exponential are two views of the same underlying process: Poisson counts how many events happen in a fixed interval, Exponential measures the waiting time between them.
The five distributions
| Distribution | Use case | Parameters | Simple estimator |
|---|---|---|---|
| Normal | Daily session length, A/B test metric residuals | mean μ, variance σ2 | Sample mean μ^=xˉ; sample variance σ^2 |
| Binomial | Whether a user converts on a feature (success/failure over n trials) | trials n, success probability p | p^=k/n (observed successes over trials) |
| Poisson | Count of clicks or errors in a fixed time window | rate λ (expected count per interval) | λ^=xˉ (sample mean of the counts) |
| Exponential | Time between failures, or time to first click | rate λ (or mean 1/λ) | λ^=1/xˉ (reciprocal of the mean inter-arrival time) |
| Uniform | Null model for randomized assignment or a simulated baseline | bounds a,b | a^=min(x), b^=max(x) |
Poisson and Exponential: two views of one process
If events (clicks, errors) occur independently at a constant average rate λ per unit time, the count of events in a fixed interval follows Poisson(λ), and the waiting time between consecutive events follows Exponential(λ). They describe the same process from two angles: "how many in this window" versus "how long until the next one." Both are memoryless in the sense that the process has no built-in trend, though the specific memoryless property (the wait for the next event doesn't depend on how long you've already waited) belongs to the Exponential distribution specifically.
Worked example: matching estimators to simulated data
Simulating each process with a pinned seed to confirm the estimators recover the true parameters:
import numpy as np
rng = np.random.default_rng(seed=52)
# Poisson: errors per minute, true rate lambda=2.4
counts = rng.poisson(lam=2.4, size=200)
lam_hat = counts.mean()
print(f"Poisson: lambda_hat={lam_hat:.3f} (true=2.4)")
# Poisson: lambda_hat=2.305 (true=2.4)
# Exponential: waiting time between the same events, same rate
interarrivals = rng.exponential(scale=1 / 2.4, size=200)
lambda_hat_exp = 1 / interarrivals.mean()
print(f"Exponential: lambda_hat={lambda_hat_exp:.3f} (true=2.4)")
# Exponential: lambda_hat=2.509 (true=2.4)
# Binomial (Bernoulli trials): conversion, true p=0.18
clicks = rng.binomial(n=1, p=0.18, size=500)
p_hat = clicks.mean()
print(f"Binomial: p_hat={p_hat:.4f} (true=0.18)")
# Binomial: p_hat=0.1860 (true=0.18)
All three estimators recover parameters close to the true simulated values, with the expected sampling noise given the sample sizes used.
Trade-offs and pitfalls
- The Normal approximation to a proportion (Binomial) breaks down near 0 or 1, or with small n; check np^ and n(1−p^) are both comfortably above about 5-10 before relying on it for a confidence interval.
- Poisson assumes the rate λ is constant and events are independent. Bursty or overdispersed count data (variance noticeably exceeding the mean) is a signal to consider a Negative Binomial instead.
- The Uniform min/max estimator is biased and sensitive to outliers; a single extreme observation redefines the estimated bound entirely, unlike the mean-based estimators above.
You are modeling clicks in a newsletter product. Each user receives 10 independent emails and each email has probability p = 0.3 of being clicked. (a) What is the probability that a user clicks at least 3 emails? (b) Compute the expected number of clicks and the variance. Show formulas and numeric answers.
Sample Answer
Quick answer
Each user receiving 10 independent emails, each clicked with probability p=0.3, is a Binomial(n=10,p=0.3) model. The probability of at least 3 clicks is easiest to get via the complement of 0, 1, or 2 clicks, which comes out to about 61.7%. The expected number of clicks is np=3, with variance np(1−p)=2.1.
Setup
X∼Binomial(n=10, p=0.3), with PMF:
P(X=k)=(k10)(0.3)k(0.7)10−k(a) P(X≥3)
Computing the complement is less arithmetic than summing 8 terms directly:
P(X≥3)=1−[P(X=0)+P(X=1)+P(X=2)]from math import comb
n, p = 10, 0.3
p0 = comb(n, 0) * p**0 * (1 - p)**10
p1 = comb(n, 1) * p**1 * (1 - p)**9
p2 = comb(n, 2) * p**2 * (1 - p)**8
p_ge3 = 1 - (p0 + p1 + p2)
print(f"P(X=0)={p0:.6f}, P(X=1)={p1:.6f}, P(X=2)={p2:.6f}")
print(f"P(X>=3)={p_ge3:.4f}")
# P(X=0)=0.028248, P(X=1)=0.121061, P(X=2)=0.233474
# P(X>=3)=0.6172
P(X≥3)≈0.6172: roughly 61.7% of users click at least 3 of the 10 emails.
(b) Expected value and variance
For a Binomial random variable:
E[X]=np,Var(X)=np(1−p) E[X]=10×0.3=3,Var(X)=10×0.3×0.7=2.1Standard deviation =2.1≈1.449.
Interpretation
An average user clicks 3 of the 10 emails, with a standard deviation of about 1.45 clicks, so a typical user's click count lands somewhere between about 1.5 and 4.5. The fact that over 60% of users clear the "at least 3" bar despite the average also being 3 reflects the right-skew-free, moderately spread-out shape of this particular Binomial (n=10, p=0.3 isn't close enough to either 0 or 1 to be heavily skewed).
Trade-offs and pitfalls
- Independence across the 10 emails is the load-bearing assumption. If a user who clicks one email is more likely to click the next (engagement momentum) or less likely (fatigue), the true distribution is overdispersed or underdispersed relative to Binomial, and the variance formula above will be wrong.
- p=0.3 is itself an estimate, not a known constant. In practice p is estimated from historical click data, and the Binomial calculation above doesn't account for uncertainty in that estimate; a more careful treatment would put a Beta prior on p and integrate over it (a Beta-Binomial), widening the resulting probability estimates.
Design a Monte Carlo simulation to test whether an observed 12% monthly metric drop could be due to random fluctuations across segments. Specify assumptions about the data-generating process (per-segment means and variances, sample sizes), number of simulations, the test statistic you would use, and how to interpret the simulation results (p-value or empirical percentile).
Sample Answer
Direct answer
Simulate the segment-level sampling distribution you believe is true under "nothing changed" (the null), aggregate each simulated draw the same way the real 12% drop was computed, and see how extreme the real observed drop is relative to that simulated null distribution - that fraction of simulations as extreme or more extreme than what was observed is the Monte Carlo p-value. This sidesteps needing a closed-form formula for the sampling distribution of a weighted, multi-segment aggregate change.
Structured elaboration
Assumptions to specify up front (the part reviewers should scrutinize hardest). Per-segment last-month means μi and standard deviations σi, and sample sizes ni, typically taken as the observed last-month values; a distributional assumption for how each segment's monthly mean varies under the null (Normal via CLT is standard when ni is reasonably large); and a choice of aggregation weights wi (usually segment size, so the simulated aggregate mirrors how the real metric was computed).
Test statistic. The same relative-change statistic used on the real data, so the simulated null distribution is comparable:
T=yˉlastyˉcurr−yˉlast,yˉ=∑iwi∑iwiyˉiSimulation procedure. For each of N replicates: draw a simulated "last month" and a simulated "this month" segment mean for every segment from Normal(μi,σi2/ni) (same μi both times, since the null says nothing changed); aggregate both with the real weights; compute T for that replicate. The empirical p-value is the fraction of simulated T values at least as extreme as the observed Tobs.
Interpreting the result. Report the empirical p-value alongside the simulated null's spread (e.g. its 2.5th/97.5th percentile), not just a pass/fail - a p-value of exactly 0 out of N simulations should be reported as "p<1/N", never as p=0, since Monte Carlo resolution is bounded by the number of replicates run.
Worked example
Four segments with given last-month means, standard deviations, and sample sizes; a 12% weighted-aggregate drop is observed this month. Simulating N=50,000 replicates under the null (pinned seed 2026):
import numpy as np
mu = np.array([40.0, 25.0, 60.0, 15.0])
sigma = np.array([12.0, 8.0, 20.0, 6.0])
n = np.array([5000, 8000, 2000, 12000])
w = n.astype(float)
se = sigma / np.sqrt(n)
last_month_weighted_mean = np.sum(w*mu)/np.sum(w)
T_obs = -0.12 # the observed 12% drop
rng = np.random.default_rng(2026)
N = 50000
T_sim = np.empty(N)
for t in range(N):
last = rng.normal(mu, se)
curr = rng.normal(mu, se) # null: same mu both months
agg_last = np.sum(w*last)/np.sum(w)
agg_curr = np.sum(w*curr)/np.sum(w)
T_sim[t] = (agg_curr - agg_last) / agg_last
p_value = np.mean(T_sim <= T_obs)
lo, hi = np.percentile(T_sim, [2.5, 97.5])
print(f"last-month weighted mean = {last_month_weighted_mean:.3f}")
print(f"observed T_obs = {T_obs*100:.2f}%")
print(f"simulated null distribution (N={N:,}): 95% range = [{lo*100:.2f}%, {hi*100:+.2f}%]")
if p_value == 0.0:
print(f"empirical one-sided p-value = {p_value:.6f} -> report as p < 1/{N:,} = {1/N:.5f}")
else:
print(f"empirical one-sided p-value = {p_value:.6f}")
Output:
last-month weighted mean = 25.926
observed T_obs = -12.00%
simulated null distribution (N=50,000): 95% range = [-0.61%, +0.62%]
empirical one-sided p-value = 0.000000 -> report as p < 1/50,000 = 0.00002
A 12% drop is roughly 20x larger than the edges of the null distribution's 95% range (about ±0.6%), so none of the 50,000 null simulations produced anything close to it - strong evidence the drop is not sampling noise given these segment-level variance assumptions.
Trade-offs & pitfalls
The Monte Carlo p-value is only as trustworthy as the assumed per-segment variances and the Normal-approximation-to-the-mean assumption; if the underlying per-user metric is heavy-tailed or the segments are small, drawing from Normal(μi,σi2/ni) understates tail risk, and a nonparametric bootstrap from raw per-user data (when available) or a Poisson/binomial-appropriate simulation for count/rate metrics is more defensible. N also has to be large enough that the reported p-value's own resolution doesn't misleadingly suggest more precision than it has - a p-value reported as exactly 0 from only 1,000 simulations is really "p<0.001," not "p=0." Finally, this test as specified ignores month-to-month correlation across segments (a market-wide shock hitting several segments together); if that's plausible, the segment draws should be simulated from a joint distribution with an estimated covariance rather than independently.
You plan a two-sided A/B test comparing conversion proportions. Baseline p0 = 0.05 and you expect a 20% relative uplift (p1 = 0.06). Using alpha=0.05 and desired power 0.8, compute the required sample size per group. Show the formula you use, numeric steps, and discuss how the calculation changes for unequal allocation or continuous metrics.
Sample Answer
Direct answer
For a two-sided test comparing two proportions with baseline p0=0.05, target p1=0.06 (20% relative lift), α=0.05, and power =0.80, the required sample size is about 8,158 per arm, computed from the standard normal-approximation formula for a two-proportion test. Unequal allocation and continuous metrics use the same underlying logic (compare the standardized difference against critical values from the desired error rates) but change the variance term and, for unequal group sizes, the optimal split between arms.
Structured elaboration
Formula (equal allocation, two-sided)
nper group=(p1−p0)2[z1−α/22pˉ(1−pˉ)+z1−βp0(1−p0)+p1(1−p1)]2where pˉ=(p0+p1)/2, z1−α/2=1.9600 for α=0.05, and z1−β=0.8416 for power =0.80.
Numeric steps
p0=0.05, p1=0.06, Δ=0.01, pˉ=0.055 2pˉ(1−pˉ)=2(0.055)(0.945)=0.32241 p0(1−p0)+p1(1−p1)=0.0475+0.0564=0.32234 numerator=1.95996(0.32241)+0.84162(0.32234)=0.90320 n=0.0120.903202=0.00010.81577=8,158(all steps reproduce exactly with plain arithmetic; verified with python3)
So you need about 8,158 users per arm, 16,320 total.
Unequal allocation
With allocation ratio r=n1/n0 (treatment vs. control), the sample size for the control arm becomes:
n0=Δ2[z1−α/2pˉ(1−pˉ)(1+1/r)+z1−βp0(1−p0)+p1(1−p1)/r]2,n1=r⋅n0with pˉ=(p0+rp1)/(1+r). The allocation ratio that minimizes the total sample size n0+n1 for a fixed variance target (Neyman allocation) is:
r∗=p0(1−p0)p1(1−p1)For this example, p0(1−p0)=0.0475 and p1(1−p1)=0.0564, so r∗=0.0564/0.0475=1.090, close to 1: with proportions this close together, equal allocation is already near-optimal, and there's little total-sample-size benefit to skewing the split. Unequal allocation becomes worth doing when one arm is deliberately capped (e.g. a risky treatment held to 10% of traffic); the cost is a larger total sample size than the balanced design would need for the same power.
Continuous metrics
Replace the proportion-variance terms with the outcome variance σ2 (estimated from historical data or a pilot):
nper group=Δ22(z1−α/2+z1−β)2σ2where Δ is the absolute mean difference you want to detect. Use a pooled or historical estimate of σ, and consider winsorizing or a log transform first if the metric is heavy-tailed, since σ2 from raw heavy-tailed data can be dominated by a handful of extreme values and understate how many "typical" observations you'll actually need.
Trade-offs & pitfalls
- For small p0 or small Δ, the normal approximation underlying this formula can be inaccurate; check that np0 and n(1−p0) are both comfortably above about 5-10, or fall back to an exact binomial/simulation-based power calculation.
- The unequal-allocation Neyman ratio only minimizes total sample size; it does not account for a fixed traffic budget per arm or for practical constraints like a hard cap on treatment exposure, so it should be treated as a starting point, not an automatic answer.
- Continuous-metric sample sizes are only as good as the variance estimate feeding them; an outdated or unrepresentative σ estimate (from before a product change, or from a different user segment) will make the planned sample size wrong in either direction.
- None of these formulas account for multiple comparisons, expected attrition, or metric seasonality; pad the computed n for real-world dropout and plan the run to cover at least one full weekly cycle regardless of what the raw sample-size number implies about duration.
Consider IID Bernoulli trials X1,...,Xn with unknown success probability p. Derive the maximum likelihood estimator (MLE) for p, show whether it is unbiased, and compute its variance and standard error formula. Explain how to form a normal-approximation 95% CI for p and mention limitations of that CI for small n or p near 0 or 1.
Sample Answer
Direct answer
For IID Bernoulli trials, the maximum likelihood estimator of p is simply the sample mean, p^=Xˉ=n1∑Xi. It's unbiased, has variance p(1−p)/n, and the standard normal-approximation 95% CI is p^±1.96p^(1−p^)/n. That CI can behave badly (even producing bounds outside [0,1]) when n is small or p is near 0 or 1, because the normal approximation to a discrete, bounded variable breaks down exactly in that regime.
Structured elaboration
Deriving the MLE
The likelihood for n IID Bernoulli(p) draws is:
L(p)=i=1∏npXi(1−p)1−XiTaking logs:
ℓ(p)=(i=1∑nXi)logp+(n−i=1∑nXi)log(1−p)Differentiating with respect to p and setting to zero:
dpdℓ=p∑Xi−1−pn−∑Xi=0Solving:
p^MLE=n1i=1∑nXi=XˉUnbiasedness
E[p^]=n1i∑E[Xi]=n1(np)=pso p^ is unbiased for every finite n.
Variance and standard error
Since the Xi are IID with Var(Xi)=p(1−p):
Var(p^)=n21i∑Var(Xi)=np(1−p) SE(p^)=np^(1−p^)(plugging in the estimate)Normal-approximation 95% CI
By the CLT, p^ is approximately N(p,p(1−p)/n) for large n, giving the Wald interval:
p^±1.96⋅np^(1−p^)Limitations
- For small n, the discrete binomial distribution is poorly approximated by a continuous normal.
- For p near 0 or 1, the sampling distribution of p^ is skewed (bounded at 0 or 1), so a symmetric normal interval can overshoot the boundary, producing a lower bound below 0 or an upper bound above 1, which is nonsensical for a probability.
- Coverage of the nominal 95% Wald interval is often noticeably below 95% in these regimes ("the interval doesn't actually cover the truth 95% of the time").
- Better alternatives: the exact Clopper-Pearson interval (guaranteed coverage, can be conservative), the Wilson score interval (better calibrated, doesn't leave [0,1]), or Agresti-Coull (a simple, well-behaved adjustment).
Worked example
Two pinned-seed simulations to show both the well-behaved and the breakdown regime:
import numpy as np
rng = np.random.default_rng(seed=11)
# Well-behaved: n=50, p_true=0.3
n = 50
x = rng.binomial(1, 0.3, n)
p_hat = x.mean() # 0.22 (11 successes / 50)
se = np.sqrt(p_hat*(1-p_hat)/n) # 0.0586
ci = (p_hat - 1.96*se, p_hat + 1.96*se) # (0.105, 0.335)
# Breakdown regime: small n, p near 0
x2 = np.array([0,0,0,0,0,0,0,0,0,1]) # p_hat = 0.1, n=10
p_hat2 = x2.mean()
se2 = np.sqrt(p_hat2*(1-p_hat2)/10) # 0.0949
ci2 = (p_hat2 - 1.96*se2, p_hat2 + 1.96*se2) # (-0.086, 0.286)
In the well-behaved case (n=50, true p=0.3), the estimate is p^=0.22 with a 95% Wald CI of (0.105,0.335), a sensible interval fully inside [0,1]. In the breakdown case (n=10, p^=0.1, one success out of ten), the Wald CI comes out to (−0.086,0.286), a lower bound that is a negative probability, which is exactly the failure mode the limitations section warns about. This is a mechanical consequence of the formula, not a coding bug: the normal approximation simply isn't valid there.
Trade-offs & pitfalls
- The MLE being unbiased doesn't mean the Wald CI built from it is well-calibrated. These are separate properties; don't conflate "unbiased point estimate" with "trustworthy interval."
- Reporting a Wald CI that includes values outside [0,1] without flagging it is a common interview red flag. A senior answer catches this and names the fix (Wilson, Clopper-Pearson, or bootstrap) rather than just reporting the number.
- Clopper-Pearson guarantees coverage but is conservative (intervals wider than necessary), which matters if you're using CI width to size a downstream decision.
- For a hypothesis test of p=p0 (rather than just a CI), you'd typically use the same test statistic, but note that the Wald-based test and the score-based test (which doesn't plug in p^ under the null) can disagree in exactly the same small-n/extreme-p regime.
Unlock Full Question Bank
Get access to all Statistical Inference and Hypothesis Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.