Advanced Experimentation Designs Questions
Experimentation techniques beyond the simple two-arm A/B test: causal inference and quasi-experiments, sequential and always-valid testing, multi-armed bandits, factorial and multivariate designs, and handling interference or network effects. Covers when randomized experiments are infeasible and difference-in-differences, instrumental variables, or switchback designs apply. The concept scope is methodology selection for hard experimental situations.
Explain aliasing in fractional factorial designs and what design resolution (III, IV, V) means. Using a 2^(3-1) half-fraction with generator C = AB, derive the full alias structure, explain concretely why main effect A ends up aliased with interaction BC, and describe how you would run a fold-over follow-up experiment to de-alias A from BC.
Sample Answer
Direct answer: Aliasing happens when a fractional factorial design cannot separate the influence of two (or more) different effects because they change in exactly the same pattern across the runs that were actually included in the reduced design; the two effects are then statistically indistinguishable from each other using that data alone. Design resolution describes how BAD the confounding is: Resolution III means main effects are aliased with two-factor interactions (2fi's); Resolution IV means main effects are clear of 2fi's but 2fi's are aliased with each other; Resolution V means main effects and 2fi's are all clear of each other (only 2fi's alias with three-factor interactions, which are usually negligible). Higher resolution costs more runs for the same number of factors.
Deriving the alias structure for the 2^(3-1) half-fraction, generator C = AB:
The half-fraction design matrix (executed):
| Run | A | B | C=A·B |
|---|---|---|---|
| 1 | -1 | -1 | +1 |
| 2 | -1 | +1 | -1 |
| 3 | +1 | -1 | -1 |
| 4 | +1 | +1 | +1 |
Since C = A·B (the defining generator), multiplying both sides by C gives C⋅C=A⋅B⋅C, and since C2=I (the identity column, all +1), the DEFINING RELATION is I=ABC.
To find what any effect is aliased with, multiply it by the defining relation and simplify using X2=I:
- A aliased with A⋅ABC=A2BC=BC
- B aliased with B⋅ABC=AB2C=AC
- C aliased with C⋅ABC=ABC2=AB
So main effect A is aliased with interaction BC precisely BECAUSE the defining relation is I=ABC: the "A" column in the half-fraction and the "BC" column (computed as B times C) are numerically identical (both equal +1,-1,-1,+1 down the four runs), so no linear model fit to this data alone can tell whether an observed effect on that column comes from A, from BC, or some mixture of both. This is a Resolution III design (shortest word in the defining relation, "ABC," has length 3).
De-aliasing via fold-over (executed simulation, single-factor construction): run a second block, a "fold-over on A": flip the SIGN of column A only, keep the B and C run-values exactly as they were in block 1 (do not re-derive C from the flipped A):
| Run | A | B | C |
|---|---|---|---|
| 1 | +1 | -1 | +1 |
| 2 | +1 | +1 | -1 |
| 3 | -1 | -1 | -1 |
| 4 | -1 | +1 | +1 |
Checking the ABC column on this second block gives -1 on every run, so its defining relation is I=−ABC (flipping one letter's sign in an odd-length word flips the word's overall sign). That sign flip is the whole trick: in block 1, A is aliased with +BC (its contrast estimates effect(A)+effect(BC)); in the fold-over block, A is aliased with −BC (its contrast estimates effect(A)−effect(BC)). Simulating a true response with A's true effect = 5 and BC's true effect = 3 (noise SD 0.5, averaged over 5000 replications): block 1's A-contrast converges to 8.00 (matching 5+3), and the fold-over block's A-contrast converges to 2.00 (matching 5−3). AVERAGING the two blocks' A-contrast estimates, (8.00+2.00)/2=5.00, isolates A alone (true A = 5, exact match); the HALF-DIFFERENCE, (8.00−2.00)/2=3.00, isolates BC alone (true BC = 3, exact match). The general mechanism: fold over on the ONE factor you need de-aliased (flip only its column, keep the others fixed), which flips that factor's alias sign but not the alias's own true effect, then the sum and difference of the two blocks' contrast estimates separate the two confounded effects cleanly.
You must test five binary product features but traffic constraints only allow 8 experimental arms. Propose a 2^(5-2) fractional factorial design, derive its full alias structure, and state which main effects are safe to interpret cleanly versus which are confounded with a two-factor interaction you should treat with caution. Then generalize your recommendation to a constrained-traffic landing-page test with 3 headline options and 2 hero-image variants where you must keep interpretability high, and briefly compare this fractional-factorial approach against running full factorial or sequential targeted experiments instead.
Sample Answer
Direct answer: With 5 binary features and only 8 arms available (2^3), use a quarter-fraction 2^(5-2) design with generators D=AB and E=AC. This yields the standard minimum-aberration Resolution III design for this factor count. A is relatively safe to interpret (its only alias is the four-letter word ABCDE, negligible in practice); D and E carry real risk since each is 1:1 aliased with a plausible two-factor interaction (AB and AC respectively).
Constructing the design (executed): start from the full 2^3 factorial on A, B, C (8 runs), then set D = A·B and E = A·C:
| Run | A | B | C | D=AB | E=AC |
|---|---|---|---|---|---|
| 1 | -1 | -1 | -1 | +1 | +1 |
| 2 | -1 | -1 | +1 | +1 | -1 |
| 3 | -1 | +1 | -1 | -1 | +1 |
| 4 | -1 | +1 | +1 | -1 | -1 |
| 5 | +1 | -1 | -1 | -1 | -1 |
| 6 | +1 | -1 | +1 | -1 | +1 |
| 7 | +1 | +1 | -1 | +1 | -1 |
| 8 | +1 | +1 | +1 | +1 | +1 |
Full alias structure (derived from the defining relations I=ABD=ACE=BCDE, the third relation is the product of the first two, word lengths 3, 3, 4, so this is Resolution III):
- A: aliased with BD (via ABD), CE (via ACE), and the weak ABCDE (via BCDE)
- B: aliased with AD, ABCE, CDE
- C: aliased with ABCD, AE, BDE
- D: aliased with AB, ACDE, BCE
- E: aliased with ABDE, AC, BCD
Interpretation guidance: A, B, and C are each cleanly separated from any OTHER main effect and only alias with 2-factor interactions (BD/CE for A, AD for B, AE for C respectively) or weaker higher-order words; if the business has no strong prior that those specific pairwise interactions are large, A/B/C main effects can be reported with reasonable confidence. D is aliased 1:1 with AB and E with AC, both entirely plausible 2-factor interactions a stakeholder might expect (e.g. if D and E were chosen as the two "riskiest" features, expect their main-effect estimates might really be picking up an AB or AC interaction): flag D and E's point estimates as "confounded, needs a confirmatory follow-up" rather than reporting them at face value.
Generalizing to the 3-headline/2-hero-image, high-interpretability landing-page case: this is a 3x2 mixed-level design (6 cells if run as a full factorial, which is small enough to run FULL rather than fractional; fractional designs earn their keep specifically when the FULL factorial's cell count would outstrip traffic, e.g. 4+ binary factors). Since interpretability is explicitly required here and 6 cells is affordable, recommend the FULL factorial over any fraction: fractional factorial trades interpretability for run-count savings, and this scenario does not need that trade.
Comparing full factorial, fractional factorial, and sequential targeted experiments (per aa4819c0's ask, when total traffic is limited across multiple segments): full factorial gives the cleanest, alias-free estimates but its cell count grows exponentially with factor count, quickly becoming infeasible; fractional factorial trades some interpretability (confounded higher-order effects) for a linear run-count, and is the right default once a full factorial would need more cells than traffic supports; sequential targeted experiments (test the highest-uncertainty factor first, then the next) avoid aliasing entirely but pay the SPEED cost quantified in S0 (roughly half the power per unit time, or double the calendar time, versus a concurrent design at the same total budget) and cannot detect interactions between factors tested in different sequential waves. Recommendation: default to full factorial while cells stay affordable (roughly up to 3-4 binary factors depending on traffic), switch to a minimum-aberration fractional design once traffic-per-cell would otherwise drop below your power floor, and reserve pure sequential testing for genuinely independent, low-interaction-risk factors or truly traffic-starved situations.
Describe how to design and analyze a switchback (crossover) experiment where a marketplace or product alternates between control and treatment on a time-block schedule, for example weekly. Explain how you would randomize the block-assignment schedule, choose a washout period and detect carryover effects, control for time trends and seasonality in the analysis, and determine how many blocks you need to reliably detect a treatment effect. Give at least two concrete cases where a crossover/switchback design would be inappropriate.
Sample Answer
Direct answer: A switchback experiment assigns TIME BLOCKS (for example, whole weeks or days for an entire market or region), not individual users, to control or treatment, alternating on a schedule; it is the standard fix when individual-level randomization is invalid because treated and control units share a common resource pool (drivers, inventory, ad auction) that lets treatment "leak" into control, the SUTVA-violation problem this topic's sibling a-b-test-design-and-statistical-rigor closes for USER-level interference but does not cover for TIME-block designs.
Randomizing the schedule: do not simply alternate control-treatment-control-treatment on a fixed clock; randomize the block ORDER (e.g. randomly assign each week to control or treatment subject to a balance constraint, roughly half of each) so that any systematic weekly pattern (day-of-week mix, holiday timing) is not perfectly confounded with the treatment schedule, and randomize the STARTING phase across different markets/regions if running the switchback in parallel across several geographies, so a market-specific event does not land on the same treatment/control alignment everywhere.
Washout period and carryover: a washout period is a buffer at the START of each new block (for example, the first hour of a new week) during which data is DISCARDED from analysis, to let effects from the PRIOR block's treatment (driver repositioning, cached recommendations, habituation) dissipate before the new block's measurement window begins. Choose the washout length from domain knowledge of how fast the system's state resets (e.g. driver-supply rebalancing might take 30-60 minutes in a ride-hailing market, materially shorter than the week-long block itself). To DETECT carryover statistically: fit a model with a lagged treatment indicator (whether the PREVIOUS block was treatment) alongside the current block's treatment indicator; a significant lagged-treatment coefficient is evidence of carryover that the washout period is not fully absorbing, and the washout should be lengthened.
Controlling time trends and seasonality (executed simulation): naive difference-in-means between treatment and control blocks is BIASED whenever the treatment schedule is correlated with time, which is common in real switchback schedules (e.g. an operationally simpler "2 weeks control, 2 weeks treatment" alternating pattern, rather than fully independent per-week randomization). I simulated 40 weekly blocks on a 2-week-alternating schedule, a true treatment effect of 0.05, a linear time trend of 0.004/week, and noise SD 0.05, and repeated 3,000 times:
- naive diff-in-means: mean bias = +0.0077 (about 15% relative bias on the true 0.05 effect), RMSE = 0.0175
- regression adjusting for the linear time trend (fitting y=β0+β1⋅week+β2⋅treat+ε and reading β2): mean bias = -0.0003 (essentially unbiased), RMSE = 0.0158
This confirms the analysis MUST include a time-trend term (or week fixed effects, or a seasonal/day-of-week control), not a raw difference in block means, whenever the treatment schedule is not perfectly independent of time, which it rarely is in operationally realistic switchback schedules.
Sample size / number of blocks: switchback power depends on the NUMBER OF INDEPENDENT TIME BLOCKS, not the number of users within a block (users within a block are correlated through the shared time-block treatment, so they do not each contribute independent information the way users in a standard parallel A/B test do); as a rule of thumb, plan for at least 15 to 20 independent blocks per arm to get a usable estimate of the block-to-block variance needed for valid inference (analogous to needing enough clusters in a cluster-randomized trial), and use a block-level (not user-level) t-test or randomization-inference test on the block averages, whose within-block sample size mainly reduces MEASUREMENT NOISE per block rather than increasing the effective number of independent observations.
Two cases where a crossover/switchback design is inappropriate: (1) when the treatment has a PERSISTENT, long-lasting effect that does not wash out within a reasonable block length, for example a pricing change that shifts long-term customer trust or a feature that permanently changes user habits, the "carryover" here is not a short-term artifact to buffer away but the actual long-run effect you are trying to measure, and a switchback would systematically underestimate it; (2) when the unit being switched (a whole market or region) is small enough, or the treatment disruptive enough, that ALTERNATING it repeatedly causes operational instability or a bad user experience (users noticing prices or features flip back and forth week to week), which is both a product-quality risk and a threat to validity if users start changing their OWN behavior in anticipation of the alternation.
Design a 2x2 factorial experiment testing a pricing change (A vs B) and a UX layout change (old vs new) on purchase conversion, given 100k eligible users per day, a 2% baseline conversion rate, target power 80%, and alpha 0.05. Compute the required sample size per cell, describe how you would allocate traffic, and explain how you would analyze and interpret the main effects and the interaction term.
Sample Answer
Direct answer: With a 2% baseline conversion rate, 100k eligible users/day, 80% power, and alpha 0.05, and assuming a 15% relative minimum detectable effect (the question does not state one explicitly, so I am stating this assumption up front, a common convention for this kind of sizing exercise; new baseline = 2.30%), the design needs roughly 18,346 users PER CELL to power the main effects, or about 36,693 users per cell to ALSO power the interaction at the same MDE. Traffic should be split evenly, 25% into each of the 4 cells (control/control, pricing-B/layout-old, pricing-A/layout-new, pricing-B/layout-new), randomized independently on pricing and layout so every user has an equal chance of any of the 4 combinations.
Sample-size computation (executed, two-proportion z-test):
For a MAIN EFFECT comparison (pooling the 2 cells that share a pricing level against the other 2), p1=0.02, p2=0.023 (15% relative lift):
n=(p2−p1)2(zα/22pˉ(1−pˉ)+zβp1(1−p1)+p2(1−p2))2gives n = 36,693 per POOLED arm (i.e. per 2-cell group). Since each pooled arm is 2 of the 4 cells, that is 18,346 users per CELL for main-effect power, 73,386 total across all 4 cells, which at 100k/day takes about 0.73 days.
For the INTERACTION contrast (an unpooled, single-cell-vs-single-cell comparison, since the interaction estimate uses the +1/-1/-1/+1 contrast across all four individual cells rather than pooled halves), the SAME formula applied per cell (no pooling) requires 36,693 users per cell, 146,771 total, about 1.47 days at 100k/day, exactly 2.00x the main-effect total. This is the general result cited in S0: interaction detection at the same MDE costs about 2x a main effect's sample size.
Traffic allocation: randomize pricing (A vs B) and layout (old vs new) INDEPENDENTLY via two separate hash-based coin flips per user (verified orthogonal in S9), producing four roughly-equal 25% cells automatically; do not manually stratify into 4 named buckets, since independent-factor randomization already balances the design and keeps it robust to later adding a third factor.
Analysis and interpretation: fit
y=β0+β1⋅Pricing+β2⋅Layout+β3⋅(Pricing×Layout)+εon the binary conversion outcome (linear probability model or logistic regression with an interaction term). β1 is the pricing main effect averaged over both layouts, β2 the layout main effect averaged over both pricing versions, and β3 the interaction: how much the pricing effect CHANGES when layout changes (equivalently, how much the layout effect changes when pricing changes). Run the interaction test FIRST: if β3 is significant, report cell-specific conversion rates (all 4 cells) rather than a single "average" pricing effect, since the average would be misleading if the effect flips sign across layouts. If the interaction is not significant, report the two main effects as if from independent tests and recommend shipping whichever combination of winning main effects performs best. Present results to stakeholders as: (1) the interaction test verdict first, in plain language, (2) a 2x2 table of observed conversion rates per cell with confidence intervals, (3) the recommended cell to ship, and (4) the guardrail-metric check on that winning cell (per this topic's sibling a-b-test-design-and-statistical-rigor, which owns guardrail-metric mechanics).
When should you choose a factorial or multivariate (MVT) design instead of running a sequence of separate A/B tests on individual features? Discuss the trade-offs in speed of learning, statistical power for main effects versus interactions, interpretability, and operational complexity, then give a concrete business scenario where a factorial design is clearly the better choice.
Sample Answer
Direct answer: Choose a factorial (or multivariate/MVT) design over a sequence of separate A/B tests when you have two or more candidate changes you plausibly want to ship TOGETHER, you suspect (or must rule out) an interaction between them, and you have enough concurrent traffic to split into cells without starving any one cell's power. Choose sequential single-factor A/B tests when changes are independent, traffic is scarce, or you need to ship and learn on the first change before the second is even designed.
Trade-offs
- Speed of learning: factorial wins when the alternative is truly SEQUENTIAL (test A, wait, then test B). For the same total sample budget, a k-factor design estimates all k main effects roughly as precisely as k separate two-arm tests run one after another with the SAME total budget split across them, but the factorial gets all k answers CONCURRENTLY instead of one-after-another, and gets the interaction essentially for free.
- Statistical power for main effects versus interactions: each main effect in a factorial is estimated by comparing half the cells against the other half (pooling over the other factor), so main-effect power is comparable to a standalone two-arm test at the same total N. The INTERACTION contrast is NOT pooled the same way: it needs the same per-cell sample size as a single unpooled cell-vs-cell comparison, so detecting an interaction of the same absolute size as a main effect costs roughly 2x the total sample of detecting that main effect alone (verified numerically in S1).
- Interpretability: a single-factor sequential test has a trivial reading ("does this metric move"). A factorial with a significant interaction requires explaining conditional effects ("A only works when B is on"), which is harder to communicate and can confuse stakeholders unless you plan for it (see S6).
- Operational complexity: factorial designs need every combination of treatments to be simultaneously deployable and measurable, more engineering coordination, more QA surface (interaction bugs, cross-feature conflicts), and a bigger simultaneous "blast radius" if something goes wrong, versus shipping and validating one change at a time.
Worked confirmation (executed simulation, not just cited): simulating 2 factors with true additive effects of size 0.12 (in outcome-SD units) each, total budget N=4000, alpha=0.05, over 4000 replications:
- FACTORIAL (4 cells of 1000 each): power(A) = 0.965, power(B) = 0.961
- SEQUENTIAL (2 separate 2-arm tests, same total budget split in half, 1000/arm each): power(A) = 0.771, power(B) = 0.759
For the identical total sample budget, the factorial design delivered materially higher power on BOTH main effects (0.96+ vs 0.77) because every unit in a factorial cell contributes to estimating EVERY factor's main effect (a unit in the (A=1,B=1) cell informs both the A comparison and the B comparison), while the sequential approach wastes half its budget on a control arm that is redundant across the two separate tests.
Business scenario where factorial is clearly preferable: a subscription product wants to test (1) a new pricing page copy and (2) a redesigned onboarding flow. Product suspects the pricing framing only lands if users have already been walked through the new onboarding (an interaction hypothesis: pricing-copy effect is conditional on onboarding version). Running these sequentially would answer "does new pricing help" and "does new onboarding help" separately but could never detect that new pricing ONLY helps when paired with new onboarding, exactly the finding the business needs to decide whether to ship them together, separately, or not at all. A 2x2 factorial answers all three questions (both main effects plus the interaction) in one concurrent test.
When NOT to use factorial: if you have 4+ factors, a FULL factorial's cell count (2^4 = 16 cells) can quickly outstrip available traffic; that is exactly the constrained-traffic case that motivates fractional factorial designs (S2, S3) rather than abandoning multi-factor testing altogether.
Unlock Full Question Bank
Get access to all 11 Advanced Experimentation Designs interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.