Experimentation Platforms and Infrastructure Questions
Infrastructure for A/B testing and experimentation: assignment/bucketing, metric pipelines for experiments, guardrail and variance-reduction plumbing, and experiment result storage. Covers building the platform that powers trustworthy online experiments at scale. Distinct from the statistics of experiment analysis.
Your experimentation platform runs thousands of tests a year and everyone trusts its p-values without question. How would you actually verify, at the platform level, that its statistical machinery is calibrated correctly, rather than trusting it because nobody has complained? Describe how you would use A/A tests for this, including concrete numeric thresholds (for example, what fraction of A/A runs showing p < 0.05 should worry you) and what you would do if calibration fails.
Sample Answer
Direct answer
Verify platform-level calibration by running periodic A/A tests (splitting real traffic into two identical, no-difference groups) and checking that the distribution of p-values across many A/A runs is close to uniform, that no more than roughly the nominal false-positive rate (about 5% at alpha=0.05) come back significant, and that assignment stays balanced throughout; if any of those checks fail, the statistical machinery itself is miscalibrated and every real experiment's p-value is suspect until it's fixed.
Structured elaboration
The logic: under a true null (no real difference, which an A/A test guarantees by construction), a correctly implemented significance test should produce a p-value that is uniformly distributed between 0 and 1, and should come back "significant" at alpha=0.05 roughly 5% of the time, purely by chance, no more and no less. Running many A/A tests (across different metrics, different traffic segments, different times) and checking this empirically is a direct test of whether the platform's variance estimation, its handling of dependent observations, or its multiple-testing correction are actually correct, independent of any real product question.
Concrete numeric thresholds worth watching: if noticeably more than 5% of A/A runs come back significant at alpha=0.05 (say, consistently landing above 8 to 10%), that's a strong signal the variance estimate is too small somewhere, commonly from treating correlated observations (repeated visits from the same user, or clustered users) as independent. If the p-value distribution is visibly non-uniform (bunched near 0, or bunched near 1) across many A/A runs, that points at a systematic bug in the test implementation itself rather than random noise.
Worked example
Running 100 independent A/A tests over a quarter, 14 come back significant at alpha=0.05, roughly triple the expected 5. Investigating further shows the variance calculation was using a simple per-event standard error rather than a per-user one, meaning power users who generated many events were effectively counted multiple times in the variance estimate, understating the true standard error and inflating the false-positive rate. This is exactly the kind of miscalibration that would silently make every real experiment's result look more confident than it actually is, until this specific bug is found and fixed.
Trade-offs and pitfalls
Running enough A/A tests to detect a modest calibration problem with confidence takes real, otherwise-unused traffic and time; a platform under pressure to maximize real experiment throughput can be tempted to skip this. The other pitfall is treating a single A/A test's non-significant result as proof of calibration: any one A/A run coming back non-significant tells you almost nothing (that's the expected outcome 95% of the time regardless), the signal only comes from the AGGREGATE behavior across many runs.
Your product team reports unexpected metric contamination after several rapid rollouts and overlapping feature flags. Walk through the operational, step-by-step plan you would run to identify, quantify, and mitigate the contamination sources while minimizing disruption to the teams shipping features.
Sample Answer
Direct answer
Run a step-by-step operational response: first quantify the contamination's actual scope (which experiments, which time window, roughly how many users affected), then decide whether affected experiments need to pause, be re-analyzed with an explicit contamination adjustment, or be allowed to continue with the contamination noted as a caveat, and only after containment turn to the longer-term fix (a coordination policy or a traffic-layering change) that prevents the same class of collision from recurring.
Structured elaboration
- Scope the problem first, before acting: pull the actual user-overlap and timing data across the flags and rollouts implicated, rather than reacting to the qualitative complaint alone. Contamination reports are often vaguer than the actual data ("something feels off") and the first job is turning that into a specific, bounded claim.
- Triage by severity: an overlap that's small and unlikely to meaningfully bias either experiment's primary metric can often be left running with a documented caveat; a large, high-confidence contamination affecting a decision that's about to be made should pause the affected experiments immediately.
- Communicate to the teams whose experiments are affected: they need to know before they make a ship decision on data that might be compromised, not after.
- Fix the immediate cause: this is often an unplanned interaction between two independently-launched flags or a rollout schedule collision rather than a single experiment's fault, so the fix may involve coordinating a schedule change with a different team, not just editing one experiment's configuration.
- Address the systemic cause afterward: only once the immediate situation is contained does it make sense to invest in the longer-term prevention, whether that's an automated pre-launch overlap check, a shared calendar of high-traffic-surface launches, or a formal coordination/prioritization policy for a specific contested surface (a busy city, a shared page) where conflicts recur.
Worked example
Two rapid feature-flag rollouts on the same page overlap for four days before anyone notices, and a third team's active experiment shares that page too. Quantifying the actual overlap shows it affects roughly 12% of one experiment's traffic during a 4-day window out of its planned 14-day run. Given the size and duration, the team decides to exclude that 4-day window from the final analysis (using the remaining 10 days, which is still enough for the planned power) rather than pausing and restarting, which would have cost the experiment more total time than simply excluding the contaminated window.
Trade-offs and pitfalls
Reacting to every contamination report by immediately pausing everything involved is the safest but most disruptive option, and overusing it trains teams to under-report ambiguous cases to avoid the disruption; underreacting (dismissing reports without actually quantifying scope) risks shipping a decision based on genuinely compromised data. The resolution is making the FIRST step, quantifying actual scope, fast and cheap enough that a team doesn't have to choose between "ignore it" and "stop everything" before they even know how bad it is.
What is the minimum feature set you would expect from a self-serve experiment configuration and rollout system? Cover experiment definition, targeting, percentage rollout, pre-launch checks, and rollback.
Sample Answer
Direct answer
At minimum a self-serve experiment configuration and rollout system needs: a way to define the experiment (hypothesis, metrics, eligible population), targeting rules (who is eligible and how they're split), a percentage-rollout mechanism (start small, ramp gradually), automated pre-launch checks (does the configuration even make sense before it touches real traffic), and a rollback path (turn it off instantly if something goes wrong).
Structured elaboration
- Experiment definition: name, hypothesis, primary and guardrail metrics, owner, and a status field. Without a required hypothesis and metric field at creation time, teams launch experiments they can't actually read out later.
- Targeting: an eligibility filter (which users can be included at all, for example, only users in a specific country or app version) plus the variant-allocation percentages.
- Percentage rollout: the ability to start an experiment at 1 or 5 percent of eligible traffic and ramp it up over defined steps rather than launching straight to 50/50, so a bad configuration only affects a small fraction of users before anyone notices.
- Pre-launch checks: automated validation that the configuration is internally consistent (allocations sum to 100%, the targeting filter isn't empty, the metrics referenced actually exist) before the "launch" button is allowed to do anything.
- Rollback: a single action that instantly reverts every affected user back to the control experience, independent of whether the original deploy that shipped the experimental code is still safe to run.
Worked example
A marketing team wants to test a new signup flow. Without pre-launch checks, they could accidentally launch with a targeting filter that matches zero users (a typo in a country code), burn a week believing the experiment is running, and only discover during read-out that it collected no data. A pre-launch check that flags "this targeting rule currently matches 0 users" catches this before the button is even pressed.
Trade-offs and pitfalls
Making the self-serve system too permissive (letting any team launch to 100% immediately) trades safety for speed and is how a single bad configuration change can affect the entire user base at once; making it too restrictive (requiring platform-team sign-off on every launch) defeats the entire point of self-serve and creates a bottleneck. The usual resolution is a risk-tiered gate: launches under some traffic threshold with no schema or backend changes can go fully self-serve, while anything touching payments, a schema migration, or traffic above a threshold requires an explicit approval step.
Propose an automated guardrail system that flags harmful regressions shortly after an experiment launches. What statistical tests, thresholds, and alert tiers would you use, and when would the system take automated action (pausing the ramp or rolling traffic back) versus just paging a human?
Sample Answer
Direct answer
An automated guardrail system needs a defined set of guardrail metrics with pre-agreed statistical tests and thresholds, a tiered alerting scheme (a small dip pages a human, a large or fast-moving dip triggers automatic action), and an explicit distinction between pausing the ramp (stop exposing new users, keep existing ones as-is) and rolling back (revert everyone), because those are different remediations for different severities.
Structured elaboration
- Statistical tests: a one-sided test against a pre-agreed harm threshold (not just "any statistically significant change") is usually the right framing for a guardrail, since the goal is catching real harm fast, not detecting any change in either direction with maximum power.
- Thresholds and tiers: a common design is two tiers, a "watch" threshold (a smaller, statistically detected regression) that pages the on-call owner for a human decision, and a "stop" threshold (a larger or faster regression, or a regression combined with high statistical confidence) that triggers automated action without waiting for a human.
- Automated protections: pausing the ramp (freezing the traffic percentage, or reverting new users to control while leaving already-exposed users in place) is the lower-risk default action; a full rollback (reverting everyone immediately) is reserved for guardrails severe enough that leaving anyone exposed even briefly is unacceptable, like a payment-failure spike.
- Avoiding false alarms: because this system runs continuously across every guardrail metric on every experiment, it needs its own multiple-testing discipline (checking dozens of metrics many times a day inflates the false-alarm rate) and typically a minimum-duration or minimum-sample-size floor before an alert can fire at all, so early noise in a tiny sample doesn't trigger a spurious automatic rollback.
Worked example
A login-flow experiment is being monitored for guardrail metrics including a technical one (login success rate) and a business one (session-start rate). Login success rate drops sharply and quickly past the "stop" threshold within the first hour of a ramp. The system automatically pauses new exposure and pages the on-call engineer, rather than waiting for a scheduled daily report, because an hour of a login-flow regression at scale is costly enough that the automated action has to happen faster than any human review cycle could. The response here is deliberately a roll-FORWARD (pause new exposure, keep serving the same code path to already-affected users while a fix ships) rather than a full rollback, since the team judges that a fast fix is available and re-exposing the already-affected cohort to yet another code change would add more risk than it removes; a different guardrail (a payment-failure spike, say) would instead call for an immediate full rollback, which is exactly why the decision between rolling forward and rolling back has to be a judgment call encoded per guardrail severity, not a single fixed response for every breach.
Trade-offs and pitfalls
Setting the automatic-action threshold too aggressive causes false rollbacks on ordinary noise, which erodes trust in the system and tempts teams to disable it; setting it too conservative means real harm runs for longer before anything happens. The typical resolution is to start conservative (a genuinely severe, high-confidence threshold for full automatic rollback) while using the "watch" tier liberally, since a false page to a human is a much cheaper mistake than a false automatic rollback that a team then has to understand and re-enable.
Given exposures(exposure_id, user_id, experiment_id, variant, assigned_at) and events(event_id, user_id, event_type, value, occurred_at), write a SQL query that computes the conversion rate (share of exposed users who generated a purchase event) per variant using a 14-day post-exposure window. Output variant, exposed_users, converters, conversion_rate.
Sample Answer
Direct answer
Compute conversion rate per variant by counting distinct exposed users per variant as the denominator, and counting distinct users who generated a purchase event within 14 days AFTER their own exposure timestamp as the numerator, being careful that the window is bounded on BOTH sides: it must start at the user's own assigned_at (excluding any purchase that happened before they were ever exposed) and end 14 days later (excluding anything past the window), anchored per-user, not to a single fixed calendar cutoff for everyone.
Structured elaboration
SELECT
e.variant,
COUNT(DISTINCT e.user_id) AS exposed_users,
COUNT(DISTINCT CASE
WHEN ev.event_type = 'purchase'
AND ev.occurred_at >= e.assigned_at
AND ev.occurred_at <= e.assigned_at + INTERVAL '14 days'
THEN e.user_id END) AS converters,
CAST(COUNT(DISTINCT CASE
WHEN ev.event_type = 'purchase'
AND ev.occurred_at >= e.assigned_at
AND ev.occurred_at <= e.assigned_at + INTERVAL '14 days'
THEN e.user_id END) AS REAL)
/ COUNT(DISTINCT e.user_id) AS conversion_rate
FROM exposures e
LEFT JOIN events ev ON ev.user_id = e.user_id
GROUP BY e.variant
ORDER BY e.variant;
Key points, complexity, and edge cases
The window is bounded on BOTH sides from each row's own assigned_at: occurred_at >= e.assigned_at excludes any purchase that happened before the user was exposed at all (a real risk for a repeat customer who purchased weeks earlier, unrelated to this experiment), and occurred_at <= e.assigned_at + INTERVAL '14 days' excludes anything past the attribution window. A query missing the lower bound will silently inflate the converter count with pre-exposure purchases; this is exactly the kind of bug that looks fine on a clean worked example (no test data with a pre-exposure purchase) and only surfaces once real, messy production data (a customer with purchase history predating the experiment) hits it. COUNT(DISTINCT ...) inside the CASE, rather than a plain COUNT, correctly avoids double-counting a user who made two purchases within the window. Complexity is a single-pass join and group-by, O(n) in the size of the joined exposures-and-events dataset; a user with no purchase events correctly contributes to exposed_users but not to converters, since the LEFT JOIN preserves them with no matching purchase row.
I executed this query in SQLite against a constructed test: two users per variant, one purchase inside the window (control, converts), one purchase exactly 15 days after exposure for the other variant (treatment, one day past the 14-day window). The query returned control: 2 exposed / 1 converter / 0.5 rate, and treatment: 2 exposed / 0 converters / 0.0 rate, correctly excluding the late purchase. I then added an adversarial case: a user exposed on 2026-02-10 with a purchase logged on 2026-01-05 (over a month before exposure). With only an upper-bound date filter (no occurred_at >= assigned_at), that pre-exposure purchase was incorrectly counted as a conversion (rate 1.0 for a single-user group); adding the lower bound above correctly excludes it (in a 3-user control group with that case added, converters = 1, not 2, giving a 0.333 rate instead of the wrong 0.667).
Trade-offs and pitfalls
A common mistake is joining events to exposures without restricting to event_type = 'purchase' inside the conversion condition, which would count ANY event (a click, a page view) as a conversion. An equally common and easy-to-miss mistake, confirmed above, is forgetting the LOWER bound on the attribution window entirely: filtering only on occurred_at <= assigned_at + 14 days looks correct and passes any test case that doesn't include a pre-exposure purchase, but silently inflates the converter count in production the moment a user with prior purchase history is exposed. Another mistake is using a single WHERE clause date filter (occurred_at <= '2026-01-15') instead of anchoring the cutoff per row to each user's own assigned_at, which silently breaks the moment users are exposed on different days, which they almost always are in a real rolling experiment.
Unlock Full Question Bank
Get access to all Experimentation Platforms and Infrastructure interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.