Experimentation Platforms and Infrastructure Questions
Infrastructure for A/B testing and experimentation: assignment/bucketing, metric pipelines for experiments, guardrail and variance-reduction plumbing, and experiment result storage. Covers building the platform that powers trustworthy online experiments at scale. Distinct from the statistics of experiment analysis.
As the head of the experimentation platform, create a 6-month plan to onboard 20 product teams onto a centralized platform while preserving statistical rigor and still letting teams move on their own schedule. Include the technical migration steps, training, templates, governance phases, success metrics for the rollout, and which platform features have to exist before you can onboard the first team.
Sample Answer
Direct answer
Sequence the 6-month onboarding so the platform's must-have features (self-serve config, automated validity checks, a shared metric catalog, and a set of standard templates) are proven with a small pilot group of teams before broad rollout, run technical migration and training in parallel tracks rather than serially, and measure success not just by "teams onboarded" but by whether those teams' experiments are actually meeting the statistical rigor bar, since onboarding a team that then runs sloppy experiments isn't real success.
Structured elaboration
- Months 1 to 2, pilot: onboard 3 to 4 representative teams (chosen to cover different risk profiles: a low-risk marketing team, a higher-risk payments-adjacent team) onto the platform, using this period to surface real gaps in the self-serve tooling, templates, and governance model before scaling to everyone.
- Templates: ship a small set of standardized, versioned templates so a team's onboarding lift is filling in a form, not inventing a process from scratch, an experiment-configuration template pre-wired to the shared metric catalog and default guardrails, a launch-readiness checklist template covering pre-registration and SRM setup, and a results-readout template that presents lift, confidence interval, and guardrail status in the platform's standard format. Teams that start from these templates inherit governance-compliant defaults automatically, which is a materially cheaper and more consistent onboarding path than each team rebuilding its own launch checklist and report format independently.
- Months 2 to 4, phased rollout: onboard the remaining teams in waves, each wave including migration of any existing homegrown experimentation the team was running (mapping their old ad hoc metrics into the shared catalog and their own launch process into the standard templates, not just their code), plus mandatory training on the platform's guardrails (why pre-registration matters, how to read a sequential-testing dashboard, how to fill out the launch-readiness template).
- Months 4 to 6, autonomy and governance maturity: shift from platform-team-assisted launches to genuine self-service for low-risk experiments using the standard templates, while the risk-tiered approval workflow (discussed elsewhere) takes over for anything higher-stakes, so the platform team's ongoing role becomes maintaining shared infrastructure and templates rather than hand-holding every launch.
- Required platform features before the FIRST team onboards: at minimum, self-serve configuration with pre-launch validation, SRM and basic guardrail monitoring, a shared metric catalog, and the experiment-configuration and launch-readiness templates; without these, the pilot teams would be testing an incomplete platform and the lessons learned wouldn't generalize.
- Success metrics: not just "number of teams onboarded" but the rate of properly pre-registered experiments, the rate of SRM-clean launches, template adoption rate (are teams actually using the standard templates or working around them), and (later) whether experiment velocity for onboarded teams actually increased versus their prior ad hoc process, since a slower but far more rigorous platform is a real trade-off worth being honest about, not something to hide behind a raw "teams onboarded" count.
Worked example
A team's existing homegrown pricing-experiment scripts define "conversion" slightly differently from the shared metric catalog. Migrating that team means more than pointing their launches at the new platform, it means an explicit reconciliation step where the team, together with the platform team, maps their historical definition to the standard one, documents where the two differ, and moves their launch process onto the standard experiment-configuration and results-readout templates, so historical experiment results and newly-onboarded ones remain comparable rather than silently using two different definitions of the same word and two different report formats.
Trade-offs and pitfalls
Rushing every team onto the platform in month one to hit an aggressive "teams onboarded" target risks skipping the reconciliation, training, and template-adoption work that actually makes the migration successful, producing teams that are nominally "onboarded" but still running experiments with the same rigor gaps they had before. The pilot-then-wave approach costs calendar time upfront but catches integration, governance, and template gaps while the blast radius is still small, which is cheaper than discovering them after all 20 teams have already migrated.
Given exposures, events, and users tables, describe how you would compute normalized lift and a bootstrap confidence interval for the metric purchase_value per variant over a 14-day window. Outline the SQL for the aggregation step and sketch the Python for the bootstrap, including how you handle user-level aggregation and stratified resampling.
Sample Answer
Direct answer
Aggregate purchase_value per user per variant in SQL first (so each user contributes exactly one row), then bootstrap by resampling users WITH replacement independently within each arm, recomputing the relative lift on each resample, and taking the 2.5th and 97.5th percentiles of the resulting distribution as the 95% confidence interval; user-level resampling, not event-level, is what correctly respects the actual unit of randomization.
Structured elaboration
SQL aggregation sketch:
SELECT e.variant, e.user_id,
COALESCE(SUM(CASE WHEN ev.event_type='purchase'
AND ev.occurred_at >= e.assigned_at
AND ev.occurred_at <= e.assigned_at + INTERVAL '14 days'
THEN ev.value ELSE 0 END), 0) AS purchase_value
FROM exposures e LEFT JOIN events ev ON ev.user_id = e.user_id
GROUP BY e.variant, e.user_id;
Both bounds of the attribution window matter: the upper bound (assigned_at + 14 days) caps how long after assignment a purchase still counts, but the lower bound (ev.occurred_at >= e.assigned_at) is just as necessary, without it the join would sum a user's entire purchase history, including purchases made long before the experiment even existed, into purchase_value, since the CASE condition only ever checked the upper bound. That silently inflates purchase_value for any user with pre-experiment purchase history and corrupts the metric everything downstream depends on.
Python bootstrap sketch (executed and verified):
def bootstrap_lift_ci(control_values, treatment_values, n_boot=2000, seed=42):
rng = random.Random(seed)
n_c, n_t = len(control_values), len(treatment_values)
lifts = []
for _ in range(n_boot):
c_sample = [control_values[rng.randrange(n_c)] for _ in range(n_c)]
t_sample = [treatment_values[rng.randrange(n_t)] for _ in range(n_t)]
c_mean, t_mean = sum(c_sample)/n_c, sum(t_sample)/n_t
if c_mean != 0:
lifts.append((t_mean - c_mean) / c_mean)
lifts.sort()
lo, hi = lifts[int(0.025*len(lifts))], lifts[int(0.975*len(lifts))]
point = (sum(treatment_values)/n_t - sum(control_values)/n_c) / (sum(control_values)/n_c)
return point, lo, hi
User-level aggregation matters because the randomization unit was the USER, not the event; resampling individual purchase events would treat a high-frequency purchaser's many events as many independent observations, when they're really one randomized unit. Stratified resampling (resampling within each variant independently, as done here) preserves each arm's own sample size in every bootstrap draw, keeping the comparison apples-to-apples.
Worked example (executed)
On a synthetic dataset with a known true +20% lift (control drawn from a distribution with mean 20, treatment from one with mean 24, 3,000 users per arm), the bootstrap produced a point estimate of approximately 0.19 with a 95% CI of roughly [0.17, 0.22], correctly containing the true 0.20 lift. On a null case (both arms drawn from the identical distribution), the CI straddled zero (approximately [-0.02, 0.02]), correctly showing no significant difference.
Trade-offs and pitfalls
Bootstrapping at the event level instead of the user level is the most common version of this mistake, and it understates the true variance because it implicitly treats a user's multiple purchases as independent randomized units when they aren't. A second common mistake, seen in an earlier draft of the SQL above, is an attribution window that only enforces its upper bound and forgets the lower one, silently pulling in a user's entire purchase history rather than just the post-assignment window; always write and test both bounds explicitly. The other pitfall is a small sample size making the percentile bootstrap itself noisy (with very few users per arm, 2,000 resamples of a tiny population doesn't manufacture information that isn't there); in that regime a parametric or bias-corrected bootstrap method is usually more appropriate than the plain percentile method used here.
You need to add a new mutually-exclusive experiment into live production traffic without reassigning users already exposed to other experiments, preserving exposure stickiness. What data structures and assignment ordering would you use, and how do you keep this cheap in state and latency on the evaluation path?
Sample Answer
Direct answer
Insert the new mutually-exclusive experiment into a dedicated traffic layer that only newly-eligible users (those not already committed to another experiment in that layer) can enter, using a priority or first-come ordering recorded per user the first time they're evaluated against that layer, so existing users already exposed to a different experiment are never reassigned.
Structured elaboration
The key data structure is a per-user, per-layer assignment record: the first time a user is evaluated against a mutually-exclusive layer, whichever experiment currently owns that layer slot for that user gets recorded, and every subsequent evaluation for that user in that layer reads the recorded assignment rather than recomputing it. Inserting a brand-new experiment into the layer means: any user with no existing record for that layer becomes eligible for the new experiment (hashed in as usual); any user WITH an existing record keeps it, full stop, regardless of what the new experiment's allocation logic would have computed for them.
To keep this cheap in state and latency, the assignment record doesn't need to be a giant table looked up synchronously on every request; it can be derived on first-write and then cached the same way any other assignment is cached, with the "first eligible slot wins and is sticky" rule enforced at write time rather than checked on every read.
Worked example
A layer currently runs Experiment A at 100% of its slot. A new Experiment B needs to launch into the same layer at a 20% share. New users hashed into that 20% get Experiment B; users who were already recorded as being in Experiment A (even if a fresh hash of their id would have landed them in B's 20%) keep A, because their assignment record already exists. This is what "preserving exposure stickiness" means concretely: existing users' experience doesn't change just because the layer's occupancy changed.
Trade-offs and pitfalls
The subtlety that trips people up is conflating "the layer's allocation percentages changed" with "existing users should be re-hashed against the new percentages." Re-hashing existing users the moment allocations change reintroduces exactly the instability this design is meant to prevent (a user who was in A suddenly finding themselves in B mid-experiment, with no clean causal interpretation of either experiment's result for that user). The cost of correctness here is that a layer's actual traffic split can drift from its nominal configured percentages for a while, since existing users are grandfathered in at their original assignment; this is expected and should be visible in the platform's reporting, not treated as a bug.
You must join exposure logs from one service with conversion events from a different service to compute experiment metrics, and the two systems use different timezone conventions and session-windowing heuristics. Design a reconciliation process: canonical timestamps, session windowing, deduplication keys, handling of late-arriving events, idempotency, and the tests you would write to prove the joined metric is correct.
Sample Answer
Direct answer
Design the reconciliation around canonical UTC timestamps computed as close to the source as possible, a shared, explicit session-windowing definition both services agree to use (rather than each inferring sessions independently), deduplication keys based on a stable event id rather than a recomputed hash, idempotent processing so a retry doesn't double-count, and automated tests that assert the joined output on a known input produces a known, hand-verified result.
Structured elaboration
- Canonical timestamps: convert every timestamp to UTC at the earliest possible point (ideally at the source service, not downstream), and store the original timezone/offset alongside it for debugging, since a downstream conversion error is much harder to spot after the original context is lost.
- Session windowing: if Service A and Service B use different session-timeout definitions (say, 30 minutes of inactivity versus a fixed calendar day), events that should logically belong to the same session can get split differently by each service; the reconciliation process needs a single, explicitly agreed session definition applied consistently across both, computed at reconciliation time rather than trusting each service's own internal session boundary.
- Deduplication keys: use a stable, source-assigned event id (not a hash of mutable fields, which can change if any field is corrected or reprocessed) so a legitimate retry is recognized as the same event rather than treated as new.
- Late-arriving events: define an explicit lateness tolerance (how long after a session's nominal end an event can still arrive and be included) consistent with the storage-tiering and watermarking approach used elsewhere in the pipeline.
- Idempotency: the join and aggregation logic should produce the same result whether run once or replayed multiple times over the same input, which usually means writing to an idempotent sink (an upsert keyed by the same deduplication key) rather than a naive append.
- Tests for correctness: a golden-file style test with hand-constructed events from both services, including a case that straddles a timezone boundary and a case with a legitimate retry, verifying the joined output matches a manually-computed expected result.
Worked example
Service A logs an exposure at 23:50 local time in a timezone that's UTC+9, while Service B logs the matching conversion event using a server timestamp already in UTC. Without converting Service A's timestamp to UTC before comparison, a naive join could conclude the conversion happened BEFORE the exposure (since 23:50 local time is actually 14:50 UTC the same day, a nine-hour difference that a naive string or local-time comparison would miss entirely), silently excluding a legitimate conversion from the metric.
Trade-offs and pitfalls
Trusting each service's own internal session or timestamp handling instead of establishing one shared, explicit definition at reconciliation time is the root cause of most cross-service join bugs like this; it's tempting because it requires no coordination between the two teams, but it's exactly the coordination that prevents silent, hard-to-detect metric corruption. The cost of building the shared reconciliation layer (agreeing on and enforcing one canonical timestamp and session definition) is real coordination overhead between two teams, but it's cheap relative to the cost of a metric quietly being wrong for months before anyone traces it back to a timezone mismatch.
Propose an automated guardrail system that flags harmful regressions shortly after an experiment launches. What statistical tests, thresholds, and alert tiers would you use, and when would the system take automated action (pausing the ramp or rolling traffic back) versus just paging a human?
Sample Answer
Direct answer
An automated guardrail system needs a defined set of guardrail metrics with pre-agreed statistical tests and thresholds, a tiered alerting scheme (a small dip pages a human, a large or fast-moving dip triggers automatic action), and an explicit distinction between pausing the ramp (stop exposing new users, keep existing ones as-is) and rolling back (revert everyone), because those are different remediations for different severities.
Structured elaboration
- Statistical tests: a one-sided test against a pre-agreed harm threshold (not just "any statistically significant change") is usually the right framing for a guardrail, since the goal is catching real harm fast, not detecting any change in either direction with maximum power.
- Thresholds and tiers: a common design is two tiers, a "watch" threshold (a smaller, statistically detected regression) that pages the on-call owner for a human decision, and a "stop" threshold (a larger or faster regression, or a regression combined with high statistical confidence) that triggers automated action without waiting for a human.
- Automated protections: pausing the ramp (freezing the traffic percentage, or reverting new users to control while leaving already-exposed users in place) is the lower-risk default action; a full rollback (reverting everyone immediately) is reserved for guardrails severe enough that leaving anyone exposed even briefly is unacceptable, like a payment-failure spike.
- Avoiding false alarms: because this system runs continuously across every guardrail metric on every experiment, it needs its own multiple-testing discipline (checking dozens of metrics many times a day inflates the false-alarm rate) and typically a minimum-duration or minimum-sample-size floor before an alert can fire at all, so early noise in a tiny sample doesn't trigger a spurious automatic rollback.
Worked example
A login-flow experiment is being monitored for guardrail metrics including a technical one (login success rate) and a business one (session-start rate). Login success rate drops sharply and quickly past the "stop" threshold within the first hour of a ramp. The system automatically pauses new exposure and pages the on-call engineer, rather than waiting for a scheduled daily report, because an hour of a login-flow regression at scale is costly enough that the automated action has to happen faster than any human review cycle could. The response here is deliberately a roll-FORWARD (pause new exposure, keep serving the same code path to already-affected users while a fix ships) rather than a full rollback, since the team judges that a fast fix is available and re-exposing the already-affected cohort to yet another code change would add more risk than it removes; a different guardrail (a payment-failure spike, say) would instead call for an immediate full rollback, which is exactly why the decision between rolling forward and rolling back has to be a judgment call encoded per guardrail severity, not a single fixed response for every breach.
Trade-offs and pitfalls
Setting the automatic-action threshold too aggressive causes false rollbacks on ordinary noise, which erodes trust in the system and tempts teams to disable it; setting it too conservative means real harm runs for longer before anything happens. The typical resolution is to start conservative (a genuinely severe, high-confidence threshold for full automatic rollback) while using the "watch" tier liberally, since a false page to a human is a much cheaper mistake than a false automatic rollback that a team then has to understand and re-enable.
Unlock Full Question Bank
Get access to all Experimentation Platforms and Infrastructure interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.