Experimentation Platforms and Infrastructure Questions
Infrastructure for A/B testing and experimentation: assignment/bucketing, metric pipelines for experiments, guardrail and variance-reduction plumbing, and experiment result storage. Covers building the platform that powers trustworthy online experiments at scale. Distinct from the statistics of experiment analysis.
Given exposures(user_id, experiment_id, variant, exposure_ts) and events(user_id, event_ts, event_type, value), write a SQL query that: (1) deduplicates exposures per user, keeping the earliest exposure; (2) computes each user's 7-day sum of purchase value after their exposure; (3) for users exposed to multiple variants, assigns them to the last-exposed variant only if that exposure occurred before their first purchase; and (4) treats an event as attributable only if event_ts falls within 7 days of the exposure. Output user_id, assigned_variant, purchase_7d_sum.
Sample Answer
Direct answer
Deduplicate exposures per user by keeping the earliest one by default, but reassign to the LAST-exposed variant specifically when that later exposure happened before the user's first purchase, since the spec's intent is attributing the purchase to whichever variant the user was actually under at the time they converted, not simply whichever variant they saw first; then compute the 7-day sum of purchase value strictly AFTER, and within 7 days of, whichever exposure timestamp ends up being the assigned one (the later exposure's timestamp if reassigned, the earliest exposure's timestamp otherwise).
Approach and code
Build a per-user candidate assignment using window functions to identify both the earliest and the latest exposure, decide between them based on where the first purchase falls relative to the latest exposure's timestamp, and carry the ASSIGNED exposure's own timestamp (not always the earliest one) forward as the anchor for the 7-day attribution window, with an explicit lower bound so a purchase before that anchor is never included.
WITH first_purchase AS (
SELECT user_id, MIN(event_ts) AS first_purchase_ts
FROM events WHERE event_type = 'purchase' GROUP BY user_id
),
ordered_exposures AS (
SELECT e.*,
ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY exposure_ts ASC) AS rn_earliest,
ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY exposure_ts DESC) AS rn_latest
FROM exposures e
),
final_assignment AS (
SELECT a.user_id,
CASE WHEN le.exposure_ts != a.exposure_ts
AND fp.first_purchase_ts IS NOT NULL
AND le.exposure_ts < fp.first_purchase_ts
THEN le.variant ELSE a.variant END AS assigned_variant,
CASE WHEN le.exposure_ts != a.exposure_ts
AND fp.first_purchase_ts IS NOT NULL
AND le.exposure_ts < fp.first_purchase_ts
THEN le.exposure_ts ELSE a.exposure_ts END AS assigned_exposure_ts
FROM (SELECT * FROM ordered_exposures WHERE rn_earliest = 1) a
JOIN (SELECT * FROM ordered_exposures WHERE rn_latest = 1) le USING (user_id)
LEFT JOIN first_purchase fp USING (user_id)
)
SELECT fa.user_id, fa.assigned_variant,
COALESCE(SUM(CASE WHEN ev.event_type='purchase'
AND ev.event_ts >= fa.assigned_exposure_ts
AND ev.event_ts <= fa.assigned_exposure_ts + INTERVAL '7 days'
THEN ev.value ELSE 0 END), 0) AS purchase_7d_sum
FROM final_assignment fa
LEFT JOIN events ev ON ev.user_id = fa.user_id
GROUP BY fa.user_id, fa.assigned_variant;
Key points: window functions identify earliest and latest exposure per user in one pass; the CASE logic implements the "reassign only if the later exposure preceded the first purchase" rule exactly as specified, and now computes the ASSIGNED exposure's own timestamp alongside the assigned variant, rather than always keeping the earliest exposure's timestamp regardless of which variant was assigned; the 7-day window is anchored to assigned_exposure_ts (whichever exposure the user was actually attributed to) with an explicit >= lower bound so a purchase before that exposure is never counted.
Complexity and edge cases
Complexity is O(n log n) for the window-function sort per user, dominated by the exposures and events table sizes; a user with no purchases returns a 0 sum via COALESCE rather than NULL; a user with only one exposure trivially has earliest equal to latest, so the CASE always falls through to the single exposure's variant and timestamp. I re-verified this against a constructed test with disambiguating cases: (a) a user reassigned to a later variant before their first purchase, with a purchase that falls within 7 days of the LATER exposure but outside 7 days of the EARLIER one, which now correctly counts (the original query incorrectly excluded it); (b) a user with a purchase that occurred BEFORE their exposure, which now correctly returns 0 (the original query incorrectly included it); (c) a user whose later exposure arrived AFTER their first purchase, correctly NOT reassigned, with a late purchase correctly excluded by the 7-day cutoff.
Trade-offs and pitfalls
The tempting simpler approach, always keeping the earliest exposure and ignoring later ones entirely, is wrong under this spec because it would misattribute the purchase to a variant the user was no longer actually under by the time they converted. The window-function approach used here scales well in a data warehouse but assumes exposure and event timestamps are already reasonably clean and synchronized; a real production version would need the same timezone and late-arrival handling discussed for the general metrics pipeline before trusting these numbers.
Design a scalable automated validity-check engine for an experimentation platform: it should run pre-launch and continuous post-launch checks covering SRM, instrumentation drift, delayed events, outliers, and contamination or leakage between variants. Describe the architecture, how rules execute, where machine-learning-based anomaly detection helps versus rule-based checks, and the remediation workflow when a check fails.
Sample Answer
Direct answer
A validity-check engine needs a rule-execution layer that runs a defined library of checks (SRM, instrumentation drift, delayed events, outliers, contamination) both once before launch and continuously during the run, an architecture that separates fast rule-based checks from slower machine-learning-based anomaly detection, and a remediation workflow that turns a failing check into a specific, actionable next step rather than just a red dot on a dashboard.
Structured elaboration
flowchart TD
Launch[Experiment Launch Request] --> PreCheck[Pre-launch rule engine]
PreCheck -->|pass| Live[Live Traffic]
PreCheck -->|fail| Block[Block launch, report reason]
Live --> Continuous[Continuous check scheduler]
Continuous --> RuleChecks[Rule-based: SRM, delayed events, outliers]
Continuous --> MLChecks[ML-based: drift, contamination anomaly scoring]
RuleChecks --> Remediation[Remediation workflow]
MLChecks --> Remediation
Remediation --> Pause[Auto-pause]
Remediation --> Page[Page owner]
- Rule-based checks (SRM, an event-volume floor, a delayed-events percentage) are cheap, deterministic, and explainable, which makes them a good fit for anything that needs to BLOCK a launch outright.
- ML-based anomaly detection earns its keep specifically where a fixed rule would be too rigid: detecting contamination between two experiments that share partial traffic, or flagging a metric drift pattern that doesn't fit a simple threshold (a gradual, compounding drift rather than a sudden step change). The trade-off is explainability: a rule-based check can tell an engineer exactly which threshold was crossed; an ML-based flag often needs a secondary "why did this fire" explanation layer or it becomes a black box nobody trusts.
- Remediation workflows: every failing check should map to a specific downstream action (block launch, pause new exposure, flag reporting as unreliable, page an owner), not just an entry in a generic alerts feed, since an alert with no clear next action trains people to ignore the whole system.
- Naming the drift taxonomy explicitly: "drift" isn't one thing; the engine should distinguish a mean shift (the metric's average level moved), a variance increase (the metric got noisier without necessarily moving on average), and a distributional change (the shape changed even if the mean looks stable), because each implies a different root cause and a different fix, and reporting all three under one generic "drift detected" label makes triage slower, not faster.
- An incident-response playbook for false positives specifically: because this engine runs continuously across every active experiment, it will generate false alarms; a documented playbook (check the last deploy, check SDK-version breakdown, check whether the same alert fired on an unrelated experiment at the same time, which would point at a platform-wide issue rather than an experiment-specific one) keeps a false-positive investigation fast and consistent rather than reinvented each time.
Worked example
A contamination check using a rule-based threshold (say, more than 5% user overlap between two experiments) would miss a subtler case: two experiments with only 2% direct user overlap but where that 2% happens to be disproportionately high-value users who account for a large share of the metric's variance. An ML-based anomaly model trained to weight overlap by its expected influence on the outcome metric, rather than raw user-count overlap, catches this case that a flat threshold rule would pass cleanly.
Trade-offs and pitfalls
Building the ML layer before the rule-based layer is a common overreach: most of the value (catching SRM, instrumentation drift, obvious contamination) comes from cheap, explainable rules, and ML-based anomaly detection should be reserved for the residual cases rules genuinely can't express, not used as the default approach everywhere. The other pitfall is under-investing in the remediation workflow itself: a technically excellent detection engine that only produces a dashboard nobody acts on has the same practical effect as no detection engine at all.
Explain the difference between randomization and assignment in an experimentation platform. Compare simple (hash-based) randomization, stratified randomization, and block/random-permutation approaches, and give one practical scenario where each is the right choice.
Sample Answer
Direct answer
Randomization is the mechanism that decides how users get split into groups; assignment is the act of recording and looking up which group a specific user ended up in. Simple (hash-based) randomization treats every eligible user independently and equally likely to land in any variant. Stratified randomization first splits users into subgroups (country, new versus returning) and randomizes within each subgroup, so the split stays balanced across those subgroups even with a small sample. Block or permuted-block randomization forces a fixed ratio within consecutive chunks of arriving traffic, so the running split never drifts far from target even early in the ramp.
Structured elaboration
- Simple hashing: hash(user_id, experiment_id) mapped to [0, 1), then compared against a cumulative allocation threshold. Cheap, stateless, and the default choice whenever the eligible population is reasonably large and homogeneous, because with enough users the law of large numbers balances any covariate you didn't explicitly stratify on.
- Stratified randomization: randomize separately within each stratum (say, five countries), so each stratum individually gets the target split. This matters when a covariate is both unevenly distributed AND correlated with the outcome metric, and your total sample size isn't large enough to trust that simple hashing balances it by chance.
- Block/permuted-block randomization: guarantee, for example, that every consecutive batch of 100 arrivals contains exactly 50 control and 50 treatment, regardless of arrival order. This matters when you cannot tolerate a temporary imbalance even in the first hour of a ramp, for example if an early guardrail check depends on having a reasonably balanced sample immediately.
Worked example
A ride-hailing company launches a pricing experiment where 70% of daily traffic is in three home markets and 30% is spread thinly across twenty smaller markets. With simple hashing, the smaller markets might individually land at something like 40/60 splits purely by chance, because each one only has a few hundred users a day. Stratifying by market guarantees each market's own split stays close to 50/50, which matters here because the team specifically wants to report results per market, not just in aggregate.
Determinism matters here for a reason worth stating explicitly: the SAME user must land in the SAME group every time they're evaluated, or the experiment isn't measuring a stable treatment at all, and the "unit of randomization" (whatever a hash key is computed over: a user, a household, a session) is the thing that has to stay fixed for that guarantee to hold.
Trade-offs and pitfalls
Stratification and blocking add real engineering cost: every additional stratum is another dimension the assignment service must track state for (or, for hash-based stratification, another input folded into the hash key), and the analysis afterward must account for the stratification design (for example, using a stratified variance estimator) or it will understate the true precision gained. The common mistake is stratifying on a variable that turns out not to correlate much with the outcome, which pays the engineering cost without meaningfully reducing variance.
What traffic allocation strategies do experimentation platforms use, static splits, progressive rollouts, hash-based bucketing, and weighted allocation? For each, give a short example of when it is the right choice and its main trade-off.
Sample Answer
Direct answer
Static splits are simplest and right when you want a fixed, known percentage for the whole experiment duration; progressive rollouts (ramping traffic up in stages) are right when you want to limit blast radius from an undiscovered bug before committing to a larger split; hash-based bucketing is the underlying mechanism that makes any of these deterministic and consistent per user; and weighted allocation (uneven splits, like 90/10) is right when you want to minimize exposure to a risky variant while still collecting enough data to learn from it.
Structured elaboration
- Static splits (say, a fixed 50/50 for the whole run): simplest to reason about and analyze, appropriate once you already trust the variant is safe and just need a clean, stable comparison.
- Progressive rollouts (1% then 5% then 25% then 50%, each stage gated by automated checks): the right choice specifically when the variant itself is new and unproven, since it limits how many users are affected by an undiscovered bug before the team has had a chance to observe it at small scale; the trade-off is a slower path to reaching full statistical power, since the effective sample size only fully accrues once the ramp reaches its final percentage.
- Hash-based bucketing: not really an alternative to the others so much as the underlying mechanism that makes ANY of them deterministic; a progressive rollout is typically implemented as the same hash function with a shifting threshold over time, so a user who's included at 10% stays included as the ramp expands to 25%, rather than being re-randomized at each stage.
- Weighted allocation: appropriate when you want to keep exposure to a risky or resource-intensive variant small on principle (not just as a temporary ramp stage, but as the intended steady-state split, like a 90/10 test of an expensive-to-serve feature), trading some statistical power for a smaller worst-case exposure if the variant turns out badly.
Worked example
A team launching a genuinely new checkout flow chooses a progressive rollout rather than a static 50/50 split, starting at 1% for the first day specifically because the flow has never run in production before; once the first day's automated checks (SRM, guardrails, crash rate) pass cleanly, the ramp proceeds to 10%, then 50%, reaching the final split over about a week rather than immediately, which limits how many users would have been affected had a serious bug only surfaced once real traffic hit the new flow.
Trade-offs and pitfalls
Choosing a static 50/50 split for a genuinely unproven, high-risk change skips the safety benefit progressive rollout provides, exposing the maximum possible share of users to an undiscovered bug from the very first moment; choosing an overly cautious progressive rollout for a low-risk, well-understood change (a copy tweak) needlessly slows down reaching statistical power for no real safety benefit. The right choice is calibrated to how much genuine uncertainty exists about the variant's safety, not applied uniformly to every experiment regardless of risk.
As the head of the experimentation platform, create a 6-month plan to onboard 20 product teams onto a centralized platform while preserving statistical rigor and still letting teams move on their own schedule. Include the technical migration steps, training, templates, governance phases, success metrics for the rollout, and which platform features have to exist before you can onboard the first team.
Sample Answer
Direct answer
Sequence the 6-month onboarding so the platform's must-have features (self-serve config, automated validity checks, a shared metric catalog, and a set of standard templates) are proven with a small pilot group of teams before broad rollout, run technical migration and training in parallel tracks rather than serially, and measure success not just by "teams onboarded" but by whether those teams' experiments are actually meeting the statistical rigor bar, since onboarding a team that then runs sloppy experiments isn't real success.
Structured elaboration
- Months 1 to 2, pilot: onboard 3 to 4 representative teams (chosen to cover different risk profiles: a low-risk marketing team, a higher-risk payments-adjacent team) onto the platform, using this period to surface real gaps in the self-serve tooling, templates, and governance model before scaling to everyone.
- Templates: ship a small set of standardized, versioned templates so a team's onboarding lift is filling in a form, not inventing a process from scratch, an experiment-configuration template pre-wired to the shared metric catalog and default guardrails, a launch-readiness checklist template covering pre-registration and SRM setup, and a results-readout template that presents lift, confidence interval, and guardrail status in the platform's standard format. Teams that start from these templates inherit governance-compliant defaults automatically, which is a materially cheaper and more consistent onboarding path than each team rebuilding its own launch checklist and report format independently.
- Months 2 to 4, phased rollout: onboard the remaining teams in waves, each wave including migration of any existing homegrown experimentation the team was running (mapping their old ad hoc metrics into the shared catalog and their own launch process into the standard templates, not just their code), plus mandatory training on the platform's guardrails (why pre-registration matters, how to read a sequential-testing dashboard, how to fill out the launch-readiness template).
- Months 4 to 6, autonomy and governance maturity: shift from platform-team-assisted launches to genuine self-service for low-risk experiments using the standard templates, while the risk-tiered approval workflow (discussed elsewhere) takes over for anything higher-stakes, so the platform team's ongoing role becomes maintaining shared infrastructure and templates rather than hand-holding every launch.
- Required platform features before the FIRST team onboards: at minimum, self-serve configuration with pre-launch validation, SRM and basic guardrail monitoring, a shared metric catalog, and the experiment-configuration and launch-readiness templates; without these, the pilot teams would be testing an incomplete platform and the lessons learned wouldn't generalize.
- Success metrics: not just "number of teams onboarded" but the rate of properly pre-registered experiments, the rate of SRM-clean launches, template adoption rate (are teams actually using the standard templates or working around them), and (later) whether experiment velocity for onboarded teams actually increased versus their prior ad hoc process, since a slower but far more rigorous platform is a real trade-off worth being honest about, not something to hide behind a raw "teams onboarded" count.
Worked example
A team's existing homegrown pricing-experiment scripts define "conversion" slightly differently from the shared metric catalog. Migrating that team means more than pointing their launches at the new platform, it means an explicit reconciliation step where the team, together with the platform team, maps their historical definition to the standard one, documents where the two differ, and moves their launch process onto the standard experiment-configuration and results-readout templates, so historical experiment results and newly-onboarded ones remain comparable rather than silently using two different definitions of the same word and two different report formats.
Trade-offs and pitfalls
Rushing every team onto the platform in month one to hit an aggressive "teams onboarded" target risks skipping the reconciliation, training, and template-adoption work that actually makes the migration successful, producing teams that are nominally "onboarded" but still running experiments with the same rigor gaps they had before. The pilot-then-wave approach costs calendar time upfront but catches integration, governance, and template gaps while the blast radius is still small, which is cheaper than discovering them after all 20 teams have already migrated.
Unlock Full Question Bank
Get access to all Experimentation Platforms and Infrastructure interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.