Design Metrics and Impact Measurement Questions
Connecting design work to measurable outcomes: choosing success metrics, KPIs and guardrails for a specific change, writing measurable problem statements and testable hypotheses, and judging when quantitative data should and should not override design judgment. Covers instrumentation and tracking plans (event taxonomy, identity resolution across platforms, data quality and privacy constraints), reading funnels, cohorts, retention and adoption curves, and reconciling behavioral analytics with survey and qualitative signal. Experimentation is a large part of this topic: A/B, multivariate and quasi-experimental design, sample size and minimum detectable effect, stopping rules, multiple-comparison and confounding traps, and isolating a design's effect when a clean test is not possible. Also covers post-launch monitoring, rollback criteria and post-mortems, ROI and business cases for design work including design systems and design programs, and reporting impact to executives and stakeholders.
A senior PM requests statistically significant evidence for a checkout redesign affecting pricing, layout, and messaging. Draft an end-to-end experimental plan: primary hypothesis, segmentation or stratification approach, assumptions for a power calculation (baseline conversion, minimum detectable effect), monitoring and stopping rules, guardrails for negative impacts, and a rollout strategy if the experiment succeeds.
Sample Answer
Direct answer
I'd give the PM (product manager) an end-to-end plan, not just a metric name: one primary hypothesis with a pre-registered minimum detectable effect (MDE, the smallest change worth caring about), a stratified randomization so the pricing/layout/messaging bundle is tested fairly across customer segments, explicit guardrails against harm, and stopping rules decided before the test starts, not while watching the dashboard.
Structured elaboration
Primary hypothesis: "Bundling the redesigned pricing display, layout, and messaging on the checkout page increases checkout completion rate relative to the current design, without reducing average order value." Note it's framed as one bundled hypothesis, since the three changes ship together as a single redesign, not three separate levers; that has a real trade-off covered below.
Segmentation / stratification: randomize at the user level, but stratify the randomization by the two dimensions most likely to interact with the redesign: device (mobile vs. desktop, since layout changes often land very differently on each) and customer type (new vs. returning, since pricing-display familiarity differs). Stratifying means you guarantee balanced representation of each stratum in both arms rather than hoping random assignment gets there by chance, which matters when a stratum is a minority of traffic.
Power calculation assumptions: baseline checkout conversion rate 3.0%, and the team wants to detect an absolute lift of 0.3 percentage points (to 3.3%, a 10% relative lift), at a 5% two-sided significance level and 80% power. Using the standard two-proportion sample size formula:
n=(p2−p1)2(zα/22pˉ(1−pˉ)+zβp1(1−p1)+p2(1−p2))2zα/2 and zβ are just the standard-normal critical values for your chosen significance level and power (1.96 and 0.84 at the conventional two-sided 5% / 80% settings used here). Plain-English intuition for the formula itself: the numerator grows with how noisy the two conversion rates are, the denominator shrinks with how big a gap you're trying to detect, so a tiny gap in a noisy metric needs a much bigger sample. With p1=0.03, p2=0.033, this comes out to roughly 53,200 checkout sessions per arm (about 106,400 total). This is the concrete number that turns into a duration commitment once you know daily checkout traffic (e.g. at 5,000 sessions/day total, roughly 21 days minimum).
Monitoring and stopping rules: pre-register the sample size above and do not call the test before it's reached, even if the p-value dips below 0.05 earlier (early peeking inflates the real false-positive rate well past the nominal 5%). If the team wants interim looks, use a sequential testing correction (e.g. an alpha-spending boundary, a pre-set schedule for how much of the total 5% false-positive budget each interim look is allowed to spend, so repeated peeking can't quietly push the real error rate above the stated level) decided in advance, not ad hoc peeking. Stop early only for a guardrail breach, never for an early "win" on the primary metric.
Guardrails: (1) average order value, so a completion-rate win isn't secretly a discount-driven race to the bottom; (2) checkout error rate / page load time, so a layout change isn't winning by masking a performance regression that would show up elsewhere; (3) refund/chargeback rate over the following weeks, since pricing/messaging changes can shift what customers expect and drive returns.
Rollout strategy if it succeeds: ramp gradually (e.g. 50% to 75% to 100%) rather than flipping immediately, keep a small (2-5%) long-term holdback for a few weeks post-ramp to confirm the effect persists and isn't a novelty spike, and only fully retire the old checkout design after that holdback confirms stability.
Worked example
At 5,000 checkout sessions/day evenly stratified across mobile/desktop and new/returning, the test needs about 106,400 total sessions, or roughly 21-22 days at that volume. If the true lift really is 0.3pp and it holds, that's worth (at, say, a $60 average order value) an estimated $0.18 in incremental revenue per session exposed, or about $27,000/month at 150,000 monthly sessions once fully rolled out, before accounting for the guardrail checks passing.
Trade-offs & pitfalls
Bundling pricing, layout, and messaging into one variant means a win tells you the bundle works, not which piece does the work; if the team wants to attribute impact to a specific element for future iteration, a factorial design (testing the three changes both separately and combined) is more informative but needs substantially more traffic to reach the same power in every cell, often not practical at this baseline conversion rate. A second pitfall: don't let "returning customer" stratification hide a real segment-level harm, such as the redesign clearly helping new customers while quietly hurting returning customers who were used to the old pricing layout; always report segment cuts, not just the pooled result, before recommending a full rollout.
Describe the structure of a dashboard for monitoring design KPIs. For three audiences (design team, product managers, executives) list the top 4 widgets you'd include, data refresh cadence, and recommended drilldowns or filters for each audience.
Sample Answer
Direct answer
A single dashboard trying to serve executives, product managers (PMs), and the design team at once usually serves none of them well, because each audience needs a different altitude: executives want a trend and a dollar figure, PMs want a diagnosable funnel, and designers want screen-level detail they can act on this sprint. I'd build one dashboard with three views (or three separate dashboards fed by the same underlying data) rather than one view with everything crammed in.
Structured elaboration
Executives (top 4 widgets)
- North Star metric trend (e.g. conversion rate or active-user growth over the last 4-6 quarters), so the headline number is always in view.
- Quarter-over-quarter attributable design impact, in dollars or percentage points, aggregated across squads (only counting statistically validated experiment results, flagged separately from directional estimates).
- Major initiative status (RAG: red/amber/green) for the 3-5 largest design bets in flight, one line each.
- Customer sentiment trend (e.g. NPS, Net Promoter Score, a 0-10 "how likely are you to recommend us" survey rolled into a single score, or CSAT, customer satisfaction score, typically a single "how satisfied were you with X" question on a 1-5 or 1-7 scale, averaged across respondents) over time.
- Cadence: monthly, with a quarterly deep-dive. Drilldowns: minimal, maybe one click from "impact" to "which initiative drove it," since executives consume this dashboard, they don't investigate it live.
Product managers (top 4 widgets)
- Funnel conversion by step for their core flow, so a drop-off is visible at the step level, not just the aggregate.
- Active experiment status board: what's running, what's queued, what just concluded and its result.
- Feature adoption curve for recently shipped features (% of eligible users who've tried it, over time since launch).
- Top user drop-off segments, i.e. which cohort (device, acquisition channel, plan tier) is underperforming the aggregate, since that's usually where the next hypothesis comes from.
- Cadence: daily or near-real-time for the funnel and experiment board (PMs need to catch a broken experiment fast), weekly for adoption curves. Drilldowns: by segment (device, channel, cohort) and by date range, with the ability to isolate a single experiment's traffic.
Design team (top 4 widgets)
- Design system component adoption % across the product surface, since it's a leading indicator of consistency debt.
- Post-ship design defect trend (UI/UX bugs per release), to catch quality regressions early.
- Screen-level engagement snapshot (e.g. click/interaction heatmap summary or session recordings volume) for recently redesigned screens.
- Qualitative signal digest: a rolled-up count and theme breakdown of recent user feedback, support tickets, and periodic usability-session findings tagged to a screen, so quantitative and qualitative signal sit side by side rather than in separate tools.
- Cadence: daily for the defect trend (catch regressions fast), weekly for adoption and engagement. Drilldowns: by screen/component and by release version, so a designer can trace a defect spike back to a specific ship date.
Worked example
Take one underlying event ("checkout redesign shipped March 3") flowing into all three views differently: the executive view shows it as one line in the RAG initiative list, moving to green with "+0.4pp conversion, ~$24k/mo" once validated, and that dollar figure has to be re-derivable on the spot from inputs stored beside it or it does not belong on an executive tile at all: 200,000 monthly checkout sessions x 0.4pp = 800 extra completed checkouts a month, at a $30 average order value = 800 x $30 = $24,000/month; the PM view shows the checkout funnel's step-by-step conversion before and after March 3, letting them see the lift is concentrated at the payment-details step specifically; the design view shows the new checkout screen's defect count (2 minor UI bugs filed in week one, both closed by week two) and its component-adoption contribution (this screen now pulls 90% of its UI from the shared library, up from 40% pre-redesign).
Trade-offs & pitfalls
The most common mistake is giving executives PM-level drilldown widgets "just in case," which trains them to ask operational questions the dashboard wasn't built to diagnose and buries the one number they actually needed. The opposite mistake, giving designers only aggregate numbers with no screen-level cut, makes the dashboard unusable for deciding what to fix next. Keep the underlying data model shared across all three views (one source of truth) even though the widgets differ, or the three audiences will eventually see different numbers for what should be the same metric and lose trust in all three dashboards at once. The same discipline applies to every currency figure on the executive view: store the inputs (exposed sessions, average order value, the measured lift and its confidence interval) next to the number, because a dollar amount an executive cannot re-derive when challenged is the fastest way to lose the credibility the rest of the dashboard depends on, and a rounded-in-the-wrong-direction estimate is indistinguishable from a fabricated one once it is on a slide.
Design a measurement plan to evaluate whether a prototype causes sustained behavior change (for example, increasing weekly exercise frequency). Define immediate and long-term metrics, cohorting and retention strategy, instrumentation events, experiment duration, and analysis techniques to control for confounders and attribution.
Sample Answer
Direct answer
Measure two different things on purpose: whether people did the target behavior at all right after exposure, and whether that behavior actually persists once the novelty wears off. Most of the risk in a behavior-change measurement plan is mistaking an initial spike for durable change, so the design has to be built specifically to detect decay, not just a single after-the-fact snapshot.
Structured elaboration
- Immediate metrics: activation of the prototype's core feature in week 1 (for example, logging a workout through it) and completion rate of the first prompted session.
- Sustained metrics: weekly exercise frequency (sessions per week) tracked per user for 12 weeks, defined as the actual primary outcome, plus a plateau check like the share of users still logging at least 2 sessions per week during weeks 8 through 12.
- Cohorting and retention strategy: track weekly signup or exposure cohorts on their own retention curves, comparing each cohort at the same number of weeks since exposure rather than the same calendar week, so a cohort exposed during a January motivation spike isn't compared unfairly against one exposed in March.
- Instrumentation: explicit events like
workout_logged {user_id, timestamp, source, duration_min}andprompt_shown/prompt_dismissed, so "used the prototype" can be separated from "exercised at all." This matters because the prototype could simply move logging from an old path to a new one without any real increase in exercise, guard against that by also tracking total exercise frequency across all sources, not just prototype-attributed activity. - Duration: run at least 12 weeks with weekly check-ins. Behavior-change products commonly show an early spike followed by decline in just the first couple of weeks, a 2-week study would only ever catch the spike.
- Confounders and attribution: compare the treatment cohort's own pre-exposure trend against its post-exposure trend using a difference-in-differences approach, with a matched control group that never gets the prototype providing the counterfactual trend for the same calendar weeks, which absorbs shared confounds like weather or a company-wide wellness campaign hitting everyone at once. Where a clean randomized holdout isn't feasible, a regression-discontinuity design (comparing outcomes for users just above versus just below some pre-existing eligibility cutoff, treating which side of the cutoff someone lands on as effectively random) or an instrumental-variable approach (finding a third factor that affects exposure to the prototype but has no direct effect on exercise frequency itself, then using that factor to isolate the prototype's own effect) is a fallback, but a true randomized holdout of eligible users for the full study window is the cleanest option when it's available.
Worked example
Assume the control group's weekly exercise frequency plateaus around 1.8 sessions per week by week 8, with a standard deviation of 1.2 sessions across users, and the goal is to detect a sustained lift to 2.2 sessions per week at 95% confidence and 80% power:
n=δ22(zα/2+zβ)2σ2With σ=1.2, δ=0.4, and using the standard z-score lookup values for 95% confidence and 80% power (zα/2=1.96 and zβ=0.84, so zα/2+zβ=1.96+0.84=2.8): n=0.162(2.8)2(1.44)≈141.1, so 142 users per arm who are still reporting data at week 8. A 12-week study realistically loses participants to non-reporting or dropout, assuming 40% cumulative attrition by week 8, enroll with margin: 142/0.6≈237 users per arm at week 0, about 474 total enrolled.
Trade-offs and pitfalls
Measuring only prototype-attributed activity overstates impact if it's cannibalizing an existing logging path rather than creating new behavior, the total-activity guardrail above exists specifically to catch that. Running the study too short mistakes a novelty spike for durable change. And withholding a wellness feature from a randomized holdout group for 12 weeks raises a real fairness question in some contexts, a senior answer should name that trade-off explicitly (even if the resolution is "acceptable because participation is opt-in for a beta") rather than assume it away.
You need to produce a one-page summary for execs showing the impact of a recent design change. What elements do you include on that page (metrics, visuals, context) and how would you communicate uncertainty and the recommended next steps?
Sample Answer
Direct answer
The one-pager needs to work as both a launch narrative and a living reference: a headline result stated in plain business terms first, the primary metric with its size of effect and how confident we are in it, one or two supporting metrics, a guardrail check confirming nothing else broke, and a clear recommendation for what happens next. I'd design it as a dashboard-style structure that keeps getting reused as the launch matures, not a one-time slide that gets forgotten the week after it's presented.
Structured elaboration
Elements to include:
- A headline banner in business language ("Checkout redesign lifted revenue-driving conversions, no negative impact on support volume") rather than a metric name and a percentage as the first thing anyone reads.
- The primary metric shown as a before/after comparison with the size of the change and a plain-language confidence statement, not just a bare percentage.
- One or two supporting metrics that corroborate the primary result (for a conversion lift, something like time-to-checkout or drop-off rate at the step that changed).
- A guardrail line confirming the metrics that shouldn't move, didn't move badly (error rate, support tickets, adjacent-flow performance).
- A small trend chart showing the metric over the surrounding weeks, so the audience can see the change wasn't a blip before or after the launch window.
- A short qualitative note or representative user quote, since a number alone rarely persuades a non-design audience as effectively as a number paired with a human reason.
- A clearly labeled recommendation and next step (ship to 100%, iterate further, hold and investigate).
Communicating uncertainty: state the confidence interval and the sample size or duration behind it in plain language rather than showing a raw statistical output, for example "we're 95% confident the true lift is between 0.08 and 0.52 percentage points, based on 100,000 sessions over two weeks" rather than just "z = 2.68, p < 0.01." If the interval is wide, say so plainly rather than letting the headline point estimate imply more precision than the data supports.
Structuring it as an ongoing dashboard, not just a launch snapshot: the same layout, headline banner, metric tiles, trend chart, guardrail row, doubles as a recurring reference stakeholders can revisit weeks later to see whether the initial win held up, rather than a static slide that gets forgotten. Treating the one-pager as a durable dashboard rather than a one-time announcement is what makes the "successful launch" story stay true past the first week, since a redesign that looks great on day 3 but decays by day 30 is a very different result than one that holds.
Worked example
Checkout redesign, two-week window, 50,000 sessions per arm: baseline conversion 3.10% (1,550 conversions), post-change conversion 3.40% (1,700 conversions).
The lift and its interval both have to be carried on the same basis, percentage points, or the arithmetic silently stops meaning anything:
diff=3.40%−3.10%=0.30 percentage points SEdiff=n1p1(1−p1)+n2p2(1−p2)=50,0000.031×0.969+50,0000.034×0.966≈0.00112That 0.00112 is a proportion, which is 0.112 percentage points; converting it before it meets the 0.30 is the whole trick.
95% CI=0.30±1.96×0.112=0.30±0.22=(0.08, 0.52) percentage pointsThat is a 0.30 percentage-point absolute lift and, separately, a relative lift of about 9.7% (0.30 divided by 3.10). Those are two different bases and the page has to label which one it is quoting: 9.7 percentage points would be roughly thirty-two times the actual 0.30-point lift, and an exec who reads the relative number as an absolute one has misjudged the result by that factor. The confidence interval doesn't cross zero (the same standard error gives z=0.0030/0.00112≈2.68), so the headline can honestly read "conversion lift, 95% confident the true improvement is between roughly 0.1 and 0.5 percentage points" rather than a single overconfident number.
Trade-offs and pitfalls
The biggest trap is choosing metrics for the one-pager that flatter the result rather than test it; if the guardrail row is quietly dropped because nothing interesting happened there, a skeptical stakeholder should assume the omission was deliberate, not that everything was fine. The second trap is presenting a point estimate with no interval at all, which reads as more certain than the underlying data justifies and sets up an awkward conversation later if the effect doesn't hold at full rollout.
Explain Customer Effort Score (CES): how to collect it, the typical scoring formula, when it is most useful, and provide an example of how a PM would act on CES results for a support-heavy product. Discuss sampling cadence and how to correlate CES with retention or churn.
Sample Answer
Direct answer
Customer Effort Score (CES) measures how much effort a customer felt they had to expend to get something done, typically resolving a support issue, framed as a single-item survey ("The company made it easy for me to handle my issue") on a 1-7 (or 1-5) agreement scale. It's most useful right after a specific transaction, not as a general relationship health check, and it's especially actionable in a support-heavy product where high effort is a direct, fixable predictor of churn.
Structured elaboration
Collection: trigger the survey immediately after a defining event, a support ticket closes, a return is processed, an onboarding step completes, so the respondent's memory of the effort involved is fresh. A single Likert-scale question ("It was easy for me to handle my issue," strongly disagree to strongly agree) is standard.
Scoring formula: CES is simply the mean of the numeric responses across a sample:
CES=n1i=1∑nscoreiPlain-English intuition: it's an average opinion score, so a 7-point scale centered around, say, 5.5 tells you "on average, people found this moderately easy," and it only becomes actionable once you break it down by issue type rather than reading the single blended average.
State the scale direction before anyone compares two CES numbers, because two conventions are in circulation and they run in opposite directions. The modern agreement-statement form used throughout this answer ("The company made it easy for me to handle my issue," 1 = strongly disagree through 7 = strongly agree) makes a HIGHER score better. The older effort-rating form ("how much effort did you personally have to put forth?") makes a higher score WORSE. They are not interchangeable: the same underlying experience scores 2 on one and 6 on the other, and the sign of any correlation with churn flips with it. Put the convention on the dashboard axis label, or two teams will read the same 5.8 as "this is going well" and "this is our worst category."
When it's most useful: CES shines for transactional, single-interaction friction (a support call, a returns process, an account-setup step), not for overall brand sentiment, that's closer to what Net Promoter Score (NPS, a "how likely are you to recommend us" 0-10 question) is trying to capture. Using CES where NPS belongs (or vice versa) is a common misapplication.
Example PM (product manager) action for a support-heavy product: suppose a subscription product's support team logs CES per issue category. If "billing disputes" scores meaningfully worse (say a mean of 2.2 on the ease-agreement scale defined above, where higher is better) than "password resets" (mean 5.9), the PM's action isn't "improve support overall," it's specifically: audit the billing-dispute flow, look for redundant identity verification steps or unclear dispute status communication, and consider a self-service dispute-status page to cut the effort involved, since that category is both high-effort and (per the section below) has real customer-value stakes attached to it.
Sampling cadence: for a support-heavy product, survey after every closed ticket if ticket volume is modest (say, under a few hundred/week), since you want category-level granularity; at higher volumes, sample a consistent percentage (commonly 20-30%) to avoid survey fatigue while still getting a large-enough sample per issue category to trust the category-level mean.
Correlating CES with retention/churn: track, for each customer who interacted with support in a given month, their CES score and whether they churned in the following period. Comparing mean CES between the churned and retained groups (or computing a simple correlation between CES and a churn indicator) tells you how strongly effort predicts attrition for this product specifically; that relationship should be re-checked periodically, since it can weaken as competitors' switching costs change.
Worked example
Take an illustrative sample of 10 customers who contacted support last month, each rated on the ease-agreement scale defined above (1 = strongly disagree it was easy, i.e. very high effort; 7 = strongly agree it was easy, i.e. very low effort), tracked for whether they churned the following quarter:
| Customer | CES (1=high effort, 7=easy) | Churned next quarter? |
|---|---|---|
| 1 | 6 | No |
| 2 | 5 | No |
| 3 | 2 | Yes |
| 4 | 3 | Yes |
| 5 | 6 | No |
| 6 | 1 | Yes |
| 7 | 7 | No |
| 8 | 4 | No |
| 9 | 2 | Yes |
| 10 | 5 | No |
The four customers who churned have a mean CES of (2 + 3 + 1 + 2) / 4 = 2.0; the six who stayed have a mean CES of (6 + 5 + 6 + 7 + 4 + 5) / 6 = 33 / 6 = 5.5. The point-biserial correlation (a version of the standard correlation coefficient used when one variable is continuous, here CES's 1-7 score, and the other is binary, here churned yes/no coded as 1/0; it produces a number on the same -1-to-+1 scale as any correlation, where 0 means no relationship, +1 a perfect positive relationship and -1 a perfect inverse one) between CES and churn across all 10 works out to roughly r = -0.89: a strong NEGATIVE association in this toy sample, meaning low ease scores (high effort) travel with churning. The sign is the part to get right and it is purely a property of the convention you collected on, not of the customers: recorded on the older high-effort-is-a-high-number scale, the identical ten experiences would read r = +0.89 and mean exactly the same thing. Report the sign together with the scale direction, or the number is unreadable. That's the shape of evidence ("high-effort support interactions cluster with churn") a PM would use to justify prioritizing effort-reduction work, though a real analysis would need a much larger sample and should control for other churn drivers (price sensitivity, usage level) before claiming CES itself is causal.
Trade-offs & pitfalls
The main pitfall is treating CES as a churn-prediction model on its own; it's a strong correlational signal, not a controlled causal estimate, so a customer with high effort and low product usage might churn for the usage reason, not the effort. A second pitfall is over-surveying: firing a CES prompt after every single interaction (rather than after meaningfully effortful ones) causes response fatigue and depresses response rates, which then biases the remaining respondents toward people with unusually strong (often negative) opinions.
Unlock Full Question Bank
Get access to all Design Metrics and Impact Measurement interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.