Design Critique, Iteration, and Decision Rationale Questions
Improving a design through structured feedback and defending the reasoning behind it: giving and receiving critique, facilitating design reviews, running iteration cycles that turn feedback into shipped changes without losing the design's intent, and justifying a specific decision with evidence rather than taste. Covers holding a decision under pushback from a product manager, engineer, or executive, reframing a request driven by opinion or a vanity metric, deciding what to do when the evidence is contradictory, thin, or blocked by an external constraint, writing the rationale into a decision log so it survives past the meeting it was made in, and running a post-mortem when a shipped decision does not produce the result that was predicted. The evidence itself is an input here, not the subject: how to run user research and usability tests, how to define success metrics, accessibility standards, and design-system governance belong to other topics.
When you're building a case for a design decision, qualitative and quantitative evidence pull different weight. Walk through the strengths and limits of each, then give a concrete example of combining a qualitative insight with quantitative validation to take something from prototype into production.
Sample Answer
Direct answer
Qualitative evidence tells you why people behave a certain way, and quantitative evidence tells you how much that behavior matters at scale. A strong case for a design decision almost always combines both, because relying on evidence-based design (grounding a decision in what users show or tell you) instead of opinion-based design (a decision from personal taste or precedent) requires knowing both the reason and the size of the effect.
Structured elaboration
Where each type of evidence comes from
- Qualitative: interviews, moderated usability tests, session replays (recorded interactions you can watch back), and open-ended survey responses. Example: an interview reveals users feel overwhelmed by a long form; a usability test shows exactly where they stall on it.
- Quantitative: analytics, such as funnel drop-off or click-through rate, and closed-ended survey results at scale. Example: analytics show a large, consistent drop on a specific step; a survey shows most respondents want faster access to a feature.
Strengths and limits
| Qualitative | Quantitative | |
|---|---|---|
| Strength | Explains why, surfaces problems you did not know to look for | Tells you how much, and whether it holds at scale |
| Limit | Small samples, hard to generalize or compare options numerically | Tells you what happened but not why; can miss edge cases and nuance |
Worked example
Take a prototype for a new onboarding step. Moderated usability sessions with a handful of users surface a specific confusion, why sign-up asks for a phone number before explaining what it is for, and that is the qualitative insight: it tells you what to fix and why. Before rolling the fix out broadly, run a short controlled comparison in production against the current flow, measuring completion rate on that step, and that is the quantitative validation: it tells you whether the fix actually moves the number at scale, not just in a small room. Only once both line up, a specific problem identified and a real lift measured, do I treat the decision as validated enough to keep in production; if the metric does not move, that is a sign the interview finding was not the whole story, and it sends me back to figure out what else is going on.
Trade-offs and pitfalls
Quantitative data with no qualitative context tells you something moved but not why, so the next iteration is a guess. Qualitative insight with no quantitative check risks over-indexing on a vocal minority, especially since the people who agree to a research session are not a random sample of your users. The common failure is treating whichever evidence type is more convenient, a stakeholder's favorite metric or a single compelling user quote, as sufficient on its own.
You have a synthesis document where participant quotes directly contradict one another and telemetry shows high variance across sessions. As the lead designer, describe the synthesis process you would run to reduce ambiguity: how you'd code and weight evidence, resolve contradictions, present uncertainty to stakeholders, and decide on a bounded set of follow-up design experiments.
Sample Answer
Direct answer
I'd treat apparent contradiction as a segmentation signal first, not noise to average away. Most quotes that seem to conflict resolve once you check whether they're actually coming from different user segments or different contexts, rather than genuinely opposed opinions from the same kind of user in the same situation.
Structured elaboration
Coding evidence. Build a codebook with explicit, shared definitions, and code with at least two people so individual bias doesn't drive the categorization. Tag every quote with the participant's segment (device, tenure, role) and session context, not just its content, because the segment tag is what turns "contradiction" into "explanation" later.
Weighting evidence. Coding tells you what each piece of evidence SAYS; weighting decides how much it counts, and skipping it is how a synthesis ends up ruled by whoever was most quotable. Weight each piece on four things, recorded as a column next to the code, not held in someone's head:
- Directness. Observed behaviour outranks self-report, and self-report outranks recalled report. Watching someone fail a task is stronger evidence than them telling you afterwards that it was confusing, which is stronger than them remembering last month.
- Independence. Two participants recruited from the same customer, the same forum thread, or the same support escalation are close to one data point, not two. Count independent SOURCES, not raw quote volume, or a single loud account gets counted five times.
- Task relevance. A comment made while doing the real task in scope beats an unprompted aside about a different part of the product, even when the aside is more vivid.
- Corroboration across method. A code that also moves a telemetry number weighs more than one that exists only in the transcripts, because the two methods have different blind spots.
In practice that means a coded theme carries a weight, not just a count, and the write-up reports both, "12 quotes from 9 independent accounts, 7 of them observed rather than reported, corroborated by telemetry," so a reader can see why one theme outranks another instead of assuming the bigger tally wins. The one weighting input to refuse is seniority of the speaker: how senior a participant is inside their own company says nothing about how representative their experience of the interface is.
Resolving contradictions. Split apparent conflicts into three buckets: segment-driven (different users, both correct for their own segment), context-driven (the same user, describing a different task or moment), and genuine noise (a one-off, small-sample artifact). Then re-cut the telemetry PER SEGMENT rather than pooled; high variance in a pooled number is itself often a clue that two distinct groups are being averaged together.
Presenting uncertainty. Attach a confidence label (High, Medium, Low) to each finding based on how many independent sources agree and whether it holds up once segmented, and say explicitly what's still unknown rather than smoothing everything into a single confident story.
Deciding on bounded follow-up experiments. Prioritize by impact times confidence times feasibility, and pick a small, named set (two to three), each designed to resolve one specific open contradiction rather than a general "let's research more."
Worked example
A synthesis document has 24 participant quotes about a new filter interface: 9 call it confusing, 8 call it a clear improvement, 7 are neutral. Pooled telemetry shows filter usage ranging from 12% to 61% across sessions, high variance that reads, at first glance, as noisy and hard to trust.
Segmenting by device changes the picture entirely. Of 14 desktop participants, average filter usage is 54%; of 10 mobile participants, it's 18%. Re-checking the quotes against device: 8 of the 9 "confusing" quotes came from mobile participants, and 7 of the 8 "improvement" quotes came from desktop participants. This isn't a genuine contradiction, it's two separate, true stories: the redesign works on desktop and regresses on mobile, most likely a tap-target or visibility problem on small screens.
The weights, not just the counts, are what make the mobile finding the stronger of the two. The 8 mobile "confusing" quotes come from 8 separately recruited participants rather than several people at one account, 6 of the 8 are tied to an observed failure in the session rather than a general opinion, and the theme is corroborated by an independent method, the 18% usage figure from telemetry. The desktop "improvement" theme has 7 quotes but they are almost all self-reported preference at the end of the session rather than observed behaviour, and its telemetry corroboration is the same number that a novelty effect would also produce. Same order of quote count, materially different weight, and that gap is exactly what the confidence labels below are recording.
Confidence levels: High that mobile has a real usability problem (8 of 10 mobile participants plus the telemetry both point the same direction). Medium that the desktop improvement is causal rather than a short-term novelty effect (it needs a controlled comparison to confirm). Two bounded follow-ups are prioritized: a moderated usability test with 6 mobile participants targeting the filter's tap targets and visibility, and a controlled experiment on desktop that runs long enough to rule out a novelty effect before crediting the redesign with the improvement.
Trade-offs and pitfalls
- Averaging contradictory quotes into a single "mixed signal" summary is the most common and most damaging shortcut; it erases exactly the segment split that would have made the finding actionable.
- It's easy to over-trust the most articulate or most senior-sounding quotes over a quieter majority saying something less quotable but more representative. Recording a weight per coded theme, rather than only a count, is the concrete defence: it forces the question "how many INDEPENDENT sources is this, and did we watch it happen or were we told about it" every time, instead of once, at the end, when the narrative has already formed.
- Running too many follow-up experiments at once dilutes attention and budget; a bounded, prioritized set beats a broad research agenda every time.
- Some contradictions genuinely don't resolve even after segmenting; naming that honestly as unresolved is more useful to stakeholders than forcing a false resolution to look thorough.
Describe a concrete example of an iteration that later proved to have failed because of confirmation bias or p-hacking. Explain what signals indicated the failure, how the team recognized and surfaced the issue, and propose concrete changes to research and experimentation practices to prevent similar mistakes in the future.
Sample Answer
Direct answer
Here's a realistic shape this failure takes: a team runs a test, the result is a small, not-quite-significant lift, and instead of accepting an inconclusive result, someone extends the test past its planned window and slices the data by a new segment until one slice crosses the significance threshold, then reports that slice as the finding. Confirmation bias (the tendency to notice and trust evidence that matches what you already expected, while explaining away evidence that doesn't) sets the motive; p-hacking (reshaping an analysis, through extra time, extra segments, or dropped outliers, until something crosses a "statistically significant" threshold by chance, then reporting only that result) is the mechanism. Both surface the same way after the fact: a shipped change that doesn't hold up.
A concrete version of the failure
A design team redesigns onboarding and runs an A/B test (a controlled comparison of the new flow against the old one) with a pre-planned two-week window, a pre-registered significance bar of p < 0.05, and a single primary metric, day-7 activation rate. Traffic to the flow runs about 5,000 users per arm per week, so the planned readout lands on roughly 10,000 users per arm. At two weeks, day-7 activation comes in at 41.2% for the treatment arm versus 40.1% for control, a 1.1 percentage-point lift. Put through a two-proportion z-test at that sample size, that gap gives z = 1.58 and p = 0.11, short of the pre-agreed 0.05 bar. Instead of shipping "inconclusive, here's what we'd test next," the team extends the test for another week without writing that decision down anywhere, and someone also starts slicing the data by device type "just to look."
The extra week is worth watching on its own. At three weeks the arms have roughly 15,000 users each, the gap is still about the same size, 41.5% versus 40.4%, and the bigger sample alone pulls the p-value from 0.11 down to 0.053. It still has not crossed, and that near-miss is exactly what makes the next move feel reasonable. The device cut is where it gives way. Mobile is about 60% of the sample, roughly 9,000 users per arm, and mobile activation reads 44.4% for treatment against 42.8% for control, a 1.6 percentage-point lift at p = 0.03. The remaining 40% on desktop, about 6,000 per arm, reads 37.15% versus 36.8%, a 0.35 point difference at p = 0.69, which is nothing. The two slices do blend back to the overall three-week numbers, so nothing looks off on inspection: 0.60 x 44.4 + 0.40 x 37.15 = 41.5, and 0.60 x 42.8 + 0.40 x 36.8 = 40.4. That mobile slice becomes the headline of the launch readout; the extension and the segment cut are never mentioned, and the team convinces itself the mobile result is what they expected all along, which is confirmation bias explaining why nobody questioned it in the room.
Notice how little work it took. One unplanned week plus one unplanned two-way split moved a result from p = 0.11 to p = 0.03 without any new effect coming into existence. That is the entire mechanism: every additional look is another chance for noise to line up. A rough feel for the cost, treating the two device slices as independent tests, which flatters the team because in reality they share the same underlying data: two shots at a 0.05 threshold give a false-positive rate of 1 - 0.95^2, about 9.8%, and adding the unplanned extension as a third look takes it to 1 - 0.95^3, about 14%. The threshold on the readout still says 0.05. The real one is nearly three times that.
How the team recognized it
The tell came four to six weeks after full rollout. At full traffic that window covers roughly 50,000 users, compared against a matched pre-launch window of about the same size, and day-7 activation across all users sat at 40.3% against the pre-launch baseline of 40.1%: a 0.2 point difference, p = 0.52. The p-value is not the useful part; the interval behind it is. The 95% confidence interval on that difference runs from about -0.4 to +0.8 percentage points, and a genuine 1.6 point mobile lift on 60% of users would have to show up as roughly a 0.96 point lift overall, which sits outside that interval. The post-launch check carried about five times the per-arm sample of the planned two-week readout, so this is not an underpowered null that failed to detect a real effect. It is a direct contradiction of the "it worked" story from the readout.
A skeptical team member then pulled the original experiment plan and compared it line by line to what actually happened: the planned duration didn't match the actual one, the planned primary metric wasn't the one that shipped in the readout, and the mobile-only segment was never named as a hypothesis before the data existed. None of the three matched, which is the specific pattern p-hacking leaves behind: a result that only exists because the analysis kept changing until something looked significant.
Surfacing it turned out to be a separate problem from spotting it, and it is the half teams get wrong. What worked was taking the plan-versus-readout comparison back to the same forum where the win had been announced, rather than raising it privately, and presenting it as a process failure rather than as somebody's mistake: the plan allowed an undocumented extension, and the readout template had no field for "what changed after launch," so the gap was invisible by design. The claim was then explicitly retracted rather than left to quietly age out of the roadmap, and the mobile hypothesis was re-entered in the backlog as an untested idea. Framing it as "our process let this through" is what makes the next person willing to raise the same flag; framing it as "who wrote this readout" is what guarantees nobody does.
Changes to prevent it recurring
- Pre-register before launch: write down the primary metric, the planned duration or sample size, and the stopping rule in a shared doc before anyone sees results, not after.
- Hold to the stopping rule: review at the planned checkpoint and decide then, rather than checking daily and extending whenever the number looks close.
- Treat any subgroup or exploratory finding discovered after the fact as a new hypothesis, to be confirmed with its own fresh experiment, never shipped as a conclusion from the same data that generated it.
- Add a second reviewer, a researcher or a peer designer who wasn't running the test, to read the write-up against the original plan before a launch decision is made, specifically checking whether anything (the metric, the duration, the segment) changed after the data came in.
- If the team genuinely needs to look early or to cut by segment, budget for it in the threshold instead of taking the extra looks for free: pre-declare the segments you will examine and tighten the bar accordingly (with two device slices, roughly 0.025 each rather than 0.05), or use a sequential design with a spending rule that permits interim peeks at a stated cost. The point is not that extra looks are forbidden, it is that they have a price and it should be paid up front.
Trade-offs and pitfalls
Extending a test or looking at a subgroup isn't wrong by itself; both are legitimate exploratory moves. The failure is presenting an exploratory finding with the confidence of a pre-planned, confirmatory one. The real trade-off is speed against rigor: pre-registration and a fixed stopping rule cost time and will feel like friction under launch pressure, and teams will be tempted to peek early and act on whatever looks good that week. Weigh that friction against what it buys, though. The cost of the failure above was not one bad onboarding flow; it was that every earlier result the team had shipped became suspect, because none of them had a plan on file to check against. The cultural fix matters as much as the procedural one: a team has to reward "this was inconclusive, here's what we'd check next" as a legitimate, even respected outcome, not treat every test that doesn't find a lift as a failure to be quietly reworked until it does.
That is every published Design Critique, Iteration, and Decision Rationale question for Design Researcher so far. Browse the other topics in this category, or practice this one interactively.