Design Critique, Iteration, and Decision Rationale Questions
Improving a design through structured feedback and defending the reasoning behind it: giving and receiving critique, facilitating design reviews, running iteration cycles that turn feedback into shipped changes without losing the design's intent, and justifying a specific decision with evidence rather than taste. Covers holding a decision under pushback from a product manager, engineer, or executive, reframing a request driven by opinion or a vanity metric, deciding what to do when the evidence is contradictory, thin, or blocked by an external constraint, writing the rationale into a decision log so it survives past the meeting it was made in, and running a post-mortem when a shipped decision does not produce the result that was predicted. The evidence itself is an input here, not the subject: how to run user research and usability tests, how to define success metrics, accessibility standards, and design-system governance belong to other topics.
Tell me about a time in your work or portfolio when usability testing led you to reverse a design you initially thought was best. Use the STAR method: describe the Situation, Task, Action, and Result. Focus on the evidence gathered, the trade-offs you communicated, and the measurable impact after the change.
Sample Answer
Direct answer
Strong candidates use this question to show they can update their own belief when evidence disagrees with it, not to show off a big win. Pick a real story where you had genuine conviction, the test data actually contradicted it, and you can say exactly what changed your mind and what the after-state measured.
Structured elaboration
STAR stands for Situation, Task, Action, and Result: Situation sets the context, Task is what you were trying to achieve, Action is what you specifically did, and Result is the measurable outcome. For this question, weight Action and Result heavily. The interviewer is really asking three things: how much evidence did you have and was it enough to trust, what did the team give up by reversing course, and did the thing you promised actually move after the change.
Worked example
Situation. I designed a compact summary card for an account dashboard, prioritizing scannability by hiding secondary details behind a "see more" toggle. Going in, I believed it was the stronger layout.
Task. Validate the design before it shipped into the shared component library, since other product teams would reuse the pattern.
Action. I ran a moderated usability test (a facilitator watches and interviews people using the product live) with 8 participants doing a realistic task: find whether there was a pending charge on the account. 6 of the 8 participants never found the pending-charge indicator at all, because it was hidden behind the toggle; among the 2 who did find it, time-to-find averaged about 30 seconds. I ran the older, more exposed layout as a counterbalanced second condition inside the same sessions rather than quoting a remembered number for it, and there 8 of 8 found the charge, averaging roughly 10 seconds. Both of those averages are over successful finds only, and I say so whenever I quote them: a task nobody completes has no completion time, so an unqualified "average time-to-find" silently drops the failures and flatters whichever design failed more. I brought the recordings and the count, 6 of 8 missed it, to the team instead of my impression, and proposed two options: keep the compact layout and add a persistent status indicator outside the toggle, a small change, or revert to the fully expanded layout, a bigger visual cost but zero discoverability risk. I recommended the persistent indicator as the smaller, testable change.
Result. After adding the indicator, a follow-up test with 8 new participants, run under the same moderated protocol and the same task wording, found 7 of 8 could locate the pending charge, averaging about 12 seconds among those 7. The headline I reported was the find rate, 2 of 8 to 7 of 8, not the seconds, and the reason is worth stating because it is the part people get wrong: matching the participant count across rounds is not the same as matching the basis. The 30-second before average is computed over only the 2 people who succeeded, the fastest and luckiest slice of the round, while the 12-second after average is computed over 7 of 8. Read naively, that comparison actually understates the improvement, since the 6 who never found the charge contribute no time to the before number at all. If you want a timing comparison that is genuinely on one basis, either cap every unsuccessful attempt at the task time limit and average over all 8 in both rounds, or report the find rate and quote the timings only for the finders with the denominator attached, which is what I did. The team adopted the indicator into the shared component library.
Trade-offs and pitfalls
- The temptation is to tell this story as "I was wrong, I'm humble," which undersells the actual skill being tested: reading evidence rigorously enough to change a real decision, not admitting fault.
- Be honest about sample size. 6 of 8 is a real, useful signal for a usability problem this consistent, but it is not the same as proof from a large-sample A/B test (an experiment that shows two versions to different groups of real visitors and compares an outcome metric at scale). Say so if asked, rather than inflating a small study into "we proved."
- A design that loses a usability test is not automatically wrong everywhere. Name what you kept from your original idea, such as the compact scannability goal, and what specifically needed a different solution.
Design a decision-log and design-rationale repository to scale collaboration across multiple product teams. Describe the data model (fields such as problem statement, options considered, owners, status), information architecture, workflows for creating and updating entries, permissions, integrations with design and project tools (Figma, Notion, Jira), and strategies to encourage adoption.
Sample Answer
Framing
A decision-log and design-rationale repository only earns its keep if logging a decision is cheaper than skipping it, and finding a past decision is cheaper than re-litigating it from scratch. The design has to balance real traceability (so a decision can be traced from research through to what actually shipped) against enough lightweight governance to keep entries trustworthy, without becoming a compliance chore that teams route around.
Data model
| Field | Purpose |
|---|---|
| Decision ID, title | Unique reference, human-readable name |
| Problem statement | What question this decision answers |
| Options considered | Each option with its pros and cons |
| Chosen option and rationale | What was picked and why, versus the alternatives |
| Evidence links | Research artifacts, prototypes, experiment results, tied to this specific decision |
| Owner, approvers | Who's accountable, who signed off |
| Status | Proposed, Accepted, Deprecated, Revisited |
| Risk or impact tier | Drives which governance gate applies (see below) |
| Related artifacts | Figma component IDs, issue-tracker ticket links |
| Review-due date | When this decision should be revisited by default |
| Audit log | Who changed what, and when |
End-to-end traceability. A decision entry shouldn't just record an opinion; it should link forward through the full chain: the research artifact that motivated it, the hypothesis it was meant to test, the prototype built from it, the experiment and its result, and finally the shipped implementation (a pull request or release reference). A reviewer should be able to walk from "why did we test this" all the way to "what's actually live" using one decision ID, not four disconnected documents.
Information architecture
Top-level index by product area, plus a global feed of decisions above a chosen risk tier so cross-team patterns are visible without hunting. Faceted search by tag, status, owner, and date. Two entry templates: a light one for low-risk, local decisions, and a fuller one required for anything that needs a cross-team gate.
Workflows and governance gates
Roles: a Design Council (cross-product senior designers, meeting on a regular cadence) that approves shared-system-level and cross-team decisions; a Product Design Lead per team who owns local decisions and the local pattern registry; and any Author who proposes a new entry. Gate structure for CREATING an entry: low-risk, single-team decisions self-publish after a peer sign-off; anything cross-team or brand-level requires Design Council review within a defined window (48 to 72 hours) before its status flips to Accepted. UPDATING an existing entry follows the same gate at a lighter weight: any team member can propose a status change to Revisited with a short reason, but flipping it to Deprecated or materially changing the chosen option re-enters the same review path the original decision took, so a quiet, undocumented reversal can't happen.
Permissions. Read access to the repository is open to the whole design and product org by default, since discoverability is the entire point. Write access is scoped by role: any Author can propose or comment; only the entry's assigned Owner or a Design Council member can flip its status; only Council members can approve a cross-team gate. Contractors and external partners get read-only access unless explicitly added to a specific product area.
flowchart LR
A[Author drafts entry: problem, options, chosen option] --> B{Risk tier}
B -->|Low risk| C[Peer sign-off, status Accepted]
B -->|Cross-team or brand-level| D[Design Council review, 48 to 72 hour window]
D -->|Approved| C
D -->|Sent back| A
C --> E[Linked to Figma component and issue tracker ticket]
E --> F[Implementation ships]
F --> G[Quarterly audit sample]
G -->|Gap found| H[Remediation entry, links back to original decision]
G -->|Clean| I[Stays in searchable index]
Enforcement versus local adaptation. A repository that's too strictly enforced becomes a compliance chore teams quietly route around in side channels; one with no enforcement at all becomes an unmaintained wiki nobody trusts. The resolution is to enforce at the point of highest leverage rather than everywhere: require a decision-log link on the ticket template for cross-team component changes specifically, rather than gating every single ticket. Pair that lightweight, always-on check with a low-frequency but real audit (a quarterly sample, not every entry) to catch drift, while letting local teams adapt their own templates for low-risk work without full Council sign-off.
Worked example
Trace one real decision through the system end to end. Decision ID DEC-214, title "Consolidate three card-hover patterns into one shared component." Problem statement: three product teams independently built slightly different hover treatments for list-item cards (one dims the background, one raises a shadow, one does both), so the same interaction reads differently depending on which team shipped it. Options considered: (a) leave all three as-is and just document the difference, (b) pick one team's existing pattern as the standard, (c) design one new shared pattern informed by usability data from all three. Chosen option and rationale: option (c). Rejecting option (b) required evidence on all three candidates, not just some of them, so the sessions covered every existing pattern rather than stopping at two. Both single-signal patterns (dim-only and shadow-only) failed outright: users could not tell hover-affected cards were interactive at all. The third, which does both, did read as interactive on desktop, so it was the one genuinely viable standardization candidate, and it was rejected on a separate, named ground: it depends entirely on hover, so it degrades to nothing on touch, where a growing share of this list view is used. That is what makes option (b) unsafe for every existing pattern rather than only for the two that were easiest to disqualify. Had the third pattern held up on touch as well, option (b) would have been the cheaper and correct answer, and the entry says so, so a future reader can see the decision was contestable rather than foreordained. Evidence links: those usability sessions, one round per existing pattern, 18 participants combined across the three, with the touch condition run explicitly on the third, plus a support-ticket search turning up 40 tickets mentioning "can't tell what's clickable" tied to card lists specifically. Owner: the Product Design Lead for the team that raised the issue. Approvers: two peer designers plus the Design Council, since this change touches three teams' shared component. Status moves from Proposed, through the Council's 48 to 72 hour review window, to Accepted. Risk or impact tier: cross-team, so it goes through the Council gate rather than self-publishing. Related artifacts: the new component's Figma file ID and its tracking ticket in the issue tracker. Review-due date: six months out, or sooner if a usability round on the new pattern turns up a problem. Once Accepted, the Figma webhook links the shipped component version to DEC-214, so a designer opening that component later sees the decision, the evidence, and the fact that two prior patterns were explicitly rejected, instead of re-litigating the same debate from nothing. A later quarterly audit samples DEC-214 and confirms the shipped component still matches what was Accepted, closing the loop from problem to shipped reality that the traceability model above is meant to guarantee.
Tooling integrations
A Figma plugin or webhook (a small notification one system sends automatically to another the moment something happens, here: "this component just got a new version") links component versions to decision IDs automatically, so opening a component in Figma surfaces the rationale behind it. Jira (or a comparable issue tracker) templates require a decision-log link before certain ticket types can close. Notion serves as the human-readable front end, backed by structured, queryable fields rather than free-form prose pages, so search and filtering actually work.
Adoption strategy and health metrics
Track two metrics from day one, and publish them where the whole design org can see them, not just internally to whoever runs the system: turnaround time (median hours from an entry being submitted to a status decision, a proxy for whether governance is slowing people down) and adoption rate (the share of qualifying changes, such as cross-team component changes, that actually have a linked entry versus an estimate of the total that should). A governance system that can't show its own users it isn't a bottleneck will get worked around regardless of how good the data model is.
Trade-offs and pitfalls
- Over-specifying metadata kills adoption: starting with five mandatory fields and ten optional ones, then adding more only once usage proves the need, beats launching with twenty required fields nobody wants to fill in.
- Turnaround time can be gamed by rubber-stamping reviews to look fast; it needs to be paired with a periodic quality spot-check, not tracked as the only signal of health.
- A rejected option is only actually rejected if the evidence covers it. The most common quiet defect in a decision entry is a "chosen option and rationale" that disqualifies an alternative using evidence gathered on a subset of the candidates it contains; reviewers rarely catch it because the entry reads as rigorous. Requiring the entry to state, per rejected option, which evidence covers WHICH candidate is a cheap structural fix.
- The traceability chain is only as strong as its weakest link, and that's usually the final one: linking a merged, shipped change back to its decision ID, because engineers rarely see a personal incentive to maintain that link. Assigning explicit ownership for that last connection (not leaving it to whoever remembers) is what keeps the chain from silently rotting.
A recent feature release generated feedback that it increased cognitive load for power users. As the PM, outline a situational design critique process: who you'd involve, what qualitative and quantitative data you'd collect, how you'd synthesize the findings, and a four-point action plan to reduce load.
Sample Answer
Direct answer
Treat a complaint-triggered spike as a signal to run on, not proof of the problem's shape. Pull together a small cross-functional group fast, gather both what users say and what they actually do, turn the notes into a short list of named friction points, and commit to a scoped, reversible fix with an owner and a two-week check-in, rather than waiting for a full controlled experiment to "prove" the load is real before acting.
Who to involve and what to collect
Keep the working group small (five or six people), not a big review:
- The designer who owns the feature, to walk through the actual screens and decisions made.
- A UX researcher, or the product manager acting as one if no researcher is staffed, to run and interpret sessions.
- Two or three power users, or a support/customer-success teammate who talks to them directly, since they generated the complaint.
- One engineer who knows the implementation, so the action plan doesn't propose something infeasible.
Data comes from two lanes that check each other:
Qualitative: re-read the verbatim complaints for repeated phrases, then run four or five short (20 to 30 minute) sessions where power users do their real workflow in the feature, not a scripted task. Watch for hesitation, backtracking, and "wait, where is..." moments.
Quantitative: whatever is already instrumented, before-vs-after the release: task completion time, backtrack or undo rate, feature adoption or opt-out rate, and the trend in support tickets mentioning the feature. The point is not to run a new controlled experiment (that's slower than this situation calls for); it's to see whether the qualitative pattern is a widespread shift or a handful of loud outliers.
Synthesis: cluster the qualitative notes into two or three named issues (not a laundry list of quotes), then check each named issue against the quantitative trend. An issue that shows up in sessions and moves a metric is real; an issue only one person mentioned, with no metric movement, gets watched, not acted on yet.
Worked example
Suppose the feature is a redesigned settings panel that now surfaces 12 toggles at once, where the old version showed 4 and tucked the rest behind an "Advanced" link. In four of five sessions, power users pause and visibly scan the panel for five or more seconds before their first action, something nobody did in the old version. The backtrack rate on this screen (any action followed by an undo or a return to the previous state) has jumped roughly eight-fold since launch, from about 3% of sessions that touch the panel before the release to around 25% after. Support tickets mentioning "settings" are up for the first time in months. Three independent signals (session observation, backtrack rate, tickets) point the same direction, so the team acts. Worth naming the size of the move as well as its direction: an eight-fold jump on a rate that was previously stable is not a metric drifting, it is a different screen, which is what justifies acting on a directional read instead of waiting for a controlled experiment.
Four-point plan:
- Ship a density toggle that defaults new and existing power users back to something close to the old 4-item view, with the rest reachable behind one click.
- Group the remaining toggles into two or three labeled sections instead of one flat list, so scanning has structure.
- Add short inline hints on the two or three toggles sessions showed people hesitating over most, not all twelve.
- Ship this as a fast-follow patch within the week, keep watching backtrack rate and ticket volume for two weeks, and set a checkpoint against the actual pre-release number rather than a vague "improved": if backtrack rate returns to roughly its 3% baseline, call it fixed; if it lands partway, say 10%, that is a partial fix and the remaining gap is the next piece of work, not a success; if it does not move at all, the density of the panel was not the real problem and it escalates to a deeper redesign.
Trade-offs and pitfalls
The biggest wrong turn is reaching for a full controlled experiment to "prove" cognitive load went up before doing anything. That's the right instrument for a slow-moving optimization question, not for a complaint that is actively costing users time this week; a directional read from converging qualitative and quantitative signals is enough to justify a reversible patch. The second pitfall is only listening to the loudest complainers: check that the metric movement is broad, not just that a few vocal users are unhappy, since power users are often a small, non-representative slice of the base. A third pitfall is reporting a change as a multiplier without its baseline: "backtrack rate tripled" and "backtrack rate is now a quarter of sessions" are compatible only if the baseline was around 8 percent, and the two framings imply very different releases. Always carry the before and after rates together, because the multiplier alone is the number that quietly drifts as a finding gets retold up the chain. Finally, a density toggle is a fast patch, not a fix: it adds a decision (which mode am I in) rather than removing one. Treat it as buying time for a properly-scoped section redesign, not as the end state.
You run a weekly design critique. Describe the structure of an effective critique session (roles, timeboxes, artifact type, expected outcomes) and how you create a psychologically safe environment so feedback is constructive.
Sample Answer
Direct answer
A critique session works when it has a repeatable structure (clear roles, a timeboxed agenda, an artifact everyone has seen in advance, and a required output) and explicit ground rules that separate feedback on the work from judgment of the person. Structure removes ambiguity about what is supposed to happen; the ground rules are what make people willing to say what they actually think.
Structured elaboration
Roles
- Presenter: owns the artifact, states the problem and the specific question they want answered (not "what do you think" but "does this navigation pattern confuse first-time users").
- Facilitator: keeps time, enforces the ground rules, and makes sure the session ends with decisions, not just opinions.
- Recorder: captures feedback as concrete items (issue, owner, priority) rather than a transcript of the conversation.
- Participants: at least one engineer for feasibility and one researcher or data owner when the discussion depends on evidence, plus two to four other reviewers.
Timeboxed agenda (40-60 minutes)
| Segment | Time | Purpose |
|---|---|---|
| Context | 5 min | Presenter states the problem, audience, and the specific feedback they need |
| Walkthrough | 10-15 min | Demo the artifact; no interruptions |
| Clarifying questions only | 5-10 min | Understand before judging; no solutioning yet |
| Structured feedback | 15-20 min | Observations first, suggestions second, one topic at a time |
| Decisions and owners | 5-10 min | What changes, who owns it, by when |
The range comes straight out of the table: every segment at its low end is a 40-minute tight version, every segment at its high end is a 60-minute full one. That matters when you are handed a fixed slot. In a 45-minute room I spend the five minutes above the floor on structured feedback rather than the walkthrough, because the walkthrough is the part that can be compressed by sharing the artifact in advance, and the feedback block is the only segment the session actually exists for. In a 30-minute room the honest move is to cut the number of questions asked, not to shave every segment proportionally, since a five-minute feedback block produces reactions rather than critique.
Artifact type scales with the question: low-fidelity sketches or flows when validating direction, an interactive prototype or annotated screens when the question is about interaction detail, a short research summary when the question is whether the evidence supports the direction at all.
Expected outcome: not "the room liked it." Every session closes with one of keep, change, or run an experiment, a prioritized list of action items with a named owner and a date, and, when the decision is uncertain, the metric that will tell you whether it worked. I log these in a shared tracker so follow-through is visible instead of implied.
Psychological safety, concretely: it is not "be nice." It means people feel safe enough to point out a real problem and safe enough to admit their own work has weaknesses, which produces both better decisions and more empathy, since the team is looking at the work through the user's eyes instead of defending their own choices. I build it with three mechanics: ground rules stated at the top of every session, most importantly "critique the work, not the person" ("I notice new users hesitate here" instead of "this is wrong") and "park anything outside today's question" so scope creep does not turn into a personal argument about unrelated choices; the presenter opening by naming their own known weak spots, which signals that flagging problems is expected rather than an attack; and the facilitator actively inviting quieter voices so the loudest opinion in the room does not become the default decision.
Worked example
For a weekly critique on a subscription checkout flow, the presenter opens with: "I'm testing whether removing the second confirmation step increases completion without users feeling rushed; give me feedback on the pacing, not the visual style." After the walkthrough, one reviewer observes that a security-conscious user might not notice their card was already saved. That becomes a logged action item ("add a visible saved-card confirmation, owner: presenter, due: next session"), not a live redesign debate.
Trade-offs and pitfalls
A short session with no artifact shared in advance turns into unfocused chatter with no output. Skipping the timebox lets the most senior or loudest person's opinion dominate, which is the opposite of the goal. Rotating the facilitator role spreads ownership and avoids one person controlling every decision, but it costs ramp-up time for whoever is new to the role. The most common failure is capturing feedback with no accountable owner and date: the room feels productive, nothing actually changes, and the next session loses credibility.
You have qualitative interview feedback where several participants explicitly prefer Feature A, while aggregate analytics show Feature B yields higher engagement. Outline a systematic approach to reconcile these conflicting signals: what additional data you would collect, segmentation or contextual analysis you'd run, experiments you'd design, and how you'd make a defensible decision.
Sample Answer
Direct answer
Conflicting signals like this usually mean the two measures are answering different questions, stated preference versus actual behavior, so the fix is not to pick a side but to find out why they disagree: collect more context, segment the data to see who is driving each signal, and design a test that isolates the actual cause before deciding.
Structured elaboration
Additional data to collect
- Ask the interview participants why they prefer Feature A: is it about ease of use, trust, or a specific task it does better?
- Look at what happens after the click on Feature B: does higher engagement mean people are succeeding at a task, or getting stuck and clicking around?
- Check satisfaction or a short post-use survey tied to each feature, not just usage counts.
Segmentation and context
- Break the analytics down by user segment, new versus returning, task type, device: a common pattern is that Feature B's engagement is concentrated in a segment whose behavior is not representative of the interview participants.
- Map the interview participants' profiles against the segments to see whether the qualitative preference reflects a narrow, non-representative slice of users.
Experiments
- Run a short controlled comparison that separates the specific thing each group values, for example a version combining Feature A's simpler interaction with Feature B's more visible entry point, and measure both engagement and downstream task success, not just clicks.
- If resources allow, test each feature against its own most relevant segment rather than the whole population at once.
Making a defensible decision
Prioritize evidence that people actually completed what they came to do, task success, retention, over a raw engagement count, since a click is not proof of value. Require the two evidence types to agree on the outcome that matters, not necessarily on every number, before committing; if they still disagree after segmentation, ship the safer, reversible option first, a phased rollout with monitoring, rather than betting fully on either signal.
Worked example
If interviews suggest Feature A feels simpler while analytics show Feature B drives more clicks, test a version that keeps A's simpler interaction but adds a clearer call to action similar to B's, and measure task completion plus a short satisfaction check for the core user segment before deciding on a full rollout, rather than trusting either signal alone. Concretely: a team redesigning a project-management app's task list can't agree between a compact list view (Feature A) and a card view with a visible "Add subtask" button (Feature B). Eight interviews with existing power users say the compact list feels faster and less cluttered, and three of the eight specifically call the card view "busy." Production analytics from the last 30 days show the card view getting 18% more clicks per session than the list view, most of them landing on the visible add-subtask button. The team ships a hybrid to a randomly chosen 10% test group, with the remaining 90% left on the current card view as the control (the control is the group held out from the change, which is why "holdout" names the 90%, not the 10%; getting that label the wrong way round in a readout is a fast way to have your result questioned): the compact list's row height and information density, but with a small, clearly labeled "+" button in the same spot the card view's button occupies. Over the next two weeks, task completion (a user actually adds a subtask, not just clicks toward one) rises from 34% to 41% in the 10% test group, measured against the 90% control still on the existing card view over the same two weeks, and a short one-question satisfaction check ("was this easy to use, yes or no") comes back positive from 78% of the test group, up from 61% in the control still on the old card view. That combination, higher completion and higher satisfaction, is what justifies the full rollout, not the raw click count that started the disagreement.
Trade-offs and pitfalls
The main trap is treating whichever metric is easier to report, usually the quantitative one, as automatically more true; a click is not the same as value delivered. The opposite trap is dismissing analytics because a handful of interviewees said otherwise, when the interview sample may not represent the users actually driving the metric. Running too many segment cuts without a clear hypothesis first turns into fishing for a story that confirms whatever you already believed.
Unlock Full Question Bank
Get access to all 41 Design Critique, Iteration, and Decision Rationale interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.