Structured Problem Solving and Decomposition Questions
Methodical problem solving for open-ended and ambiguous situations once the problem is defined: decomposing a goal or problem into mutually exclusive, collectively exhaustive parts (issue trees, metric trees and driver breakdowns, work breakdown into subproblems and vertical slices), forming and prioritizing hypotheses (hypothesis trees and funnels, including laying out candidate explanations for a metric drop or model degradation and the cheapest test for each branch), choosing an analytical approach, and reasoning to a recommendation. Also covers turning a vague mandate into measurable, testable subproblems with owners, breaking a large initiative into workstreams, mapping dependencies, sequencing the work, deciding the first deliverable and what to defer, structuring plans that mix research, analytics and experiments, and talking through real examples of cutting a messy problem into parts. Covers explaining and adapting structured problem-solving methods across contexts, choosing and switching methods, and coaching others to structure ambiguity. Excludes turning a vague request into a scoped problem statement, named-framework business cases, root-cause techniques for failures and metric movements (running the diagnosis itself), prioritization scoring and trade-off decisions, deciding how to act under incomplete information, market sizing and estimation, and framing machine-learning problems, which are covered elsewhere.
Describe a time you reframed an ambiguous product problem into something you could test. How did you get there, and what did the test change?
Sample Answer
Direct answer
I would tell one specific story: a vague brief ("make the dashboard easier to use") that I turned into a falsifiable claim (one that some possible observation could prove wrong), tested cheaply, and let the result change what we built. The strong shape is: what made the problem ambiguous, how I narrowed it, what exactly the test could prove or disprove, and which decision moved because of it. (The details below are an illustrative story skeleton. Replace them with your own real project.)
Structured elaboration
- Situation. As the product designer on an analytics tool, I was handed "new admins find the dashboard confusing." That is a feeling, not something I could test.
- How I got to a testable problem. I (1) listed what "confusing" could mean: cannot find the first report, do not understand the chart labels, or do not trust the numbers; (2) read 20 support tickets and watched session recordings (screen captures of real users) to see which of the three showed up in behaviour; (3) wrote one claim: "New admins who cannot reach their first report in the first session are the ones who leave, and the cause is the navigation label, not the chart design."
- The test. A prototype test with two navigation variants: the existing label ("Insights") and a task-based label ("See your first report", named for what the user wants to do). Eight participants saw each variant. Prediction written beforehand: if the label is the cause, participants on the new label reach the report unaided noticeably more often; if both variants fail the same way, my claim is wrong. Task success was measured by who reached the report without help.
- What the test changed. Most participants on both variants still stalled, at the step after navigation (choosing a data source): 2 of 8 reached the report unaided with the old label and 3 of 8 with the new one, far short of the 6 of 8 bar in the worked example below. The label was not the main cause. We dropped the planned navigation rework and put the effort into a guided first-run flow for data source selection.
- Result, kept honest. A moderated test (a researcher present, giving tasks and observing) with eight participants per variant shows direction, not proof. I reported it as "strong enough to redirect a two-week design effort, not strong enough to forecast a metric", and the later activation measurement (the share of new admins reaching their first report in live use) was the real confirmation.
Worked example
Before: "dashboard is confusing" (cannot be wrong, so cannot be tested). After: "In a first session, at least 6 of 8 new admins will reach their first report unaided if the label is the cause." The after-version names who, what behaviour, and what result would count as failing. When the test showed most stalling in both variants, the claim failed cleanly, which is exactly why it was useful.
Trade-offs and pitfalls
- Common wrong turn: reframing the problem into a question that confirms the solution you already like. Writing the prediction and the failing result first prevents that.
- Over-claiming: do not quote a percentage lift from a small prototype test.
- What I would do differently: check the funnel data (counts of users reaching each step of the flow) before building the prototype. A single query would have shown the drop was after navigation and saved a variant.
Half of your interview participants say they want more customization and half say they want simplicity. How do you frame that tension into testable questions, and how do you decide what to learn next?
Sample Answer
Direct answer
First, check whether it is truly a contradiction. A 50/50 split among a small set of interviews is a prompt for questions, not a market finding. Then turn the tension into four testable questions about who wants what, what they actually do, what each word means, and whether one design can serve both. For next steps, choose the cheapest evidence that would change the design decision: re-code what you already have, then usage data, then a prototype test.
Structured elaboration
Four testable questions:
- Does the split line up with a user attribute? (for example daily users vs occasional users, expert vs novice, role). Re-code the existing interviews by that attribute: go back through your notes and tag each participant with it (for example daily or occasional user), then see whether the two camps differ. This costs nothing.
- What is behind "customization"? A specific unmet need, a wish for control, or a workaround already in use? Check what people do today, not what they say.
- What does "simplicity" mean to each person? Fewer options, fewer steps, or less to learn are different design problems.
- Can defaults plus progressive disclosure serve both? Progressive disclosure means showing a simple path by default and revealing advanced options on request. Test with a prototype that has both paths. In practice, a settings-heavy screen opens with sensible defaults (the pre-chosen settings most people never change) and one "Advanced" panel. Give each participant a task such as "set up a weekly report for your team" and watch whether occasional users finish unaided on the default path and whether frequent users find and change the advanced controls without help.
Deciding what to learn next: rank by how much the answer changes the design decision and how cheap it is. (1) Re-code existing data; (2) pull product usage: what share of users ever change a setting today; (3) prototype test of question 4. A survey comes later, and only to size a segment difference (estimate how many users fall into each group and how their needs differ) once found.
Worked example
Illustrative. Twelve interviews: 6 ask for customization, 6 for simplicity. Re-coded by usage frequency, the customization group is 5 daily users and 1 occasional user; the simplicity group is 2 daily and 4 occasional. Daily users total 7 and occasional users 5, so 12 in all. The split now looks like a frequency split, which gives a sharper hypothesis: "frequent users want control, occasional users want a clean path." Twelve interviews cannot confirm it, so I would check usage data for how many people change the default settings, then run a prototype with a simple default and an advanced panel and watch whether occasional users finish tasks unaided and frequent users find the controls.
Trade-offs and pitfalls
- Counting opinions is weak evidence. A 6 to 6 split in 12 people says nothing about the proportions in the user base.
- Do not average the two groups into a middle design that satisfies neither.
- Stated wants and behaviour diverge. What people ask for in interviews often differs from what they use.
- What would change my call: if usage data shows almost nobody changes any setting, I would lean toward simplicity and park customization.
Your VP gives you the mandate 'improve product adoption' and nothing else. How would you lead the team to turn that into measurable subproblems with owners, and how would you know the breakdown is good enough to start work?
Sample Answer
Direct answer
Run a working session, not a solo analysis. The team agrees one definition of adoption, breaks it into stages whose rates multiply back to the whole, gives every stage one metric with a baseline and one owner, and then stress-tests the breakdown before anyone builds. The VP confirms the definition. The team draws the branches.
Step 1: define adoption briefly
Adoption = the share of eligible accounts that reach the key value action (the first action that shows a user got real value, for example "shared a first report") and still use it in week 4. Eligible means accounts that could use the product. Here you need only enough definition to measure the goal.
Step 2: break it down with real numbers (illustrative)
Of 10,000 eligible new accounts: 4,000 try the product (40%), 2,000 of those reach the key action (50%), and 1,200 of those are still active in week 4 (60%). Check: 10,000 × 0.4 × 0.5 × 0.6 = 1,200, so adoption is 12%.
Because the stages multiply, +10 points at any one stage gives: try 40 to 50% yields 1,500 adopters, activate 50 to 60% yields 1,440, retain 60 to 70% yields 1,400. Arithmetic slightly favours earlier stages, but not because an extra early account is worth more: an extra account at the try stage is worth only 0.5 x 0.6 = 0.3 of an adopter, while an extra retained account is worth a full adopter. Ten points is worth more at try because it is applied to a larger base: 10 points of 10,000 eligible accounts is 1,000 extra triers (x 0.3 = 300 adopters, the 1,500), 10 points of 4,000 triers is 400 extra activated accounts (x 0.6 = 240, the 1,440), and 10 points of 2,000 activated accounts is 200 extra retained accounts (x 1.0 = 200, the 1,400). The differences are small, though (1,500 versus 1,440 versus 1,400 adopters is about one percentage point of adoption between best and worst), so research on which stage is actually feasible to improve decides.
Three stage subproblems plus one enabling workstream:
| Subproblem | Metric (baseline: today's measured value) | Accountable owner | Partner | First hypothesis |
|---|---|---|---|---|
| Try | 40% of eligible accounts start within 14 days | Growth PM | Product marketing | Eligible users never see the entry point |
| Activate | 50% of triers reach the key action | Product Designer (onboarding flow) | Design Researcher (first-use sessions) | The first screen hides the key action |
| Retain | 60% of activated still active in week 4 | PM for core experience | Technical PM, engineering | The key action does not become a habit |
| Enable: measurement | Weekly stage rates exist | Data Analyst | Technical PM | Events are not tracked today |
Run one customer segment end to end first (a vertical slice: one thin end-to-end piece of work, rather than one stage for every segment) rather than every stage for every segment, because averages can hide a segment where adoption is near zero.
Step 3: is the breakdown good enough to start?
Five tests:
- Rebuild: the parts reproduce the whole (0.4 × 0.5 × 0.6 returns 12%). If a stage cannot be written as a rate of the previous one, the tree overlaps or leaves gaps.
- Ownership: each leaf has one metric, one baseline and one owner.
- Monday test: the owner could start tomorrow without asking you what the leaf means.
- Testable: the first test finishes within one to two sprints (a sprint is a fixed one- or two-week work cycle). If not, split again or pick a leading indicator (an early signal that predicts the real outcome, such as week-1 repeat use standing in for week-4 retention).
- Missing branch: ask "if adoption rose but none of our metrics moved, what happened?" Each answer is a missing branch or a measurement gap. For example, if adoption rose from 12% to 14% while the try, activate and retain rates barely moved, the likely answers are a new partner channel that was never counted as eligible, or an event that changed what counts as the key action.
Stop splitting when a leaf passes tests 2 to 4. Going further creates busywork.
Leading the session
Have each function sketch its own branch for ten minutes, then merge and resolve overlaps. Send the VP a one-page summary: definition, funnel, owners, first tests, and what you are deliberately not doing.
Pitfalls
Decomposing alone, giving a leaf two owners or none, and splitting for weeks without starting a test.
Retention dropped for one cohort and you are the designer on the team. Lay out how you would generate hypotheses, which would be answered by data queries versus research, and which you would test first.
Sample Answer
Direct answer
A cohort is a group of users who started in the same period or through the same route. Week-4 retention is the share of a cohort still active in the fourth week after signup, and a segment is a subset of users who share an attribute, such as acquisition channel. I would generate hypotheses from three angles (who the cohort is, what they experienced, when it happened), send "what and where" questions to data queries and "why" questions to research, and test first the cheap query that localises the drop, because it tells research whom to talk to. As the designer, I also bring something others lack: the change log of what design shipped to that cohort.
Structured elaboration
Generate hypotheses, one prompt per angle:
- Who: acquisition channel, plan, device, region, job role. Is this cohort different from neighbouring cohorts?
- What they experienced: which onboarding variant, which features existed, any outage, pricing or content change in their first weeks. This is where design history matters.
- When: where on the retention curve (the share of a cohort still active at each day or week since signup) does retention separate (day 1, week 2?), and any seasonality or external event.
Sort by what can answer them:
| Question type | Best source | Example |
|---|---|---|
| How big, where, when | Data query | retention by channel; the day the curves split |
| Who is different | Data query | segment mix of this cohort vs usual |
| Why did they leave | Research | interviews with churned users (people who stopped using the product) from the localised segment |
| Did the experience confuse them | Research and recordings | session recordings, usability test of the first-week flow |
Test first: the localising query (a query that narrows down where the drop is concentrated: retention by segment, and by day). It costs hours and prunes hypotheses. Then interview 6 to 8 churned users from the affected segment: these are exploratory (to find reasons), not a count of how common each reason is. Only then test a design fix.
Worked example
Illustrative. The cohort that signed up in one week shows week-4 retention of 28% against 35% for neighbouring cohorts, a 7-point gap. Take 10,000 signups per cohort. In a usual week, paid-social (users acquired through paid ads on social networks) is 20% of signups (2,000) and retains 15% (300 users), while the other 8,000 retain 40% (3,200), so 3,500 of 10,000 is 35%. In the odd week the segment query shows paid-social jumped to 48% of signups (4,800) and still retained 15% (720), while the other 5,200 still retained 40% (2,080), so 2,800 of 10,000 is 28%. No channel got worse; the mix shifted toward the low-retention channel, which accounts for the whole 7 points. So the design-history question becomes: did that audience land on a different onboarding flow? The change log shows a campaign-specific landing page promising a feature that onboarding never mentions. Interviews with churned users in that segment confirm the expectation gap. The fix is aligning the promise and the first-run experience, and the test is whether the next such cohort retains normally.
Trade-offs and pitfalls
- Data tells you where; it rarely tells you why. Do not stop at the query.
- Do not start with interviews. You will interview the wrong users and hear plausible but unrelated complaints.
- Be wary of a story that fits too well. Check it against another cohort that had the same exposure.
- Designers can overweight design causes. Rule out acquisition and external causes first.
A redesigned onboarding prototype got glowing feedback in interviews but moved activation by nothing. How would you break down the possible reasons for the mismatch and decide which to investigate first?
Sample Answer
Direct answer
The mismatch has three families of explanation: the test did not measure what we think, the interview signal was misleading, or the redesign worked on something that is not the bottleneck. I would investigate in this order: (1) check the experiment and measurement (a data query, hours), (2) look at step-level funnel data for where the redesign did or did not move behaviour, (3) re-read the interview evidence for opinion versus behaviour. Activation means a new user reaching the first meaningful value action, such as completing a first project.
Structured elaboration
A. The test did not measure the change.
- Users never saw the new flow (an exposure bug: the assignment logic put people in the wrong group, or the rollout, the gradual release to a share of users, reached the wrong audience). The new arm, meaning the group of users assigned the redesigned flow, is the group to check first.
- Activation is defined so the redesign cannot move it, or is measured over too short a window.
- Too few users to detect a small effect (with a small sample, random noise can be as large as the real change, so a real improvement looks flat; a sample-size calculation says how many users are needed).
B. The interview signal was misleading.
- Participants are polite and say they like things (social desirability: people give answers that make them look agreeable).
- Prototype in a guided session is not the real flow, with real data and distractions.
- Recruited participants were more motivated than typical signups.
- Questions led toward praise.
C. The redesign improved something that is not the bottleneck (the step where the most potential activations are lost).
- Drop-off lives in a later step.
- Improvement increased completion but with lower-intent users, so later steps got worse (the effect is masked).
Why this order: A is a broken instrument (the measurement setup: tracking, assignment and the metric definition). If it is broken, B and C are unanswerable. It is also the cheapest. C is next because step-level data is already collected. B comes last because it needs new research, and I would run it only on the part that C narrowed.
Worked example
Illustrative numbers. Of 1,000 signups, 800 finished onboarding (80%) and 25% of finishers started a first project, so 200 activated (20%). After the redesign, 900 finish onboarding (90%), but only about 22% of finishers start a project, so about 198 activate, roughly the same 20%. Overall activation shows nothing, yet step-level data shows the redesign did what interviews praised (more people finish) while pulling in lower-intent users at the next step. The next test is a change to the step after onboarding, not further polish on the flow people already like.
A second case, where the redesign simply hit the wrong step (illustrative): of 1,000 signups, 92% finish onboarding before the redesign and 95% after, but only 20% of finishers start a project at the next step either way. Activation goes from 1,000 x 0.92 x 0.20 = 184 to 1,000 x 0.95 x 0.20 = 190, a gain of 6 users, too small to see in a flat topline. The bottleneck is the 20% step, so that is where the next test goes.
Trade-offs and pitfalls
- Do not call the interviews wrong. Interviews measure comprehension and attitude well and behaviour poorly.
- A flat topline (the single overall number on the experiment dashboard) can hide offsetting movements. Always look at steps.
- Do not rerun interviews first. It feels natural to a researcher, but it is the most expensive way to learn nothing new.
- What would change my order: if the experiment dashboard (the tool showing how many users are in each group and the metric for each) showed fewer users than planned in the new arm, I would stop and fix exposure before anything else.
Unlock Full Question Bank
Get access to all 7 Structured Problem Solving and Decomposition interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.