Structured Problem Solving and Decomposition Questions
Methodical problem solving for open-ended and ambiguous situations once the problem is defined: decomposing a goal or problem into mutually exclusive, collectively exhaustive parts (issue trees, metric trees and driver breakdowns, work breakdown into subproblems and vertical slices), forming and prioritizing hypotheses (hypothesis trees and funnels, including laying out candidate explanations for a metric drop or model degradation and the cheapest test for each branch), choosing an analytical approach, and reasoning to a recommendation. Also covers turning a vague mandate into measurable, testable subproblems with owners, breaking a large initiative into workstreams, mapping dependencies, sequencing the work, deciding the first deliverable and what to defer, structuring plans that mix research, analytics and experiments, and talking through real examples of cutting a messy problem into parts. Covers explaining and adapting structured problem-solving methods across contexts, choosing and switching methods, and coaching others to structure ambiguity. Excludes turning a vague request into a scoped problem statement, named-framework business cases, root-cause techniques for failures and metric movements (running the diagnosis itself), prioritization scoring and trade-off decisions, deciding how to act under incomplete information, market sizing and estimation, and framing machine-learning problems, which are covered elsewhere.
Retention dropped for one cohort and you are the designer on the team. Lay out how you would generate hypotheses, which would be answered by data queries versus research, and which you would test first.
Sample Answer
Direct answer
A cohort is a group of users who started in the same period or through the same route. Week-4 retention is the share of a cohort still active in the fourth week after signup, and a segment is a subset of users who share an attribute, such as acquisition channel. I would generate hypotheses from three angles (who the cohort is, what they experienced, when it happened), send "what and where" questions to data queries and "why" questions to research, and test first the cheap query that localises the drop, because it tells research whom to talk to. As the designer, I also bring something others lack: the change log of what design shipped to that cohort.
Structured elaboration
Generate hypotheses, one prompt per angle:
- Who: acquisition channel, plan, device, region, job role. Is this cohort different from neighbouring cohorts?
- What they experienced: which onboarding variant, which features existed, any outage, pricing or content change in their first weeks. This is where design history matters.
- When: where on the retention curve (the share of a cohort still active at each day or week since signup) does retention separate (day 1, week 2?), and any seasonality or external event.
Sort by what can answer them:
| Question type | Best source | Example |
|---|---|---|
| How big, where, when | Data query | retention by channel; the day the curves split |
| Who is different | Data query | segment mix of this cohort vs usual |
| Why did they leave | Research | interviews with churned users (people who stopped using the product) from the localised segment |
| Did the experience confuse them | Research and recordings | session recordings, usability test of the first-week flow |
Test first: the localising query (a query that narrows down where the drop is concentrated: retention by segment, and by day). It costs hours and prunes hypotheses. Then interview 6 to 8 churned users from the affected segment: these are exploratory (to find reasons), not a count of how common each reason is. Only then test a design fix.
Worked example
Illustrative. The cohort that signed up in one week shows week-4 retention of 28% against 35% for neighbouring cohorts, a 7-point gap. Take 10,000 signups per cohort. In a usual week, paid-social (users acquired through paid ads on social networks) is 20% of signups (2,000) and retains 15% (300 users), while the other 8,000 retain 40% (3,200), so 3,500 of 10,000 is 35%. In the odd week the segment query shows paid-social jumped to 48% of signups (4,800) and still retained 15% (720), while the other 5,200 still retained 40% (2,080), so 2,800 of 10,000 is 28%. No channel got worse; the mix shifted toward the low-retention channel, which accounts for the whole 7 points. So the design-history question becomes: did that audience land on a different onboarding flow? The change log shows a campaign-specific landing page promising a feature that onboarding never mentions. Interviews with churned users in that segment confirm the expectation gap. The fix is aligning the promise and the first-run experience, and the test is whether the next such cohort retains normally.
Trade-offs and pitfalls
- Data tells you where; it rarely tells you why. Do not stop at the query.
- Do not start with interviews. You will interview the wrong users and hear plausible but unrelated complaints.
- Be wary of a story that fits too well. Check it against another cohort that had the same exposure.
- Designers can overweight design causes. Rule out acquisition and external causes first.
Describe a time you reframed an ambiguous product problem into something you could test. How did you get there, and what did the test change?
Sample Answer
Direct answer
I would tell one specific story: a vague brief ("make the dashboard easier to use") that I turned into a falsifiable claim (one that some possible observation could prove wrong), tested cheaply, and let the result change what we built. The strong shape is: what made the problem ambiguous, how I narrowed it, what exactly the test could prove or disprove, and which decision moved because of it. (The details below are an illustrative story skeleton. Replace them with your own real project.)
Structured elaboration
- Situation. As the product designer on an analytics tool, I was handed "new admins find the dashboard confusing." That is a feeling, not something I could test.
- How I got to a testable problem. I (1) listed what "confusing" could mean: cannot find the first report, do not understand the chart labels, or do not trust the numbers; (2) read 20 support tickets and watched session recordings (screen captures of real users) to see which of the three showed up in behaviour; (3) wrote one claim: "New admins who cannot reach their first report in the first session are the ones who leave, and the cause is the navigation label, not the chart design."
- The test. A prototype test with two navigation variants: the existing label ("Insights") and a task-based label ("See your first report", named for what the user wants to do). Eight participants saw each variant. Prediction written beforehand: if the label is the cause, participants on the new label reach the report unaided noticeably more often; if both variants fail the same way, my claim is wrong. Task success was measured by who reached the report without help.
- What the test changed. Most participants on both variants still stalled, at the step after navigation (choosing a data source): 2 of 8 reached the report unaided with the old label and 3 of 8 with the new one, far short of the 6 of 8 bar in the worked example below. The label was not the main cause. We dropped the planned navigation rework and put the effort into a guided first-run flow for data source selection.
- Result, kept honest. A moderated test (a researcher present, giving tasks and observing) with eight participants per variant shows direction, not proof. I reported it as "strong enough to redirect a two-week design effort, not strong enough to forecast a metric", and the later activation measurement (the share of new admins reaching their first report in live use) was the real confirmation.
Worked example
Before: "dashboard is confusing" (cannot be wrong, so cannot be tested). After: "In a first session, at least 6 of 8 new admins will reach their first report unaided if the label is the cause." The after-version names who, what behaviour, and what result would count as failing. When the test showed most stalling in both variants, the claim failed cleanly, which is exactly why it was useful.
Trade-offs and pitfalls
- Common wrong turn: reframing the problem into a question that confirms the solution you already like. Writing the prediction and the failing result first prevents that.
- Over-claiming: do not quote a percentage lift from a small prototype test.
- What I would do differently: check the funnel data (counts of users reaching each step of the flow) before building the prototype. A single query would have shown the drop was after navigation and saved a variant.
Product wants a 'save for later' feature on web and mobile in two weeks. How would you break it down, what would you cut to make a first slice that is real and usable, and what would you tell product about what is deferred?
Sample Answer
Direct answer
Define "real and usable" as the smallest loop a user can complete: save an item on one device, come back later, find it, open it. Build exactly that loop on web and mobile against one shared server-side store, cut everything that organizes, shares or notifies, and give product a written list of what is deferred, why, and what would bring each item back.
Structured elaboration
Breakdown. Components: save and unsave control on items; saved-items storage and API (add, remove, list); saved list screen; sign-in requirement; analytics events; edge cases (item deleted after saving, duplicate saves, long lists).
Cut rule. Keep what the loop needs. Cut what adds organization (folders, tags), distribution (sharing), prompting (reminders, notifications) or fidelity (offline save, real-time cross-device push).
Kept for the first slice
- Save/unsave toggle on item detail and list rows.
- A saved list, newest first, with pagination (the list loads a page at a time, for example 20 items, instead of all at once).
- Logged-in users only: an anonymous tap leads to sign-in, then completes the save.
- Server-side storage, so the same account sees the same list on web and mobile (pull on open, not push: the app fetches the list when the screen opens instead of the server sending live updates).
- Save and unsave are idempotent (repeating the request has no extra effect) so double taps are harmless.
Assumed team and plan (state assumptions out loud). One engineer each for backend, web and mobile, 10 working days:
- Days 1-2: agree the API contract (the exact requests and responses) and the data model (how a saved item is stored: user, item, time saved); backend ships a stub (a stand-in that returns fixed sample data so web and mobile can build before the real service exists).
- Days 3-7: backend, web and mobile build in parallel against the contract.
- Days 8-9: integration, edge cases, analytics events.
- Day 10: buffer and a staged rollout (release to a small share of users first) behind a feature flag (a switch that turns the feature on or off without a new release).
Worked example
What you tell product, as a short table:
| Item | Status | Why | What brings it back |
|---|---|---|---|
| Save, unsave, saved list on web and mobile | Ships in 2 weeks | The core loop | n/a |
| Folders or collections | Deferred | Adds UI, a bigger data model (folder records) and migration work (updating stored data to the new structure) | Saved lists growing long, or users asking to group |
| Sharing a saved list | Deferred | New permissions and privacy questions | Evidence of group use cases |
| Reminders and notifications | Deferred | Needs notification infrastructure and consent design | Saved items rarely revisited |
| Offline saving on mobile | Deferred | Needs sync and conflict handling (reconciling saves made offline with the server when the two versions differ) | Many saves made on poor connections |
| Real-time cross-device updates | Deferred | List refresh on open covers most needs | Visible stale-list complaints |
Success measures to agree before launch: saves per active user, and the share of savers who return to the list within 7 days. Any estimate of how long a deferred item takes is a rough estimate to confirm after the first slice, not a promise.
A worked reading of the measures (illustrative): in launch week 2,000 users are active and 500 of them save at least once, making 3,000 saves in total, so saves per active user is 3,000 / 2,000 = 1.5. If 150 of those 500 savers open the list again within 7 days, the return rate is 150 / 500 = 30%.
Trade-offs and pitfalls
- Cutting "web or mobile" instead of cutting depth breaks the brief; cut features, keep both platforms.
- Silent scope cuts destroy trust. A deferred list with reasons makes the cut a decision product owns.
- Skipping edge cases to save time ships bugs on day one: deleted items and duplicate saves must be handled even in the thin slice.
- If the two weeks are fixed and the team is smaller than assumed, cut the web or mobile list screen polish before cutting the backend contract.
- What would flip the plan: if anonymous saving is the main use case, you would need local storage and merge-on-login, which is a bigger first slice.
Explain what makes a breakdown 'MECE' to a new analyst, using a retention decline as the example. Show one split that looks reasonable but overlaps or leaves a gap, then fix it, and say what you would do to pick the most useful way to cut the problem.
Sample Answer
Direct answer
MECE (pronounced "me-see") stands for mutually exclusive, collectively exhaustive. Mutually exclusive means no item belongs to two branches. Collectively exhaustive means every item belongs to at least one branch. Together: each item has exactly one home. A split is only useful if it is MECE and its branches lead to different actions.
Example: monthly retention fell (illustrative numbers)
Retention is the share of customers at the start of the month who are still customers at the end. Say 1,000 customers started the month; retention fell from 80% (800 kept, 200 lost) to 72% (720 kept, 280 lost), so 80 extra customers left.
A split that looks reasonable but is not MECE
Cut the 1,000 customers into: new users (300), mobile users (500), enterprise customers (150), European customers (200).
- Overlap: Priya joined 10 days ago, uses mobile, is on an enterprise plan and lives in Germany. She sits in all four branches. The counts add to 300 + 500 + 150 + 200 = 1,150, more than the 1,000 customers, which is the quick tell.
- Gap: a two-year-old desktop customer on a small-business plan in the US fits none of the four branches.
The fix: one dimension, one rule
Split by tenure at the start of the month: under 30 days (300), 30 to 179 days (400), 180 days or more (300). Each customer has exactly one start date, so exactly one bucket (exclusive), and 300 + 400 + 300 = 1,000 (exhaustive). Other dimensions such as mobile versus desktop then become a second-level split inside each tenure bucket, never mixed in at the same level.
Tests for any split
- Sum test: do the branch counts add to the total?
- One-home test: pick five real items and place each. Two homes means overlap, none means a gap.
- So-what test: does each branch imply a different action? If not, the split is correct but useless.
How I pick the most useful cut
Four families of splitting dimension:
| Family | Example | Why it is MECE |
|---|---|---|
| Formula | Retention = 1 - churn; churn = voluntary cancel + involuntary lapse (failed payment) | An identity, so each churned customer has one recorded final status |
| Process or funnel | Fulfillment throughput: received, picked, packed, shipped | Stages are sequential, so an unshipped order is in exactly one stage. The stage with the longest queue is the bottleneck candidate |
| Segment | Tenure, plan, region | Each customer has one value per attribute |
| Cause | Product, price, competitor | Hardest to keep MECE, since causes interact. Use last, and attribute each loss to a primary cause |
I choose the dimension where (a) branches differ most in outcome, (b) data exists, (c) branches map to owners, and (d) I can check it quickly. For a retention decline I would start with tenure, because "new-user problem or long-time-user problem" decides which team owns it.
Continuing the example (illustrative, same tenure mix last month): last month 60, 80 and 60 customers were lost in the three tenure buckets (200 in total). This month the losses are 140, 80 and 60 (280 in total). The extra 80 sit entirely in the under-30-days bucket, whose loss rate went from 60/300 = 20% to 140/300 = about 47%, while the other two stayed at 20%. That is the so-what: the problem is onboarding, not long-time customers, so the next split goes inside that bucket.
More partitions
- Signup conversion: never started, started and abandoned, submitted but unverified, completed.
- Onboarding non-completers, with a hypothesis per bucket: (1) never started: the prompt is not being seen; (2) started and stalled: one step is too hard or long; (3) completed but not recorded: a tracking gap undercounts completion; (4) ineligible, such as test or internal accounts: the denominator is polluted. Tie rule: eligibility is checked first, so a stalled internal account lands in bucket 4 only.
Pitfalls
- Mixing dimensions at one level (the 1,150 trap).
- An "Other" bucket that quietly holds half the data.
- Treating MECE as a pass mark: a MECE split nobody can act on is still a bad split.
Half of your interview participants say they want more customization and half say they want simplicity. How do you frame that tension into testable questions, and how do you decide what to learn next?
Sample Answer
Direct answer
First, check whether it is truly a contradiction. A 50/50 split among a small set of interviews is a prompt for questions, not a market finding. Then turn the tension into four testable questions about who wants what, what they actually do, what each word means, and whether one design can serve both. For next steps, choose the cheapest evidence that would change the design decision: re-code what you already have, then usage data, then a prototype test.
Structured elaboration
Four testable questions:
- Does the split line up with a user attribute? (for example daily users vs occasional users, expert vs novice, role). Re-code the existing interviews by that attribute: go back through your notes and tag each participant with it (for example daily or occasional user), then see whether the two camps differ. This costs nothing.
- What is behind "customization"? A specific unmet need, a wish for control, or a workaround already in use? Check what people do today, not what they say.
- What does "simplicity" mean to each person? Fewer options, fewer steps, or less to learn are different design problems.
- Can defaults plus progressive disclosure serve both? Progressive disclosure means showing a simple path by default and revealing advanced options on request. Test with a prototype that has both paths. In practice, a settings-heavy screen opens with sensible defaults (the pre-chosen settings most people never change) and one "Advanced" panel. Give each participant a task such as "set up a weekly report for your team" and watch whether occasional users finish unaided on the default path and whether frequent users find and change the advanced controls without help.
Deciding what to learn next: rank by how much the answer changes the design decision and how cheap it is. (1) Re-code existing data; (2) pull product usage: what share of users ever change a setting today; (3) prototype test of question 4. A survey comes later, and only to size a segment difference (estimate how many users fall into each group and how their needs differ) once found.
Worked example
Illustrative. Twelve interviews: 6 ask for customization, 6 for simplicity. Re-coded by usage frequency, the customization group is 5 daily users and 1 occasional user; the simplicity group is 2 daily and 4 occasional. Daily users total 7 and occasional users 5, so 12 in all. The split now looks like a frequency split, which gives a sharper hypothesis: "frequent users want control, occasional users want a clean path." Twelve interviews cannot confirm it, so I would check usage data for how many people change the default settings, then run a prototype with a simple default and an advanced panel and watch whether occasional users finish tasks unaided and frequent users find the controls.
Trade-offs and pitfalls
- Counting opinions is weak evidence. A 6 to 6 split in 12 people says nothing about the proportions in the user base.
- Do not average the two groups into a middle design that satisfies neither.
- Stated wants and behaviour diverge. What people ask for in interviews often differs from what they use.
- What would change my call: if usage data shows almost nobody changes any setting, I would lean toward simplicity and park customization.
Unlock Full Question Bank
Get access to all 13 Structured Problem Solving and Decomposition interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.