Structured Problem Solving and Decomposition Questions
Methodical problem solving for open-ended and ambiguous situations once the problem is defined: decomposing a goal or problem into mutually exclusive, collectively exhaustive parts (issue trees, metric trees and driver breakdowns, work breakdown into subproblems and vertical slices), forming and prioritizing hypotheses (hypothesis trees and funnels, including laying out candidate explanations for a metric drop or model degradation and the cheapest test for each branch), choosing an analytical approach, and reasoning to a recommendation. Also covers turning a vague mandate into measurable, testable subproblems with owners, breaking a large initiative into workstreams, mapping dependencies, sequencing the work, deciding the first deliverable and what to defer, structuring plans that mix research, analytics and experiments, and talking through real examples of cutting a messy problem into parts. Covers explaining and adapting structured problem-solving methods across contexts, choosing and switching methods, and coaching others to structure ambiguity. Excludes turning a vague request into a scoped problem statement, named-framework business cases, root-cause techniques for failures and metric movements (running the diagnosis itself), prioritization scoring and trade-off decisions, deciding how to act under incomplete information, market sizing and estimation, and framing machine-learning problems, which are covered elsewhere.
You are asked to add search to an existing mobile app. How would you break the initiative into subcomponents, map the dependencies between them, and decide which slice ships first?
Sample Answer
Direct answer
Break search into the layers a query passes through, agree the contract between the layers first so teams can work in parallel, then ship the thinnest slice that goes all the way from the search box to a tappable result. A thin vertical slice (one narrow path through every layer) beats finishing one layer completely, because only a shipped slice produces the real queries that tell you what to build next.
Structured elaboration
Subcomponents of "add search to a mobile app"
| Component | Job |
|---|---|
| Content source and index | Decide what is searchable and keep a searchable copy (an index) of it fresh |
| Search service and API | Takes a query string, returns ranked results |
| Query handling and ranking | Typo tolerance, synonyms, ordering of results |
| Mobile UI | Search box, results list, empty state, recent searches |
| Instrumentation | Log queries, taps, and searches with no results |
| Quality evaluation | A set of example queries with expected results |
| Operations | Response time, cost, access control on results |
Dependencies. The UI needs the API contract (request and response shape), not the finished API, so UI and backend can start together against a stub (a stand-in that returns fixed sample results so the screen can be built before the real service exists). The API needs the index; the index needs access to the content. Ranking improvements need query logs, which need instrumentation, which needs the UI to be live. That last chain is why the first slice should ship early.
Which slice first. Choose the slice that touches every layer and teaches the most: search one content type by title with prefix matching (what has been typed so far is matched against the start of titles, so "sho" finds "shoes"), return the first 20 results, tap to open the item, log every query. Defer filters, typo tolerance, autocomplete, voice input and personalization. Each is a bounded add-on once real queries exist.
Worked example
Slice 1 for a shopping app: "type a product name, see up to 20 matching products, tap one." Contract agreed on day one: request is a query string plus a page cursor (a marker saying where the next page of results begins); response is a list of item id, title and thumbnail. Done-criteria: the UI renders results from the stub; the real API answers within the latency target the team set; every query is logged. As one checkable sentence: typing "shoe" returns up to 20 matching products from the real API within the agreed response time, tapping one opens its detail page, and a log row with the query text and result count exists.
The same method on other initiatives, in short
- Save for later: the interfaces are client to saved-items API (save, unsave, list) and API to analytics (event schema). Agreeing those two contracts lets mobile, web and backend build in parallel.
- Four-market wallet launch: decompose by workstream (compliance, partnerships, UX, backend, fraud) and launch one market first as the thin slice. Compliance and partner approvals are external and slow, so start them first even though the engineering path looks shorter.
- Audit logging: components are what to record, capture hook, tamper-resistant storage, retention, and a viewer. Validate each by performing an action and checking one record with who, what and when appears, then trying to alter it and confirming that fails.
- ML feature pipeline (ingestion, features, training, inference, monitoring), each with one done-criterion and one risk to mitigate. Terms: a feature is an input signal given to the model; offline means on stored historical data and online means on live traffic; training teaches the model from past examples and inference is the model predicting on new cases; a label is the known right answer used to teach and score the model; a schema is the agreed list of columns and types in the data; proxy metrics are quick stand-in measures used while real labels have not arrived yet.
| Component | Done when | Risk and mitigation |
|---|---|---|
| Ingestion | A day of data lands with expected row counts | Silent schema change: validate schema on arrival |
| Features | Same feature values offline and online on a sample | Training and serving drift apart (the model sees differently computed values live than it learned from): one shared definition |
| Training | A reproducible run beats a simple baseline | Leakage of future data (the model learns from information that would not exist yet at prediction time): split by time |
| Inference | Predictions return inside the latency budget | Dependency down: fall back to a safe default |
| Monitoring | Dashboard and alert on input and output shifts | Labels arrive late: watch proxy metrics |
Trade-offs and pitfalls
- Splitting by team ("backend search, mobile search") hides dependencies; split by what each piece delivers.
- Finishing the index and ranking before any UI exists means no user learning for weeks.
- A slice must be usable, not a demo: includes an empty state and error handling.
- Do not turn each component into a long acceptance-criteria document; one checkable done-criterion is enough at this stage.
Your VP gives you the mandate 'improve product adoption' and nothing else. How would you lead the team to turn that into measurable subproblems with owners, and how would you know the breakdown is good enough to start work?
Sample Answer
Direct answer
Run a working session, not a solo analysis. The team agrees one definition of adoption, breaks it into stages whose rates multiply back to the whole, gives every stage one metric with a baseline and one owner, and then stress-tests the breakdown before anyone builds. The VP confirms the definition. The team draws the branches.
Step 1: define adoption briefly
Adoption = the share of eligible accounts that reach the key value action (the first action that shows a user got real value, for example "shared a first report") and still use it in week 4. Eligible means accounts that could use the product. Here you need only enough definition to measure the goal.
Step 2: break it down with real numbers (illustrative)
Of 10,000 eligible new accounts: 4,000 try the product (40%), 2,000 of those reach the key action (50%), and 1,200 of those are still active in week 4 (60%). Check: 10,000 × 0.4 × 0.5 × 0.6 = 1,200, so adoption is 12%.
Because the stages multiply, +10 points at any one stage gives: try 40 to 50% yields 1,500 adopters, activate 50 to 60% yields 1,440, retain 60 to 70% yields 1,400. Arithmetic slightly favours earlier stages, but not because an extra early account is worth more: an extra account at the try stage is worth only 0.5 x 0.6 = 0.3 of an adopter, while an extra retained account is worth a full adopter. Ten points is worth more at try because it is applied to a larger base: 10 points of 10,000 eligible accounts is 1,000 extra triers (x 0.3 = 300 adopters, the 1,500), 10 points of 4,000 triers is 400 extra activated accounts (x 0.6 = 240, the 1,440), and 10 points of 2,000 activated accounts is 200 extra retained accounts (x 1.0 = 200, the 1,400). The differences are small, though (1,500 versus 1,440 versus 1,400 adopters is about one percentage point of adoption between best and worst), so research on which stage is actually feasible to improve decides.
Three stage subproblems plus one enabling workstream:
| Subproblem | Metric (baseline: today's measured value) | Accountable owner | Partner | First hypothesis |
|---|---|---|---|---|
| Try | 40% of eligible accounts start within 14 days | Growth PM | Product marketing | Eligible users never see the entry point |
| Activate | 50% of triers reach the key action | Product Designer (onboarding flow) | Design Researcher (first-use sessions) | The first screen hides the key action |
| Retain | 60% of activated still active in week 4 | PM for core experience | Technical PM, engineering | The key action does not become a habit |
| Enable: measurement | Weekly stage rates exist | Data Analyst | Technical PM | Events are not tracked today |
Run one customer segment end to end first (a vertical slice: one thin end-to-end piece of work, rather than one stage for every segment) rather than every stage for every segment, because averages can hide a segment where adoption is near zero.
Step 3: is the breakdown good enough to start?
Five tests:
- Rebuild: the parts reproduce the whole (0.4 × 0.5 × 0.6 returns 12%). If a stage cannot be written as a rate of the previous one, the tree overlaps or leaves gaps.
- Ownership: each leaf has one metric, one baseline and one owner.
- Monday test: the owner could start tomorrow without asking you what the leaf means.
- Testable: the first test finishes within one to two sprints (a sprint is a fixed one- or two-week work cycle). If not, split again or pick a leading indicator (an early signal that predicts the real outcome, such as week-1 repeat use standing in for week-4 retention).
- Missing branch: ask "if adoption rose but none of our metrics moved, what happened?" Each answer is a missing branch or a measurement gap. For example, if adoption rose from 12% to 14% while the try, activate and retain rates barely moved, the likely answers are a new partner channel that was never counted as eligible, or an event that changed what counts as the key action.
Stop splitting when a leaf passes tests 2 to 4. Going further creates busywork.
Leading the session
Have each function sketch its own branch for ten minutes, then merge and resolve overlaps. Send the VP a one-page summary: definition, funnel, owners, first tests, and what you are deliberately not doing.
Pitfalls
Decomposing alone, giving a leaf two owners or none, and splitting for weeks without starting a test.
When a team plans a large feature, what is the difference between slicing it vertically and slicing it horizontally, and why does it matter? Use a search feature as the example and tell me which you would default to and when you would not.
Sample Answer
Direct answer
A sprint is a fixed work period, often two weeks. A horizontal slice is a piece of work along one technical layer (the database, then the service, then the screen). A vertical slice is a thin piece that cuts through every layer and delivers something a user can actually use. It matters because only vertical slices produce feedback and real risk information early. I default to vertical slicing and avoid it in a few specific cases: a shared foundation that must exist first, specialist teams that need an agreed interface to work in parallel, and work with no user-visible increment.
Search feature, two ways
Terms: the search index is the lookup structure that makes searching fast, the API is the service the screen calls to run a search, typo tolerance matches "shose" to "shoes", and an increment is a piece of working product a user can use.
| Horizontal plan | Vertical plan |
|---|---|
| Sprint 1: build the search index | Slice 1: keyword search on product titles only, simple results list, no filters, working end to end |
| Sprint 2: build the search API | Slice 2: filters (price, category) |
| Sprint 3: build the results screen | Slice 3: typo tolerance and synonyms |
| Nothing is usable until sprint 3 | Slice 4: ranking by popularity |
| Slice 5: autocomplete |
Why it matters
- Early feedback: after slice 1 real people can type queries, and the team learns what they search for before building filters nobody uses.
- Risk shows early: speed with real catalog size, or the index missing half the products, appears in week one, not after three layers are "done".
- Optionality (keeping your choices open): if priorities change after slice 2, the team stopped with a working feature, not three unfinished layers.
- Honest progress: "layers done" says little about whether a user can do anything.
The test for a good slice
Can a real user do something useful with this slice alone? Can it be demoed? Does it touch every layer, even thinly?
When I would not default to vertical
- A risky shared foundation several features will depend on, such as migrating the search engine or a data model hard to change later. I would still make the first step a thin "walking skeleton" (one query end to end) to prove the layers connect.
- Specialist teams working in parallel, where agreeing the interface first (a time-boxed horizontal step) lets them proceed independently. The interface is the agreed contract between the teams, for example: the screen sends the query text and a page number, and the service returns a list of product id, title and price. That is a two-day agreement session, not a two-sprint build.
- Work with no user-visible increment, such as security hardening or compliance changes. A horizontal first step here looks like a one-sprint (two-week) task such as "enforce encrypted connections on the search service", finished and verified on its own, with a visible done-test (for example a scan that reports no unencrypted endpoints).
Second example: a reporting dashboard
Two independently valuable vertical slices: (1) a "weekly sales by region" chart for regional managers, covering the data pull, calculation, chart, access control and export, using data that already exists; (2) a "late deliveries" table for operations, with its own query, definition of "late", alert threshold and permissions. Either can ship alone and be judged by its users, which a plan of "build all queries, then all charts" cannot offer.
Pitfalls
- A "vertical" slice that is really one layer, such as a screen with fake data and no backend.
- Slices so thin they carry no user value (a search box that returns nothing).
- Skipping the foundation for ever, then paying for it in rework.
Explain what makes a breakdown 'MECE' to a new analyst, using a retention decline as the example. Show one split that looks reasonable but overlaps or leaves a gap, then fix it, and say what you would do to pick the most useful way to cut the problem.
Sample Answer
Direct answer
MECE (pronounced "me-see") stands for mutually exclusive, collectively exhaustive. Mutually exclusive means no item belongs to two branches. Collectively exhaustive means every item belongs to at least one branch. Together: each item has exactly one home. A split is only useful if it is MECE and its branches lead to different actions.
Example: monthly retention fell (illustrative numbers)
Retention is the share of customers at the start of the month who are still customers at the end. Say 1,000 customers started the month; retention fell from 80% (800 kept, 200 lost) to 72% (720 kept, 280 lost), so 80 extra customers left.
A split that looks reasonable but is not MECE
Cut the 1,000 customers into: new users (300), mobile users (500), enterprise customers (150), European customers (200).
- Overlap: Priya joined 10 days ago, uses mobile, is on an enterprise plan and lives in Germany. She sits in all four branches. The counts add to 300 + 500 + 150 + 200 = 1,150, more than the 1,000 customers, which is the quick tell.
- Gap: a two-year-old desktop customer on a small-business plan in the US fits none of the four branches.
The fix: one dimension, one rule
Split by tenure at the start of the month: under 30 days (300), 30 to 179 days (400), 180 days or more (300). Each customer has exactly one start date, so exactly one bucket (exclusive), and 300 + 400 + 300 = 1,000 (exhaustive). Other dimensions such as mobile versus desktop then become a second-level split inside each tenure bucket, never mixed in at the same level.
Tests for any split
- Sum test: do the branch counts add to the total?
- One-home test: pick five real items and place each. Two homes means overlap, none means a gap.
- So-what test: does each branch imply a different action? If not, the split is correct but useless.
How I pick the most useful cut
Four families of splitting dimension:
| Family | Example | Why it is MECE |
|---|---|---|
| Formula | Retention = 1 - churn; churn = voluntary cancel + involuntary lapse (failed payment) | An identity, so each churned customer has one recorded final status |
| Process or funnel | Fulfillment throughput: received, picked, packed, shipped | Stages are sequential, so an unshipped order is in exactly one stage. The stage with the longest queue is the bottleneck candidate |
| Segment | Tenure, plan, region | Each customer has one value per attribute |
| Cause | Product, price, competitor | Hardest to keep MECE, since causes interact. Use last, and attribute each loss to a primary cause |
I choose the dimension where (a) branches differ most in outcome, (b) data exists, (c) branches map to owners, and (d) I can check it quickly. For a retention decline I would start with tenure, because "new-user problem or long-time-user problem" decides which team owns it.
Continuing the example (illustrative, same tenure mix last month): last month 60, 80 and 60 customers were lost in the three tenure buckets (200 in total). This month the losses are 140, 80 and 60 (280 in total). The extra 80 sit entirely in the under-30-days bucket, whose loss rate went from 60/300 = 20% to 140/300 = about 47%, while the other two stayed at 20%. That is the so-what: the problem is onboarding, not long-time customers, so the next split goes inside that bucket.
More partitions
- Signup conversion: never started, started and abandoned, submitted but unverified, completed.
- Onboarding non-completers, with a hypothesis per bucket: (1) never started: the prompt is not being seen; (2) started and stalled: one step is too hard or long; (3) completed but not recorded: a tracking gap undercounts completion; (4) ineligible, such as test or internal accounts: the denominator is polluted. Tie rule: eligibility is checked first, so a stalled internal account lands in bucket 4 only.
Pitfalls
- Mixing dimensions at one level (the 1,150 trap).
- An "Other" bucket that quietly holds half the data.
- Treating MECE as a pass mark: a MECE split nobody can act on is still a bad split.
A redesigned onboarding prototype got glowing feedback in interviews but moved activation by nothing. How would you break down the possible reasons for the mismatch and decide which to investigate first?
Sample Answer
Direct answer
The mismatch has three families of explanation: the test did not measure what we think, the interview signal was misleading, or the redesign worked on something that is not the bottleneck. I would investigate in this order: (1) check the experiment and measurement (a data query, hours), (2) look at step-level funnel data for where the redesign did or did not move behaviour, (3) re-read the interview evidence for opinion versus behaviour. Activation means a new user reaching the first meaningful value action, such as completing a first project.
Structured elaboration
A. The test did not measure the change.
- Users never saw the new flow (an exposure bug: the assignment logic put people in the wrong group, or the rollout, the gradual release to a share of users, reached the wrong audience). The new arm, meaning the group of users assigned the redesigned flow, is the group to check first.
- Activation is defined so the redesign cannot move it, or is measured over too short a window.
- Too few users to detect a small effect (with a small sample, random noise can be as large as the real change, so a real improvement looks flat; a sample-size calculation says how many users are needed).
B. The interview signal was misleading.
- Participants are polite and say they like things (social desirability: people give answers that make them look agreeable).
- Prototype in a guided session is not the real flow, with real data and distractions.
- Recruited participants were more motivated than typical signups.
- Questions led toward praise.
C. The redesign improved something that is not the bottleneck (the step where the most potential activations are lost).
- Drop-off lives in a later step.
- Improvement increased completion but with lower-intent users, so later steps got worse (the effect is masked).
Why this order: A is a broken instrument (the measurement setup: tracking, assignment and the metric definition). If it is broken, B and C are unanswerable. It is also the cheapest. C is next because step-level data is already collected. B comes last because it needs new research, and I would run it only on the part that C narrowed.
Worked example
Illustrative numbers. Of 1,000 signups, 800 finished onboarding (80%) and 25% of finishers started a first project, so 200 activated (20%). After the redesign, 900 finish onboarding (90%), but only about 22% of finishers start a project, so about 198 activate, roughly the same 20%. Overall activation shows nothing, yet step-level data shows the redesign did what interviews praised (more people finish) while pulling in lower-intent users at the next step. The next test is a change to the step after onboarding, not further polish on the flow people already like.
A second case, where the redesign simply hit the wrong step (illustrative): of 1,000 signups, 92% finish onboarding before the redesign and 95% after, but only 20% of finishers start a project at the next step either way. Activation goes from 1,000 x 0.92 x 0.20 = 184 to 1,000 x 0.95 x 0.20 = 190, a gain of 6 users, too small to see in a flat topline. The bottleneck is the 20% step, so that is where the next test goes.
Trade-offs and pitfalls
- Do not call the interviews wrong. Interviews measure comprehension and attitude well and behaviour poorly.
- A flat topline (the single overall number on the experiment dashboard) can hide offsetting movements. Always look at steps.
- Do not rerun interviews first. It feels natural to a researcher, but it is the most expensive way to learn nothing new.
- What would change my order: if the experiment dashboard (the tool showing how many users are in each group and the metric for each) showed fewer users than planned in the new arm, I would stop and fix exposure before anything else.
Unlock Full Question Bank
Get access to all 11 Structured Problem Solving and Decomposition interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.