Structured Problem Solving and Decomposition Questions
Methodical problem solving for open-ended and ambiguous situations once the problem is defined: decomposing a goal or problem into mutually exclusive, collectively exhaustive parts (issue trees, metric trees and driver breakdowns, work breakdown into subproblems and vertical slices), forming and prioritizing hypotheses (hypothesis trees and funnels, including laying out candidate explanations for a metric drop or model degradation and the cheapest test for each branch), choosing an analytical approach, and reasoning to a recommendation. Also covers turning a vague mandate into measurable, testable subproblems with owners, breaking a large initiative into workstreams, mapping dependencies, sequencing the work, deciding the first deliverable and what to defer, structuring plans that mix research, analytics and experiments, and talking through real examples of cutting a messy problem into parts. Covers explaining and adapting structured problem-solving methods across contexts, choosing and switching methods, and coaching others to structure ambiguity. Excludes turning a vague request into a scoped problem statement, named-framework business cases, root-cause techniques for failures and metric movements (running the diagnosis itself), prioritization scoring and trade-off decisions, deciding how to act under incomplete information, market sizing and estimation, and framing machine-learning problems, which are covered elsewhere.
You have five workstreams: auth, payments, client validation, analytics and fraud detection, with dependencies among them. How would you decide what order to build them in, where the fastest learning is, and what could run in parallel?
Sample Answer
Direct answer
Separate two orders that are often confused: the order things must be built in (set by dependencies) and the order you learn in (set by uncertainty). Build order follows the critical path; learning order starts the riskiest unknowns early with cheap spikes (short, time-boxed experiments that answer one question and are thrown away, not production code), even if the full build comes later. Anything with no dependency between them runs in parallel.
Structured elaboration
Assumed dependencies (state them, then confirm with the team)
- Auth (who the user is) feeds payments, because a payment belongs to an identity.
- Payments produces the transactions that fraud detection scores.
- Analytics events are inputs to fraud features, and analytics can start with instrumentation early.
- Client validation (checking input in the app before sending it) depends on nothing but helps payments forms.
Critical path is the longest chain of dependent work: auth, then payments, then fraud detection. Everything off that chain has slack (time it can slip without delaying the end).
Worked example
Illustrative durations in weeks: auth 2, payments 3, fraud detection 4, analytics 2, client validation 1.
- Critical path: 2 + 3 + 4 = 9 weeks.
- Schedule on the critical path: auth weeks 1-2, payments weeks 3-5, fraud detection weeks 6-9.
- Analytics must be finished by the end of week 5, when fraud starts. It takes 2 weeks, so its latest start is week 4; starting in week 1 instead gives 4 - 1 = 3 weeks of slack.
- Client validation must be finished by the end of week 5, when payments goes live. It takes 1 week, so its latest start is week 5; starting in week 1 gives 5 - 1 = 4 weeks of slack.
- So weeks 1-2 run auth, analytics and (if capacity allows) validation in parallel.
Fastest learning. Fraud detection is last to build but the most uncertain: is there signal, how are labels obtained? Run a shallow offline analysis on historical data and a payments-provider sandbox spike in week 1, while building in dependency order.
The same method on two other problems
- Five ML workstreams, four weeks, one engineer. One engineer means no parallelism, so cut each to a thin slice and sequence by what unblocks and teaches most. For example (data pipeline, labeling, training, serving, monitoring): week 1 data plus a minimal labeled set and a baseline, week 2 a first model with offline evaluation, week 3 serving, week 4 hardening with monitoring reduced to logging inputs and predictions, full alerting deferred.
- Model quality versus a 100 ms latency budget. Latency is the time from request to response; budget it per stage, illustratively 20 for network and orchestration (the code that coordinates calls between services), 30 for feature fetch (looking up the input values the model needs), 35 for inference (the model computing its prediction), 15 for post-processing (total 100). Levers: architecture (a smaller model, or a distilled one trained to imitate a larger one), features (precompute expensive ones), caching (reuse answers for repeated inputs, at the cost of staleness), batching (grouping several requests into one model call: higher throughput, meaning more requests handled per second, but waiting to fill a batch eats budget). Measure the quality lost per millisecond saved (illustratively, a distilled model that cuts inference from 35 to 20 ms, 15 ms saved, while accuracy falls from 91% to 90%, costs about 0.07 points per ms; caching that saves 10 ms at no accuracy cost should be taken first) and judge at the 99th percentile, the slowest 1% of requests, not the average.
Trade-offs and pitfalls
- Dependency order can hide the riskiest item at the end; spikes fix that without changing build order.
- Parallel work needs people; with fewer people than branches, the critical path still goes first.
- A mocked upstream lets parallel work start, but integration surprises move to the end unless you schedule an early integration.
- What would flip the order: if fraud signal turns out to be weak in the spike, you reconsider the whole workstream before payments is built.
Your VP gives you the mandate 'improve product adoption' and nothing else. How would you lead the team to turn that into measurable subproblems with owners, and how would you know the breakdown is good enough to start work?
Sample Answer
Direct answer
Run a working session, not a solo analysis. The team agrees one definition of adoption, breaks it into stages whose rates multiply back to the whole, gives every stage one metric with a baseline and one owner, and then stress-tests the breakdown before anyone builds. The VP confirms the definition. The team draws the branches.
Step 1: define adoption briefly
Adoption = the share of eligible accounts that reach the key value action (the first action that shows a user got real value, for example "shared a first report") and still use it in week 4. Eligible means accounts that could use the product. Here you need only enough definition to measure the goal.
Step 2: break it down with real numbers (illustrative)
Of 10,000 eligible new accounts: 4,000 try the product (40%), 2,000 of those reach the key action (50%), and 1,200 of those are still active in week 4 (60%). Check: 10,000 × 0.4 × 0.5 × 0.6 = 1,200, so adoption is 12%.
Because the stages multiply, +10 points at any one stage gives: try 40 to 50% yields 1,500 adopters, activate 50 to 60% yields 1,440, retain 60 to 70% yields 1,400. Arithmetic slightly favours earlier stages, but not because an extra early account is worth more: an extra account at the try stage is worth only 0.5 x 0.6 = 0.3 of an adopter, while an extra retained account is worth a full adopter. Ten points is worth more at try because it is applied to a larger base: 10 points of 10,000 eligible accounts is 1,000 extra triers (x 0.3 = 300 adopters, the 1,500), 10 points of 4,000 triers is 400 extra activated accounts (x 0.6 = 240, the 1,440), and 10 points of 2,000 activated accounts is 200 extra retained accounts (x 1.0 = 200, the 1,400). The differences are small, though (1,500 versus 1,440 versus 1,400 adopters is about one percentage point of adoption between best and worst), so research on which stage is actually feasible to improve decides.
Three stage subproblems plus one enabling workstream:
| Subproblem | Metric (baseline: today's measured value) | Accountable owner | Partner | First hypothesis |
|---|---|---|---|---|
| Try | 40% of eligible accounts start within 14 days | Growth PM | Product marketing | Eligible users never see the entry point |
| Activate | 50% of triers reach the key action | Product Designer (onboarding flow) | Design Researcher (first-use sessions) | The first screen hides the key action |
| Retain | 60% of activated still active in week 4 | PM for core experience | Technical PM, engineering | The key action does not become a habit |
| Enable: measurement | Weekly stage rates exist | Data Analyst | Technical PM | Events are not tracked today |
Run one customer segment end to end first (a vertical slice: one thin end-to-end piece of work, rather than one stage for every segment) rather than every stage for every segment, because averages can hide a segment where adoption is near zero.
Step 3: is the breakdown good enough to start?
Five tests:
- Rebuild: the parts reproduce the whole (0.4 × 0.5 × 0.6 returns 12%). If a stage cannot be written as a rate of the previous one, the tree overlaps or leaves gaps.
- Ownership: each leaf has one metric, one baseline and one owner.
- Monday test: the owner could start tomorrow without asking you what the leaf means.
- Testable: the first test finishes within one to two sprints (a sprint is a fixed one- or two-week work cycle). If not, split again or pick a leading indicator (an early signal that predicts the real outcome, such as week-1 repeat use standing in for week-4 retention).
- Missing branch: ask "if adoption rose but none of our metrics moved, what happened?" Each answer is a missing branch or a measurement gap. For example, if adoption rose from 12% to 14% while the try, activate and retain rates barely moved, the likely answers are a new partner channel that was never counted as eligible, or an event that changed what counts as the key action.
Stop splitting when a leaf passes tests 2 to 4. Going further creates busywork.
Leading the session
Have each function sketch its own branch for ten minutes, then merge and resolve overlaps. Send the VP a one-page summary: definition, funnel, owners, first tests, and what you are deliberately not doing.
Pitfalls
Decomposing alone, giving a leaf two owners or none, and splitting for weeks without starting a test.
You inherit a roadmap with 20 or more subprojects across three quarters and no clear order. How do you decompose and resequence it into a plan that balances bets, customer needs and capacity?
Sample Answer
Direct answer
I would not try to rank 22 projects against each other directly. I would first put every project into exactly one purpose bucket, map dependencies and capacity, and then sequence the buckets against three constraints: promises with dates, enablers that unlock other work, and a limited number of uncertain bets running at once. The output is a plan whose total load fits capacity with room left over, plus a short list of what is explicitly deferred and why.
Structured elaboration
- Inventory and normalise. For each subproject write one line: the outcome it delivers, the owner, a rough size in squad-quarters (one team working for one quarter), what it depends on, and whether anyone outside the team is waiting on it.
- Decompose by primary purpose so every project lands in exactly one bucket (mutually exclusive) and none is left out (collectively exhaustive, together called MECE):
- Commitments: contractual, regulatory or customer-promised, usually with a date. Example: EU data residency for a signed customer, due 31 March.
- Enablers: platform or infrastructure work whose value is unlocking other projects. Example: rebuilding event tracking, which usage dashboards and an AI-summaries feature both need.
- Bets: uncertain growth or product ideas whose payoff is unknown. Example: AI-generated summaries in the dashboard.
- Keep-the-lights-on: reliability, maintenance and debt. Example: a database upgrade that must finish before vendor support ends.
- Draw the dependency graph and mark the longest chain. Anything on it that starts late delays the end date.
- Compute capacity honestly. Subtract planned leave and ongoing support first, then plan to about 80% of what remains and keep the rest for surprises. The 80% is a planning assumption to tune per team, not a benchmark.
- Sequence. Early: long-lead enablers and commitments with early dates, plus small experiments that shrink the uncertainty of the biggest bets. Middle: commit to the bets that survived. Late: items with slack. Limit how many bets run in parallel (for example three at once with three squads, so each gets roughly a squad's attention; the limit is illustrative and tuned to the team) so each gets enough attention to learn something.
- Cross-cadence teams. If one team releases weekly and another quarterly, tie them together through agreed interfaces, not shared dates. The weekly team ships its part behind a feature flag (a switch that keeps new behaviour off) before the quarterly team's cutoff, and the quarterly release is the single date at which flags turn on. Write the cutoff and the interface version into both plans.
- Scaling the method down to a three-month analytics roadmap. Three tracks: early wins in month one (fix the dashboards people already distrust), mid-term in months two and three (the core metric definitions and models), and an infrastructure track running alongside (pipeline reliability and data quality checks) so the mid-term work does not sit on a weak base. Each track gets a capacity share, not a start date alone.
Worked example
Three squads for three quarters give 9 squad-quarters. Planned leave and ongoing support take about 1.0 (illustrative), leaving 8.0. Planning at 80% gives 8.0 x 0.8 = 6.4.
| Bucket | Projects | Demand (squad-quarters) | Scheduled | Deferred |
|---|---|---|---|---|
| Commitments | 5 | 2.5 | 5 projects, 2.5 | none |
| Enablers | 4 | 2.0 | 3 projects, 1.5 | 1 project, 0.5 |
| Keep-the-lights-on | 3 | 1.0 | 3 projects, 1.0 | none |
| Bets | 10 | 4.0 | 3 projects, 1.2 | 7 projects, 2.8 |
| Total | 22 | 9.5 | 14 projects, 6.2 | 8 projects, 3.3 |
Each of the 22 projects sits in exactly one purpose row. The usage dashboards named in the sequence below are one of the five commitments (promised to a customer in a contract), which is why they are scheduled in quarter 2 and are not counted among the three scheduled bets. Scheduled load is 2.5 + 1.5 + 1.0 + 1.2 = 6.2, under the 6.4 ceiling. Total demand is 9.5 against 9.0 raw capacity, which is why 8 projects are deferred (0.5 + 2.8 = 3.3, and 9.5 - 3.3 = 6.2). The conversation with leadership is then concrete: to add a deferred project, name what moves out.
Sequence for the scheduled projects:
- Quarter 1: EU data residency (the commitment with the earliest date), the event-tracking rebuild (the enabler the dashboards and the summaries bet both wait on), the database upgrade, and one cheap experiment on the AI-summaries bet.
- Quarter 2: the remaining commitments, the usage dashboards that need the new event data, and a go or stop decision on the summaries experiment.
- Quarter 3: the AI-summaries bet if it survived, the two other scheduled bets, and slack for incidents.
The longest dependency chain is event rebuild (Q1), then dashboards (Q2), then summaries (Q3), so a slip in the first moves the last.
Trade-offs and pitfalls
- Sequencing by who asks loudest rewards politics. Buckets plus capacity make trade-offs visible.
- A plan loaded to 100% breaks on the first incident. The buffer is a feature.
- If an enabler has only one dependent project, question whether it is an enabler at all or just part of that project.
- This would change if a commitment's date moved: commitments are the bucket I treat as fixed by default. A keep-the-lights-on item with a hard external date, such as the database upgrade that must finish before vendor support ends, is treated the same way.
You have six weeks to deliver a prototype of multimodal search. Break the work into workstreams that do not overlap or leave gaps, name the critical dependencies, and define the minimum viable version.
Sample Answer
Direct answer
Multimodal search means search across more than one type of content or query, here assumed to be text and image queries returning images from a product catalog. Four terms recur: retrieval is finding the items that match a query; an embedding model converts text or an image into a list of numbers so that similar items sit close together; an index is the lookup structure that finds the nearest items quickly; ranking is ordering the matches so the best come first. I would split the work by what each workstream produces, so each deliverable has one owner and the list covers the whole prototype. Critical path: data, then retrieval, then API, then demo, with evaluation running alongside and gating the model choice. The minimum viable version (the smallest version that still proves the core idea) is text-to-image search on a fixed catalog that meets an agreed quality bar, with image-upload search as the stretch.
Structured elaboration
Workstreams (mutually exclusive by output, collectively exhaustive for a demo):
| # | Workstream | Produces | Also owns (so cross-cutting items are not orphaned) |
|---|---|---|---|
| 1 | Data and corpus | cleaned catalog, image ingestion, metadata | image licensing and privacy check |
| 2 | Retrieval core | embedding model choice, index, ranking | latency budget (the longest one search may take before users notice the wait) |
| 3 | Query experience | API and UI for text box, image upload, results | usability smoke test |
| 4 | Evaluation | judged query set (test searches with human-marked correct results), quality metric (one score for how often search finds them), test plan | regression checks each week |
| 5 | Platform | hosting, deployment, cost tracking | access and demo environment |
Critical dependencies:
- The evaluation set (workstream 4) must exist by end of week 2, or the embedding comparison has no yardstick.
- The corpus (1) must be ingested before the index (2) can be built.
- The API contract (the agreed request and response format between the UI and the search backend; owned by workstream 3, signed off by workstream 2) must be fixed by week 3 so UI work does not wait.
- Platform (5) must provide a stable environment before week 5 integration.
Milestones with acceptance criteria and test owners:
| Week | Milestone | Acceptance criterion | Tested by |
|---|---|---|---|
| 1 | Corpus plan, evaluation set drafted, API contract draft | catalog loaded; 100 judged queries (illustrative size) | Evaluation owner |
| 2 | Text baseline plus embedding comparison | one model chosen on the judged set | Evaluation owner |
| 3 | Index and API; midpoint go/no-go (a checkpoint where the team decides to continue or stop) | queries return results end to end | Retrieval owner |
| 4 | UI with text and image upload | demo path works unaided | Query experience owner |
| 5 | Integration, quality and latency checks | meets quality bar agreed with the product lead | Evaluation and Platform owners |
| 6 | Demo, hardening, handoff notes | demo script runs three times cleanly | Whole team |
Minimum viable version: text query to top 10 image results over the fixed catalog, hitting the agreed quality bar (for example, one judged query is "red leather ankle boots" with 12 catalog items marked relevant; the bar could be at least one relevant item in the top 10 results for 80 of the 100 judged queries, with the number illustrative and set with the product lead). Image-upload search is added only if the week 3 go/no-go shows the index and API on schedule (the week 4 check confirms it).
Communication plan for the senior lead: a one-page update each week (done, next, risks, decisions needed from you), a go/no-go review at the end of week 3, and the demo in week 6.
Image-upload decision point. The week 4 milestone builds the text path first and the image-upload path only if the week 3 go/no-go shows the index and API on schedule. If week 3 slips, upload is dropped from the week 4 milestone and from the minimum viable version, and the week 4 acceptance criterion becomes "text query demo path works unaided". So the single decision point for the stretch is the end of week 3, and the week 4 check confirms it rather than deciding it. Likewise, the week 1 row means the catalog extract is loaded and the full ingestion and cleaning plan is written, with the full ingestion finished before the index build in week 3.
Trade-offs and pitfalls
- Splitting by technology layer alone leaves cross-cutting issues (latency, licensing, testing) with no owner. The right-hand column of the table prevents that.
- Building the UI before the evaluation set produces a demo that looks good and cannot be judged.
- Scope pressure: if week 3 slips, drop image upload before cutting evaluation.
Nightly pipelines deliver on time 85% of the time and the goal is 99% in three months. How would you structure the work: what to learn first, what subproblems to split off, and how you would explain the plan to leadership in one page?
Sample Answer
Direct answer
Treat "85% to 99%" as a measurement problem first and an engineering problem second. Spend the first week pinning down what "on time" means and why the late runs were late, then split the work along the causes that account for the most late runs, and give leadership a one-page plan with a baseline, a few workstreams, monthly targets and the decisions you need from them.
Structured elaboration
What to learn first (week 1)
- The definition. A pipeline is a scheduled job that moves and transforms data. "On time" needs a deadline (say 06:00), a measurement point (job finished, or data actually visible to consumers) and a scope (all pipelines, or only the ones people depend on). Count in one unit: a pipeline-night is one pipeline's run on one night.
- Who is hurt by lateness. Tier the pipelines. One feeding a 07:00 executive dashboard matters more than an archive job. The 99% may only need to hold for the top tier, which is a question to settle with leadership early.
- A cause log. For every late pipeline-night in the last several weeks, record the first thing that went wrong. Rule: each late run goes in exactly one bucket, so the buckets cannot overlap.
Subproblems to split off. The split below is MECE (mutually exclusive: no run in two buckets; collectively exhaustive: every late run lands somewhere), because the first-cause rule decides placement in time order.
| Subproblem | What it means | Typical levers |
|---|---|---|
| Started late | Upstream data or a dependency was not ready at start time | Agreed delivery times with upstream owners, sensors (checks that wait for the upstream data to appear and then start the job, instead of running at a fixed clock time), fallback to last good data |
| Failed | The run errored and recovery took too long | Automatic retries, safe reruns (idempotent: running twice gives the same result), alert routing to on-call (the engineer currently responsible for responding to pages) |
| Ran too long | Started on time but overran the window | Profile the slowest stage (measure where the time goes), fix skew (one worker getting far more data than the others, so everything waits for it), resize resources (give the job more compute), reorder jobs on the critical path (the longest chain of jobs that depend on each other, which sets the finish time) |
Worked example
All numbers are illustrative. Take 20 pipelines over 30 nights: 600 pipeline-nights. At 85% on time, 510 are on time and 90 are late. At 99%, at most 6 may be late, so about 84 of the 90 late runs (roughly 93%) must stop happening.
Cause log: started late 45 (upstream late 36, dependency chain with no slack 9 (job B waits for job A with no spare time between them, so any delay in A passes straight through)), failed 27, ran too long 18. The total is 90. So half the problem (45 of 90) is upstream timing and 30 percent (27 of 90) is failures. That ordering decides the plan: fixing run time first would attack only 20% of the lateness.
One run traced through the rule: a pipeline due at 06:00 gets its upstream data at 05:30 instead of 04:00, so it starts late, and it then also fails on a timeout. It is logged once, under started late, because that was the first thing that went wrong; the failure goes in a notes column. Counting it in both buckets would make the buckets sum to more than the 90 late runs.
Monthly targets: 91% means at most 54 late runs, 96% at most 24, 99% at most 6.
The one page for leadership (excerpt):
- Goal: tier-1 pipelines on time 99% of nights by month 3, measured as data visible by 06:00. Baseline today: 85% across all 20 pipelines (the tier-1 baseline comes from cutting the same cause log by tier, and the 99% applies to that tier-1 slice).
- Diagnosis: 90 late runs in 30 nights: 50% started late, 30% failed, 20% ran long.
- Workstreams: (1) upstream delivery agreements and fallback data, (2) failure recovery and alerting, (3) runtime on the critical path.
- Measure: weekly on-time rate per tier, plus late runs by cause.
- Milestones: 91% month 1, 96% month 2, 99% month 3.
- Asks and risks: a named owner on each upstream team; if upstream will not commit, we deliver stale-marked data on time instead of fresh data late.
The same one-page shape for a latency goal such as cutting data latency by 50% in six months. Define latency as source event to queryable data, measured at the 95th percentile (sort the runs from fastest to slowest; 95 of every 100 finish at or under this value, so it shows the slow tail rather than the typical run). Break it by stage and give each a budget. Illustrative: ingest 15, transform 30, load 10, refresh 5 minutes (60 total) becomes 8, 14, 5, 3 (30 total). Caveat: the 95th percentile of the total is not the sum of the stage 95th percentiles. With independent, light-tailed stages the sum usually overstates it, because the slowest run in one stage is rarely the slowest in every stage; with heavy-tailed stages (a few very long runs) it can understate it, so neither direction is guaranteed. Treat the stage budgets as targets and measure the end-to-end 95th percentile directly. The page then lists goal, one component per stage, a measure per stage and checkpoints roughly every month.
Trade-offs and pitfalls
- A single average hides the tail: report on-time rate per tier, not one blended number.
- Do not promise 99% before the cause log exists. After week 1, state the plan with real data.
- Upstream lateness cannot be fixed by your team alone, so the plan must contain a fallback you control.
- Fixing the loudest incident instead of the largest bucket is the common wrong turn.
- If two buckets turn out to share one cause (for example a shared cluster), say so on the page and treat it as one workstream.
Unlock Full Question Bank
Get access to all 11 Structured Problem Solving and Decomposition interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.