Structured Problem Solving and Decomposition Questions
Methodical problem solving for open-ended and ambiguous situations once the problem is defined: decomposing a goal or problem into mutually exclusive, collectively exhaustive parts (issue trees, metric trees and driver breakdowns, work breakdown into subproblems and vertical slices), forming and prioritizing hypotheses (hypothesis trees and funnels, including laying out candidate explanations for a metric drop or model degradation and the cheapest test for each branch), choosing an analytical approach, and reasoning to a recommendation. Also covers turning a vague mandate into measurable, testable subproblems with owners, breaking a large initiative into workstreams, mapping dependencies, sequencing the work, deciding the first deliverable and what to defer, structuring plans that mix research, analytics and experiments, and talking through real examples of cutting a messy problem into parts. Covers explaining and adapting structured problem-solving methods across contexts, choosing and switching methods, and coaching others to structure ambiguity. Excludes turning a vague request into a scoped problem statement, named-framework business cases, root-cause techniques for failures and metric movements (running the diagnosis itself), prioritization scoring and trade-off decisions, deciding how to act under incomplete information, market sizing and estimation, and framing machine-learning problems, which are covered elsewhere.
A junior analyst on your team freezes when problems are ambiguous. How would you teach them to structure a problem, and how would you know it is working?
Sample Answer
Direct answer
Freezing usually means the junior has no safe first move, not that they lack ability. I would teach one small, repeatable routine, model it myself out loud, have them do it on low-stakes problems with a checkpoint at each step, and fade my involvement. I would know it is working from observable changes: they bring a structure sooner, their hypotheses are testable, and they need fewer rescues.
Structured elaboration
First, diagnose with a short conversation: what do they do in the first ten minutes of an ambiguous ask? Common causes are fear of being wrong, an unclear boundary of the task, or too many possible starts.
The routine (five steps, each with a coaching checkpoint)
- Write the question as one sentence including the decision it supports. Checkpoint: I read it before any data work.
- Draw a tree with three or four branches that do not overlap and together cover the question (MECE: mutually exclusive, collectively exhaustive). Checkpoint: I check that every idea they have fits exactly one branch.
- Write a guess for each branch, a hypothesis, plus what data would disprove it. Checkpoint: each is falsifiable (it can be shown wrong by some specific data), with at least three in total. Wrong guesses are allowed and expected; unfalsifiable ones are not.
- Falsifiable: "Users who skip the setup screen are less likely to return in week two. This is wrong if skippers and non-skippers return at about the same rate."
- Not falsifiable: "Onboarding could be better." No data could ever show it wrong, so there is nothing to test.
- Pick the cheapest test first and predict its result. Checkpoint: I review the plan before they run it.
- Time-box and share: a time-box is a fixed amount of time set in advance (for example, 90 minutes). When it ends they show whatever they have, even half-formed.
Applying the routine to a concrete case: users not returning in week two
For "why do users in our new vertical not return in week two?", the junior builds branches by funnel stage (the ordered steps a user passes through, from arriving to coming back): wrong users arrive, right users never reach value, users reach value but have no reason to return, and the measurement itself is off. Under each, they write hypotheses and the disproving data. A short excerpt a junior might bring:
| Branch | Hypothesis | Disproved if |
|---|---|---|
| Right users never reach value | Fewer than half of new users finish the first task | Most new users finish it and still do not return |
| Reach value, no reason to return | Users who finish the task once have nothing pending, so no prompt to come back | Users with pending items return no more often than users without |
| Measurement is off | The week-two event stopped logging on one platform | Event counts per platform are flat across the release |
Coaching checkpoints: review the tree before any query, review the hypothesis list, then review the first result and what it changes.
Worked example of progress
Week 1: they bring nothing for two days, then a long query. Week 3: within a few hours they bring a one-page tree and three hypotheses, one of which is rejected by a ten-minute check. Week 6: they bring the tree to me already tested against one hypothesis and ask a sharper question ("branch two looks right, should I get session recordings?").
How I would know it is working
- Time to first written structure falls from their baseline.
- Hypotheses are falsifiable and each has a stated test.
- Fewer "what do I do next" messages; more "here is what I plan and why".
- Fewer restarts caused by a misunderstood question.
- I ask them to teach the routine to a peer in a month; being able to explain it is a sign it has stuck.
Trade-offs and pitfalls
- Handing over my own tree teaches them my tree, not how to make one.
- Praising only correct answers makes freezing worse. I praise a clear structure and a well-chosen test, even when the guess turns out wrong.
- If after several weeks nothing changes, I look at the work environment (is it safe to be wrong here?) before assuming a skills problem.
You own a nine-month, revenue-critical migration to the cloud. How would you split it into top-level workstreams that cover everything without overlap, and how would you decide the order and the dependencies between them?
Sample Answer
Direct answer
Split by one dimension only, the kind of deliverable each workstream produces, so every task has exactly one home. I would use six workstreams: discovery and planning, cloud foundation, workload migration, operations readiness, decommission and exit, and program and change management. Then order them by the critical path (the longest chain of dependent work, which sets the earliest finish date) and move the revenue-critical systems last, after the process has been proven on safer ones.
Structured elaboration
Why MECE matters here. MECE means mutually exclusive (no task in two workstreams) and collectively exhaustive (nothing falls between them). Overlap causes double work and blurred ownership; gaps are where a revenue-critical migration gets hurt late. Test it: take ten random tasks (set up VPN, train support staff, shut off the old database server) and confirm each lands in exactly one workstream. Then ask "where do security, cost and training live?" and make sure each has one named home plus an explicit boundary.
Terms used in the table: a wave is a group of applications migrated together, and the wave plan says which applications go in which wave and when. Identity means who can sign in and with what permissions. The security baseline is the minimum set of security settings every account must have. Tagging is attaching labels (team, cost center) to cloud resources so cost and owners can be traced. A region is a geographic location where the cloud provider runs data centers. Runbooks are step-by-step instructions for handling a known situation. Decommissioning is shutting down and removing the old systems. Rollback is returning an application to the old environment if the move fails.
| Workstream | Covers | Boundary rule |
|---|---|---|
| 1 Discovery and planning | Inventory of applications, dependencies and data; target design; wave plan | Produces plans, builds nothing |
| 2 Cloud foundation | Accounts, network connectivity, identity, security baseline, cost budgets and tagging (the "landing zone", the pre-built environment workloads move into) | Environment only; per-application security checks belong to workstream 3 |
| 3 Workload migration | Moving applications and their data in waves, including cutover (the moment traffic switches) and rollback plans | One owner per application |
| 4 Operations readiness | Monitoring, alerting, runbooks, on-call, deployment pipelines, cost review routines | Defines how the new environment is run; the shared runbook format, on-call and alerting live here, while each application's own rollback steps stay in workstream 3 |
| 5 Decommission and exit | Retiring old systems, data retention, contract and facility exit | Starts only after a wave is stable |
| 6 Program and change management | Governance, risk register, communication, training | Cross-cutting, owns no technical deliverable |
Worked example
Illustrative nine-month order (months 1-9):
| Workstream | Months | Depends on |
|---|---|---|
| Discovery and planning | 1-2 | None |
| Cloud foundation | 1-3 | Early target design from discovery (regions, compliance needs) |
| Operations readiness | 2-5 | Foundation; minimum version before the first production workload |
| Workload migration | pilot month 4, waves in months 5-8 | Foundation and operations readiness |
| Decommission and exit | 7-9 | Each wave stable first |
| Program and change management | 1-9 | None |
Operations readiness has a minimum version that must exist before the month-4 pilot: monitoring and alerts on the pilot app, a tested rollback runbook for the pilot app (its steps come from the workstream 3 rollback plan; operations readiness checks that on-call can actually run them) and a named on-call person. It is finished by the end of month 3; the rest (cost review routines, deployment pipelines for later waves) continues through month 5 alongside the early waves.
Waves: pilot (a low-risk internal app, month 4), wave 1 low risk (month 5), wave 2 (month 6), wave 3 medium risk (month 7), wave 4 revenue-critical systems (month 8), leaving month 9 as buffer and for finishing decommissioning. Order criteria: start long-lead items first (tasks with slow waiting time that cannot be compressed: network connectivity, compliance approvals meaning sign-offs that the new setup meets legal or regulatory rules, large data transfers), keep an easy rollback until each wave is stable, and avoid cutovers in business peak periods.
Trade-offs and pitfalls
- Splitting by team or by application type mixes dimensions and breaks MECE; a mixed split looks fine on a slide and fails at the first overlapping task.
- Moving the revenue-critical system last trades a later benefit for lower risk. If the cost of the old environment is the burning problem, you might move one revenue-critical but simple component earlier with a tested rollback.
- Big-bang cutover is the wrong turn for revenue-critical systems; migrate in waves with a rollback path.
- Operations readiness is the workstream most often forgotten, and the first production incident exposes it.
- What would flip the order: a hard external date (data center contract ending) moves decommission planning to the front and compresses waves.
You inherit a roadmap with 20 or more subprojects across three quarters and no clear order. How do you decompose and resequence it into a plan that balances bets, customer needs and capacity?
Sample Answer
Direct answer
I would not try to rank 22 projects against each other directly. I would first put every project into exactly one purpose bucket, map dependencies and capacity, and then sequence the buckets against three constraints: promises with dates, enablers that unlock other work, and a limited number of uncertain bets running at once. The output is a plan whose total load fits capacity with room left over, plus a short list of what is explicitly deferred and why.
Structured elaboration
- Inventory and normalise. For each subproject write one line: the outcome it delivers, the owner, a rough size in squad-quarters (one team working for one quarter), what it depends on, and whether anyone outside the team is waiting on it.
- Decompose by primary purpose so every project lands in exactly one bucket (mutually exclusive) and none is left out (collectively exhaustive, together called MECE):
- Commitments: contractual, regulatory or customer-promised, usually with a date. Example: EU data residency for a signed customer, due 31 March.
- Enablers: platform or infrastructure work whose value is unlocking other projects. Example: rebuilding event tracking, which usage dashboards and an AI-summaries feature both need.
- Bets: uncertain growth or product ideas whose payoff is unknown. Example: AI-generated summaries in the dashboard.
- Keep-the-lights-on: reliability, maintenance and debt. Example: a database upgrade that must finish before vendor support ends.
- Draw the dependency graph and mark the longest chain. Anything on it that starts late delays the end date.
- Compute capacity honestly. Subtract planned leave and ongoing support first, then plan to about 80% of what remains and keep the rest for surprises. The 80% is a planning assumption to tune per team, not a benchmark.
- Sequence. Early: long-lead enablers and commitments with early dates, plus small experiments that shrink the uncertainty of the biggest bets. Middle: commit to the bets that survived. Late: items with slack. Limit how many bets run in parallel (for example three at once with three squads, so each gets roughly a squad's attention; the limit is illustrative and tuned to the team) so each gets enough attention to learn something.
- Cross-cadence teams. If one team releases weekly and another quarterly, tie them together through agreed interfaces, not shared dates. The weekly team ships its part behind a feature flag (a switch that keeps new behaviour off) before the quarterly team's cutoff, and the quarterly release is the single date at which flags turn on. Write the cutoff and the interface version into both plans.
- Scaling the method down to a three-month analytics roadmap. Three tracks: early wins in month one (fix the dashboards people already distrust), mid-term in months two and three (the core metric definitions and models), and an infrastructure track running alongside (pipeline reliability and data quality checks) so the mid-term work does not sit on a weak base. Each track gets a capacity share, not a start date alone.
Worked example
Three squads for three quarters give 9 squad-quarters. Planned leave and ongoing support take about 1.0 (illustrative), leaving 8.0. Planning at 80% gives 8.0 x 0.8 = 6.4.
| Bucket | Projects | Demand (squad-quarters) | Scheduled | Deferred |
|---|---|---|---|---|
| Commitments | 5 | 2.5 | 5 projects, 2.5 | none |
| Enablers | 4 | 2.0 | 3 projects, 1.5 | 1 project, 0.5 |
| Keep-the-lights-on | 3 | 1.0 | 3 projects, 1.0 | none |
| Bets | 10 | 4.0 | 3 projects, 1.2 | 7 projects, 2.8 |
| Total | 22 | 9.5 | 14 projects, 6.2 | 8 projects, 3.3 |
Each of the 22 projects sits in exactly one purpose row. The usage dashboards named in the sequence below are one of the five commitments (promised to a customer in a contract), which is why they are scheduled in quarter 2 and are not counted among the three scheduled bets. Scheduled load is 2.5 + 1.5 + 1.0 + 1.2 = 6.2, under the 6.4 ceiling. Total demand is 9.5 against 9.0 raw capacity, which is why 8 projects are deferred (0.5 + 2.8 = 3.3, and 9.5 - 3.3 = 6.2). The conversation with leadership is then concrete: to add a deferred project, name what moves out.
Sequence for the scheduled projects:
- Quarter 1: EU data residency (the commitment with the earliest date), the event-tracking rebuild (the enabler the dashboards and the summaries bet both wait on), the database upgrade, and one cheap experiment on the AI-summaries bet.
- Quarter 2: the remaining commitments, the usage dashboards that need the new event data, and a go or stop decision on the summaries experiment.
- Quarter 3: the AI-summaries bet if it survived, the two other scheduled bets, and slack for incidents.
The longest dependency chain is event rebuild (Q1), then dashboards (Q2), then summaries (Q3), so a slip in the first moves the last.
Trade-offs and pitfalls
- Sequencing by who asks loudest rewards politics. Buckets plus capacity make trade-offs visible.
- A plan loaded to 100% breaks on the first incident. The buffer is a feature.
- If an enabler has only one dependent project, question whether it is an enabler at all or just part of that project.
- This would change if a commitment's date moved: commitments are the bucket I treat as fixed by default. A keep-the-lights-on item with a hard external date, such as the database upgrade that must finish before vendor support ends, is treated the same way.
Nightly pipelines deliver on time 85% of the time and the goal is 99% in three months. How would you structure the work: what to learn first, what subproblems to split off, and how you would explain the plan to leadership in one page?
Sample Answer
Direct answer
Treat "85% to 99%" as a measurement problem first and an engineering problem second. Spend the first week pinning down what "on time" means and why the late runs were late, then split the work along the causes that account for the most late runs, and give leadership a one-page plan with a baseline, a few workstreams, monthly targets and the decisions you need from them.
Structured elaboration
What to learn first (week 1)
- The definition. A pipeline is a scheduled job that moves and transforms data. "On time" needs a deadline (say 06:00), a measurement point (job finished, or data actually visible to consumers) and a scope (all pipelines, or only the ones people depend on). Count in one unit: a pipeline-night is one pipeline's run on one night.
- Who is hurt by lateness. Tier the pipelines. One feeding a 07:00 executive dashboard matters more than an archive job. The 99% may only need to hold for the top tier, which is a question to settle with leadership early.
- A cause log. For every late pipeline-night in the last several weeks, record the first thing that went wrong. Rule: each late run goes in exactly one bucket, so the buckets cannot overlap.
Subproblems to split off. The split below is MECE (mutually exclusive: no run in two buckets; collectively exhaustive: every late run lands somewhere), because the first-cause rule decides placement in time order.
| Subproblem | What it means | Typical levers |
|---|---|---|
| Started late | Upstream data or a dependency was not ready at start time | Agreed delivery times with upstream owners, sensors (checks that wait for the upstream data to appear and then start the job, instead of running at a fixed clock time), fallback to last good data |
| Failed | The run errored and recovery took too long | Automatic retries, safe reruns (idempotent: running twice gives the same result), alert routing to on-call (the engineer currently responsible for responding to pages) |
| Ran too long | Started on time but overran the window | Profile the slowest stage (measure where the time goes), fix skew (one worker getting far more data than the others, so everything waits for it), resize resources (give the job more compute), reorder jobs on the critical path (the longest chain of jobs that depend on each other, which sets the finish time) |
Worked example
All numbers are illustrative. Take 20 pipelines over 30 nights: 600 pipeline-nights. At 85% on time, 510 are on time and 90 are late. At 99%, at most 6 may be late, so about 84 of the 90 late runs (roughly 93%) must stop happening.
Cause log: started late 45 (upstream late 36, dependency chain with no slack 9 (job B waits for job A with no spare time between them, so any delay in A passes straight through)), failed 27, ran too long 18. The total is 90. So half the problem (45 of 90) is upstream timing and 30 percent (27 of 90) is failures. That ordering decides the plan: fixing run time first would attack only 20% of the lateness.
One run traced through the rule: a pipeline due at 06:00 gets its upstream data at 05:30 instead of 04:00, so it starts late, and it then also fails on a timeout. It is logged once, under started late, because that was the first thing that went wrong; the failure goes in a notes column. Counting it in both buckets would make the buckets sum to more than the 90 late runs.
Monthly targets: 91% means at most 54 late runs, 96% at most 24, 99% at most 6.
The one page for leadership (excerpt):
- Goal: tier-1 pipelines on time 99% of nights by month 3, measured as data visible by 06:00. Baseline today: 85% across all 20 pipelines (the tier-1 baseline comes from cutting the same cause log by tier, and the 99% applies to that tier-1 slice).
- Diagnosis: 90 late runs in 30 nights: 50% started late, 30% failed, 20% ran long.
- Workstreams: (1) upstream delivery agreements and fallback data, (2) failure recovery and alerting, (3) runtime on the critical path.
- Measure: weekly on-time rate per tier, plus late runs by cause.
- Milestones: 91% month 1, 96% month 2, 99% month 3.
- Asks and risks: a named owner on each upstream team; if upstream will not commit, we deliver stale-marked data on time instead of fresh data late.
The same one-page shape for a latency goal such as cutting data latency by 50% in six months. Define latency as source event to queryable data, measured at the 95th percentile (sort the runs from fastest to slowest; 95 of every 100 finish at or under this value, so it shows the slow tail rather than the typical run). Break it by stage and give each a budget. Illustrative: ingest 15, transform 30, load 10, refresh 5 minutes (60 total) becomes 8, 14, 5, 3 (30 total). Caveat: the 95th percentile of the total is not the sum of the stage 95th percentiles. With independent, light-tailed stages the sum usually overstates it, because the slowest run in one stage is rarely the slowest in every stage; with heavy-tailed stages (a few very long runs) it can understate it, so neither direction is guaranteed. Treat the stage budgets as targets and measure the end-to-end 95th percentile directly. The page then lists goal, one component per stage, a measure per stage and checkpoints roughly every month.
Trade-offs and pitfalls
- A single average hides the tail: report on-time rate per tier, not one blended number.
- Do not promise 99% before the cause log exists. After week 1, state the plan with real data.
- Upstream lateness cannot be fixed by your team alone, so the plan must contain a fallback you control.
- Fixing the loudest incident instead of the largest bucket is the common wrong turn.
- If two buckets turn out to share one cause (for example a shared cluster), say so on the page and treat it as one workstream.
Your team's delivery velocity has dropped over two quarters. Leadership suspects technical debt, others blame unclear requirements, others blame process. How would you structure the question so the competing explanations can actually be separated?
Sample Answer
Direct answer
Make the three explanations predict different, measurable fingerprints before looking at any data. First fix what "velocity" means (items finished per period, or cycle time per item), then split the problem into three non-overlapping parts: how much capacity there is, what the capacity is spent on, and how efficiently an item flows through stages. Technical debt (shortcuts in the code that make later changes slower), unclear requirements and process each leave a different mark on the flow stages.
Structured elaboration
Step 1: define and check the measurement. Story points (the relative size units a team gives each item) can drift (teams re-estimate) so I would use count of finished items and cycle time, the days from "started" to "done". Confirm the metric did not change definition.
Step 2: a MECE top split (mutually exclusive, collectively exhaustive).
- Capacity: headcount, attrition (people leaving), new joiners ramping (still learning, so not yet at full speed), time lost to meetings or on-call (the rotating duty to respond to production incidents).
- Work mix: share of effort on planned features vs bugs, incidents and unplanned work.
- Flow efficiency: how long an item takes per stage: clarify (requirements ready), build, review, test and release.
Step 3: map each suspect to a fingerprint.
| Explanation | Where it would show | Specific evidence |
|---|---|---|
| Technical debt | Build stage and work mix | build time up on a few high-churn modules (parts of the codebase that are edited very often, unrelated to customer churn); rising share of bugs and rework; engineers' notes naming those modules |
| Unclear requirements | Clarify stage and rework | time before work starts; tickets reopened or changed after start; scope changes (the definition of the item growing or shifting) mid-item |
| Process | Waiting time between stages | review queue time, approval steps, handoffs, meeting load |
| Capacity (a fourth rival the tree adds) | Capacity and work mix | attrition, ramping hires, more on-call |
The tree has three top-level branches. The three camps map onto them (debt into work mix and build time, requirements into the clarify stage, process into waiting between stages), and capacity is the branch none of the camps named, which is why the table has a fourth row.
Step 4: write predictions first, then sample. Take 40 completed items from each quarter, record time in each stage and flags (reopened, scope changed, touched module).
Worked example
Illustrative sample: median cycle time rose from 6 days to 10 days. By stage, days went: clarify 1 to 2, build 3 to 4, review 1 to 3, release 1 to 1. Check: 1 + 3 + 1 + 1 = 6 before and 2 + 4 + 3 + 1 = 10 after. Of the 4 extra days, 2 are review waiting (process), 1 is clarify (requirements) and 1 is build (possible debt). The finding is that no single camp is right: process explains about half and the other two roughly a quarter each. The next step is to check whether the build increase is concentrated in a few modules, which would separate debt from general slowdown.
Why a stage maps to a cause: review time is mostly an item waiting for a person, which is process. Clarify time is an item not being ready to start, which is requirements. Build time is the work itself, where debt would show, but only if it concentrates in particular modules, so the build day stays "possible debt" until that check. The other two top-level branches are measured differently. Capacity (illustrative): with 8 engineers, if 2 are new and work at about half speed while ramping, effective capacity is 6 + 2 x 0.5 = 7, which is 7 / 8 = 87.5% of full strength, 12.5% less. Work mix: count finished items by type. If unplanned work (bugs, incidents) grew from 15% to 30% of finished items, only 70% of the team's output is planned features instead of 85%.
Caveats on this illustrative sample. The stage list in Step 2 names five stages (clarify, build, review, test, release); the example above folds test time into review and release so that four stages sum to the cycle time. Also, stage medians do not in general add up to the median of total cycle time, so a real readout would compute the stage times per item and average or take medians of those. Finally, with 40 items per quarter a one-day shift in a stage can be inside normal spread, so the "about half" and "roughly a quarter each" split is a lead that sets the next check, not a verdict on the three camps; look at the spread (for example the 25th to 75th percentile range) for each stage before naming a winner.
Trade-offs and pitfalls
- Do not argue the three explanations in the abstract. Each camp can tell a plausible story, so the data must be framed to discriminate.
- Capacity and work mix are easy to forget and often explain a lot (a new hire ramp, more incident work).
- Sample, do not average everything. A hand-read sample of 40 items per quarter reveals reasons that dashboards hide.
- Avoid blame framing in the readout. Present stages, not teams.
That is every published Structured Problem Solving and Decomposition question for Engineering Manager so far. Browse the other topics in this category, or practice this one interactively.