Data-Driven Business Decision-Making Questions
Moving from data to a defensible business recommendation and being transparent about the evidence behind it. Covers weighing conflicting or weak evidence sources, including third-party reports and vendor claims, reconciling figures that disagree and checking whether a headline movement can be trusted, documenting assumptions and sensitivity, sizing an impact from limited inputs with stated assumptions, quantifying and communicating uncertainty in business terms, choosing between options when the data is incomplete or a short-term cost trades against a longer-term gain, deciding how much precision a decision needs, making a recommendation reproducible and auditable, defending it to skeptical stakeholders, and revisiting it when later results contradict it. Tests whether a candidate can reach and justify a recommendation from evidence rather than intuition. Building a full business case or financial model, experiment and causal-inference methods, metric definition, query writing, dashboard building, and presentation craft are covered elsewhere.
You have researched three vendor or partner options and need to hand leadership a one-page recommendation they can act on without you in the room. What goes on that page, how do you show the evidence and its weak spots, and what would change your call?
Sample Answer
What goes on the page
A reader with five minutes and no access to you needs seven things:
- The decision being asked and by when.
- The recommendation in one sentence, with conditions.
- The options compared on the same criteria.
- The evidence behind each score and how strong it is.
- The weak spots and risks.
- What would change the call.
- Next steps with an owner for each.
Excerpt (illustrative: three vendors for a customer-messaging platform)
Recommendation: choose Vendor B, conditional on its security review passing before contract signature.
Scores run from 1 (poor) to 5 (strong). Weights sum to 100. Total = sum of weight x score, divided by 5, so the maximum is 100. Cost is scored from the 3-year total: lowest cost scores 5, up to 20% above lowest scores 4, 20 to 40% above scores 3, over 40% above scores 2. For example, C ($360k) is the lowest and scores 5; B ($430k) is 19% above C, so it scores 4; A ($540k) is 50% above C, so it scores 2. Viability means whether the vendor is likely to still be around, funded and supporting the product in three years.
| Criterion (weight) | A | B | C |
|---|---|---|---|
| Fit to requirements (30) | 5 | 4 | 3 |
| 3-year cost (25): A $540k, B $430k, C $360k | 2 | 4 | 5 |
| Integration effort and risk (20) | 3 | 4 | 3 |
| Security and compliance (15) | 4 | 4 | 3 |
| Viability and support (10) | 4 | 3 | 2 |
| Weighted total | 72 | 78 | 68 |
Showing the evidence and its weak spots
Label each score by evidence strength: tested (we ran it), corroborated (reference calls or documents), or claimed (vendor statement only). For B: fit is tested in a trial; cost is a written quote; integration is corroborated by one reference call and a demo; security rests on the vendor's own security questionnaire (a standard form of questions about how it protects data), which is claimed. Saying this on the page is what lets leadership trust the rest.
What would change my call
- B's integration score falls from 4 to 2: B drops to 70 and A (72) wins. A drop to 3 leaves B ahead, 74 to 72.
- Moving 15 weight points from fit to cost leaves B on top, 78 to C's 74.
- B failing the security review ends the recommendation outright, whatever the score.
Next steps: the security lead completes the review; procurement (the team that negotiates and signs vendor contracts) negotiates B's price; the engineering manager runs a 2-week integration spike (a short, time-boxed trial that connects B to our real systems to find out how hard the integration is).
Leadership will act on your analysis next quarter, and an auditor or a new analyst may later ask how you reached it. What would you put in place so your recommendation can be reproduced and audited, and how would that differ for a quick pilot versus a decision with large financial or regulatory exposure?
Sample Answer
Direct answer
A recommendation is reproducible if a different person, given the same inputs and your materials, gets the same numbers. It is auditable if a reviewer can trace each figure back to its source and see who made each judgement and why. Build both in from the start, but scale the effort to what is at stake: a light version for a pilot, a controlled version for a decision with large financial or regulatory exposure.
What to put in place
- The question and decision. What decision this supports, who owns it, and the date.
- Data lineage: a record of where each input came from and when it was extracted, plus a frozen snapshot (a read-only copy of the data as it stood on that date), so a later refresh of the source cannot change your result.
- Code or steps that run end to end from raw inputs to the final table, kept under version control (a tool such as Git that saves every version of a file and who changed it), not clicked together by hand.
- An assumptions log: each assumption, its source, who approved it.
- Pinned environment: a written list of the exact tool and library versions used (for example, the Python version and the packages), so a rerun months later behaves the same even after the tools update.
- Independent check: a second person reruns it.
- Change log and sign-off: what changed between versions and who approved.
The test: hand it to a new analyst with no help. If they reproduce your headline figures, it is reproducible.
How the bar differs
| Element | Quick pilot | Large financial or regulatory exposure |
|---|---|---|
| Data | Note the source and extract date | Frozen snapshot stored with access controls and a retention period (how long the record must be kept, set by company policy or regulation) |
| Code | One saved script or notebook | Version-controlled, peer-reviewed, reruns from raw data |
| Assumptions | Short list in the write-up | Formal log with owner and approval |
| Review | Self-check, optional peer glance | Independent reproduction and sign-off |
| Record | A page of notes | Decision record kept for the audit period |
Worked example (illustrative): a two-week pricing pilot is documented in a one-page note with the query, extract date, three assumptions, and the result, and is rerun by a colleague only if it scales up. A pricing change touching all customers and revenue reporting gets a frozen dataset, a reviewed analysis, a log of ten assumptions with owners, an independent rerun that matched to the cent, and approval recorded before launch. One assumptions-log entry might read: "A4: 18% of customers accept the new price. Source: the two-week pilot. Owner: pricing lead. Approved by Finance on the sign-off date." The rerun by a second analyst from the frozen snapshot reproduced the headline figure of $412,350 exactly, and that match was recorded in the decision record.
Trade-off and judgement: heavy process on a pilot slows learning, while light process on a high-exposure decision leaves you unable to defend it. Start light, and add rigour as soon as the decision becomes harder to reverse.
Market research says one thing, partner feedback says another, and sales data suggests a third conclusion. How would you resolve the conflict, decide what to believe, and communicate the uncertainty without sounding indecisive?
Sample Answer
Direct answer
I do not pick a winner among the three sources. I first find out why they disagree, because the disagreement is usually information: each source measures a different population, time or kind of behavior. Once one explanation fits all three, I decide on that basis, state my confidence in plain words, name the evidence that would change my mind, and commit to a decision date. Being decisive about the next step is different from claiming certainty about the facts.
Structured elaboration
1. Make a disagreement table. For each source write down who was measured, what was measured (stated intent or actual behavior), when, and who benefits from the answer.
| Source | Measures | Typical bias |
|---|---|---|
| Market research (survey) | What people say they would do | Stated intent overstates action; sample may not match buyers |
| Partner feedback | What partners hear and want | Partners are intermediaries with their own margin and incentives; they speak for the customers they serve, which may be only one segment |
| Sales data | What customers did buy | Reflects the current offer, price and who sales chose to call |
2. Calibrate against history. Calibrating means scaling stated intent by how past stated intent turned out. If a past launch exists, compare what the survey predicted with what happened. That ratio is better than a generic discount.
3. Reconcile by segment and time. Often the sources agree inside one segment and diverge across segments.
4. Decide and communicate with calibrated language: "We think X (about 70% confident). The main risk is Y. We will know more by date Z, and if A turns out, we will switch."
Worked example
Illustrative: should we add a premium tier? "Top-box" means the share of respondents who chose the strongest answer on the scale ("definitely would buy"). To calibrate means to scale stated intent by how past stated intent turned out.
- Survey of 1,000 people: 60% "interested", 14% top-box (140 people).
- Partners: "customers do not ask for it."
- Sales: a comparable add-on (a different product from last year's add-on used for calibration below) sold in 450 of 5,000 deals, 9%.
Calibration: for last year's add-on the survey top-box was 20% and actual sales were 11%, so reality was 11 / 20 = 55% of the stated figure. Apply that to the new survey: 14% x 0.55 = 7.7%. That is close to the 9% in the sales data, so the survey and sales agree once the survey is discounted. This is a cross-check against a proxy product, not an independent measurement of the premium tier, and it applies one ratio (0.55) from a single past launch to every segment, although overstatement may differ between large and small accounts.
Now reconcile by segment. Splitting the same data:
| Segment | Survey respondents | Top-box | Calibrated (x 0.55) | Sales on similar add-on |
|---|---|---|---|---|
| Large accounts | 400 | 25% (100) | 13.75% | 280 of 2,000 deals = 14% |
| Small accounts | 600 | 6.7% (40) | 3.7% | 170 of 3,000 deals = 5.7% |
| All | 1,000 | 14% (140) | 7.7% | 450 of 5,000 = 9% |
In large accounts the calibrated survey (13.75%) and sales (14%) agree closely, so both point to real demand there. In small accounts both are low. The partners we asked sell mainly to small accounts, so "customers do not ask for it" is consistent with the data: they are describing the segment where demand is weakest, not contradicting the large-account signal.
Decision: launch the premium tier to large accounts first, not to everyone. Planning figure for large accounts: 12 to 14% uptake (the calibrated survey and sales numbers). I would say: "I am fairly confident (about two in three) that uptake in large accounts lands between 10% and 15%. That range is wider than the 13.75% and 14% estimates because the calibration rests on a single past launch. If the first 90 days show under 10%, we pause. Review on a set date." For a launch to everyone the same logic gives 7.7% to 9%, which is why I would not plan around a blended figure.
Trade-offs and pitfalls
- Averaging the three numbers hides the cause of the disagreement; the table finds it.
- Do not dismiss the partners because they disagree; check whether they see a segment you do not.
- To avoid sounding indecisive, give the decision first, the confidence second and the trigger that would change it third.
You are evaluating a new business segment with very little historical data, no direct benchmark, and inconsistent customer feedback. How would you make a recommendation anyway, and what guardrails or pilot design would you put in place to manage the uncertainty?
Sample Answer
Direct answer
With little data I still recommend, but I recommend a bounded, reversible step instead of a full commitment. I make the assumptions explicit, size the pilot so that its result can actually change my mind, and set the spending cap and the stop rule before it starts. Thin data changes the shape of the decision (small, staged, checkable), not whether I give a view.
Structured elaboration
Act now or validate first? Four tests.
- Reversibility: can we undo it cheaply? A pilot yes, a multi-year contract no. Reversible decisions can be made on thin evidence.
- Cost of waiting: is a window closing (competitor, season)? If delay is cheap, gather more data first.
- Value of information: would more data plausibly flip the decision? If every outcome of a study leads to the same action, skip the study. Example: if the segment is worth entering whether the study shows 5% or 15% conversion, the study has no value.
- Cost of the test versus the cost of being wrong: do not spend $200k to avoid a $50k mistake. Applied here: a capped $150k pilot is reversible and the full launch is not, so the pilot passes tests 1 and 4.
State the assumptions the decision rests on. An assumption is something we believe but have not shown. Typically: customers in the new segment have the problem, will pay roughly what we charge, and can be reached at an acceptable cost. I rank them by how much the decision depends on each, and test the riskiest first.
Communicate the limits of the sample. Inconsistent feedback from 9 customers, 6 positive, is not "two thirds are positive". The 95% interval (the range of true shares that fit the data) for 6 of 9 runs from roughly 35% to 88%, which includes "barely half" and "nearly everyone". I say that plainly, and also say who the 9 were (who picked them, who is missing).
Worked example
Illustrative: a new small-business segment for a product sold to enterprises. No history, no benchmark, 9 mixed interviews.
Pilot design:
- Question: can we win qualified small-business leads at a rate that makes the segment worthwhile?
- Scope: 60 qualified leads over 8 weeks, with a separate small team and a hard spend cap of $150k.
- Success bar (set now): at least 6 of 60 convert (10%) and the sales cycle (time from first contact to signed deal) under 45 days. A qualified lead is a prospect who fits the target customer profile and has a real need and budget.
- Stop rule: if fewer than 3 convert by week 6, stop and write up why.
- Owner and check date: the segment lead, with a review at week 4.
Is 60 leads enough? This is a power question: power is the chance the pilot reaches the bar when the segment really is good. The binomial calculation (the standard way to count how often a given number of successes appears in 60 independent tries at a fixed rate) gives the chance of seeing 6 or more conversions: if the true rate were 3%, under 1%; at 5%, about 8%; at 10%, about 56%; at 15%, about 90%. Read it this way: if the segment is really poor (3%), a pass by luck is very unlikely, so a pass is meaningful evidence the rate is not tiny. If the segment is really at 10%, we still fall short of 6 about 44% of the time (100% minus 56%), so a miss is not proof of failure. I therefore treat a result of 3 to 5 conversions as "extend the pilot" and not as a verdict.
Check the week-6 stop rule the same way, because by week 6 only about three quarters of the leads (roughly 45 of 60) have been worked. At a true 10% rate, the chance of fewer than 3 conversions among 45 leads is about 16%, so the stop rule would wrongly end a segment that is exactly on the bar about one time in six. At a true 5% it fires about 61% of the time and at 15% only about 3%. That is acceptable for a spending guard on a $150k cap, but I say so: week 6 is a cheap-stop check for clearly poor results, and the launch decision waits for the full 60 leads at week 8.
Recommendation to leadership: approve the capped pilot, not the segment launch. Decision on launch at week 8 using the bar above.
Trade-offs and pitfalls
- The recommendation changes if a cheap, fast study (for example 10 more interviews with the right buyers) is available and could flip the decision; then validate first.
- Pitfall: a pilot too small to detect the effect you care about. Compute its power (as above) before running it.
- Pitfall: moving the success bar after seeing results.
A skeptical VP doubts that a proposed product change is worth piloting. Write the one-page plan you would put in front of them: what you would test, the range of impact you expect, and exactly which results would make you proceed, change course or stop.
Sample Answer
One-page pilot plan: guided setup checklist for new accounts
(Illustrative example of a product change; every number is a stated assumption.)
The ask: a 6-week pilot, with the go / change course / stop rule below agreed before we start.
What we test: half of new accounts get the checklist (treatment), half do not (control), assigned at random. One change only.
Metric: 14-day activation (an account completes its first key action within 14 days of signup). Baseline: 40%.
Why I expect it to work: new accounts today stall at setup; the checklist removes the step where most of them drop off. This is a hypothesis, which is why it is a pilot.
Range of impact I expect: from 0 to +4 points, most likely about +2 (a judgment from comparable onboarding changes; I will say if the evidence is thin).
Size and duration: to detect a 2-point difference (40% vs 42%) with 95% confidence and 80% power, a standard two-proportion calculation needs about 9,500 accounts per group, roughly 19,000 in total. Confidence of 95% means that if the checklist did nothing, we would be fooled into seeing a difference (in either direction) only about 1 time in 20; a false lift specifically, which is what the Proceed row requires, about 1 time in 40. Power of 80% means that if the true lift is 2 points, we would catch it 8 times in 10. The calculation:
n per group = (1.96 x sqrt(2 x 0.41 x 0.59) + 0.84 x sqrt(0.40 x 0.60 + 0.42 x 0.58))^2 / 0.02^2
= about 9,500
1.96 and 0.84 are the standard values for 95% confidence and 80% power; 0.41 is the average of the two rates. A smaller true effect would need far more accounts, and a baseline of 25% instead of 40% needs fewer (about 7,500 per group). At 5,000 new accounts a week that is under 4 weeks of enrolment; I run 4 full weeks (20,000 accounts, 10,000 per group) plus 2 weeks for the 14-day window to complete: 6 weeks.
Decision rule (set now, not after seeing results)
(The 95% interval is the range of true lifts consistent with the data; "entirely above zero" means even its low end is a real improvement. A guardrail is a metric that must not get worse, defined under Guardrails below.)
| Result | Decision |
|---|---|
| Lift of +2.0 points or more, 95% interval entirely above zero, no guardrail breached | Proceed: roll out to everyone |
| Lift from +1.0 up to +2.0 points | Change course: find where treated accounts still drop off, iterate once, retest |
| Lift under +1.0 point, or any guardrail breached | Stop: not worth building and maintaining |
Why +1.0 as the floor: it is about 50 extra activated accounts a week at 5,000 signups, roughly the size below which maintenance cost outweighs the gain (assumption).
Guardrails (things that must not get worse): support tickets per 1,000 new accounts up by more than 10% relative; 30-day retention of treated accounts lower than control by more than 1 point, with the 95% interval for the difference entirely below zero. A bare "below control" is not a breach: with about 10,000 accounts per group, chance alone puts the treated group below control about half the time even if the checklist does nothing, so that rule would stop the pilot on a coin flip. "Any guardrail breached" in the table means these thresholds.
Timing of the retention guardrail: 30-day retention needs 30 days after signup, so it is not complete at week 6 for accounts enrolled late. I read it on the first two weeks of enrolment (about 10,000 accounts) at about week 7 and on everyone at about week 9, and I do not release the full rollout until the first read is in. The activation read stays at week 6.
What would change the plan: if weekly signups fall well below 5,000, the pilot runs longer rather than being declared early.
Addressing the skeptic directly: the plan costs 6 weeks of one team's time, can be switched off at any point, and ends in one of three named outcomes, none of them "more analysis".
Unlock Full Question Bank
Get access to all 23 Data-Driven Business Decision-Making interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.