Business Acumen and Commercial Context Questions
Reasoning about how a business earns and spends money, and judging the commercial consequence of a decision from the candidate's own seat, at a generic (company-agnostic) level. Covers translating technical, product or data work into revenue, cost, margin and risk terms: model accuracy, precision and recall choices and decision thresholds priced in dollars; latency, availability and incident downtime turned into revenue at stake; build-versus-buy, vendor, architecture, multi-region, single-tenant and technical-debt decisions weighed commercially. Also covers trade-offs such as growth versus profitability, speed versus cost, engagement versus revenue, and opportunity cost and the cost of delay between competing investments; how an engineering, data or platform function creates and demonstrates business value, including indirect contributions; connecting functional plans and OKRs to company strategy and revenue goals; and explaining the commercial case to non-technical leaders and finance partners, with stated assumptions and uncertainty. Tests whether a candidate can tie day-to-day choices to commercial outcomes and explain that link clearly. Full investment business cases, KPI design, computing unit economics, pricing decisions, case-interview frameworks and researching a specific company are covered elsewhere.
How would you connect a service's reliability to commercial outcomes? Walk through how a longer recovery time after an outage becomes an estimated revenue loss.
Sample Answer
Direct answer
Reliability connects to revenue through a short chain: how often the service fails, how long customers are affected (the recovery time), how much revenue flows through the affected part during those minutes, how much of that comes back later, and what extra costs the outage triggers (service-level credits, support load). A longer recovery time multiplies the minutes of impact, so the lost-revenue part of the loss grows in proportion to it; the fixed per-incident costs do not grow, so in the example below tripling recovery time from 30 to 90 minutes raises the total cost per incident about 1.65 times ($17,753 to $29,260), not 3 times.
The chain, step by step
- Frequency. Number of customer-visible incidents per year (6 in the example).
- Duration. MTTR, mean time to recovery: the average time from the start of customer impact to restoration, so it includes the time taken to detect the problem as well as the time to fix it. In the example it moves from 30 to 90 minutes.
- Revenue exposure per minute. Average revenue per minute (annual online revenue / minutes in the year), adjusted for when the outage happens (a peak factor of 1.5, an assumption that your incidents tend to land in busier-than-average hours; if outages struck at uniformly random times the factor would be 1.0, so take it from the hour of day of your own past incidents) and for the share of revenue the outage actually blocks (80%, for example because 20% of orders come through a path the outage does not touch, such as the mobile app on a separate service).
- Recoverable share. Some customers retry later: here 30% of blocked purchases return.
- Other costs. SLA credits (refunds promised in a service-level agreement), support tickets and engineer time: $12,000 per incident, illustrative and treated as fixed per incident here (in practice credits often grow with duration).
- Total. Loss per minute x minutes + other costs, then x incidents per year.
Worked example
Annual online revenue is $120 million.
ANNUAL_ONLINE_REVENUE = 120_000_000
HOURS_PER_YEAR = 8760
PEAK_FACTOR = 1.5 # an outage hits busier-than-average hours
SHARE_AFFECTED = 0.8 # share of revenue flows the outage blocks
RETURN_RATE = 0.3 # share of blocked purchases that happen later anyway
SLA_CREDITS_AND_SUPPORT = 12_000 # dollars per incident, illustrative
avg_per_minute = ANNUAL_ONLINE_REVENUE / HOURS_PER_YEAR / 60
loss_per_minute = avg_per_minute * PEAK_FACTOR * SHARE_AFFECTED * (1 - RETURN_RATE)
print(f"average revenue per minute ${avg_per_minute:,.2f}; lost revenue per outage minute ${loss_per_minute:,.2f}")
for mttr in (30, 90):
per_incident = loss_per_minute * mttr + SLA_CREDITS_AND_SUPPORT
print(f"MTTR {mttr:>2} min: ${per_incident:>9,.0f} per incident, 6 incidents a year ${6 * per_incident:>9,.0f}")
Output:
average revenue per minute $228.31; lost revenue per outage minute $191.78
MTTR 30 min: $ 17,753 per incident, 6 incidents a year $ 106,521
MTTR 90 min: $ 29,260 per incident, 6 incidents a year $ 175,562
If the factor were 1.0 (incidents at random hours) lost revenue per outage minute would be $127.85 and an incident at 30 minutes would cost $15,836, so the peak-factor assumption moves the per-incident cost by about 12%. Lost revenue per outage minute at the 1.5 factor is $191.78, so extending recovery from 30 to 90 minutes adds 60 x $191.78 = $11,507 per incident (from $17,753 to $29,260) and about $69,000 a year over six incidents (from $106,521 to $175,562).
Where the numbers come from, and the caveats
- Revenue per minute should come from payment logs for the same weekday and hour of past incidents, not from a yearly average.
- The return rate comes from comparing order volume in the hours after past outages with a normal day. For example, if 1,000 orders were blocked during an outage and the following three hours ran 300 orders above a normal day, the return rate is 300 / 1,000 = 30%.
- The share blocked comes from payment logs too: the fraction of normal orders in the affected hours that went through the failed component.
- Revenue that is not minute-by-minute (annual subscriptions) is not lost during an outage; the damage shows up as churn, credits and trust, which need separate estimates.
- The model leaves out longer-term effects (customers who leave), so it is a floor for the damage of a single long outage.
Why this matters for decisions
It turns an engineering metric into a price: "each minute we shave off recovery is worth about $190 per incident", which you can set against the cost of better alerting or runbooks (written step-by-step recovery procedures).
Leadership asks you to estimate what an hour of downtime costs the business, quickly enough that responders can use it to prioritise during an incident. What inputs would you use, how would you handle the uncertainty, and how would you present the number?
Sample Answer
Direct answer
Build the estimate from four components: revenue that is actually lost (not just delayed), the profit on it, contractual credits owed, and the fixed cost of the response. Compute it ahead of time as a small table by time-of-day band so a responder can look it up in seconds, show it as a range with a midpoint, and label it as an estimate. In the illustration below one hour costs about $12,200 in the evening peak, $7,700 in the day and $4,100 overnight, with a peak range of $9,650 to $14,810.
Inputs
Definitions: revenue at risk is what the service would have earned that hour. Unrecovered share is the fraction of that revenue that never comes back (customers who buy later, retry or wait have only delayed their purchase). Contribution margin is revenue minus variable costs, so lost contribution is the real profit lost. SLA credits (service-level agreement credits) are refunds promised to customers when availability drops under the contract.
Illustrative inputs for an online store (replace with your own): AOV $80; orders per hour 600 at weekday-evening peak, 300 in the day, 60 overnight; contribution margin 25%; 12 engineers at $100 per hour loaded (salary plus benefits and overhead, divided by hours worked); 0.25 support tickets per lost order at $6 each; $2,000 of contractual credits per outage hour.
The calculation (peak hour, midpoint 70% unrecovered)
| Component | Calculation | Result |
|---|---|---|
| Revenue at risk | 600 orders x $80 | $48,000 |
| Lost revenue (70% unrecovered) | $48,000 x 70% | $33,600 |
| Lost contribution | $33,600 x 25% | $8,400 |
| Responder time | 12 x $100 | $1,200 |
| Extra support | 600 x 70% x 0.25 x $6 | $630 |
| SLA credits | assumed | $2,000 |
| Total cost of the hour | $12,230 |
The lookup table for responders
| Band | Orders per hour | Revenue at risk | Cost at 50% / 70% / 90% unrecovered |
|---|---|---|---|
| Evening peak | 600 | $48,000 | $9,650 / $12,230 / $14,810 |
| Daytime | 300 | $24,000 | $6,425 / $7,715 / $9,005 |
| Overnight | 60 | $4,800 | $3,845 / $4,103 / $4,361 |
Handling the uncertainty
- Give a range, and move only the inputs that are uncertain. Here the unrecovered share is the main unknown; measure it from past outages by comparing the hours after recovery with a normal day. For example, if an outage removed 600 orders and the three hours after recovery ran 180 orders above the normal level for that time of day, 30% of the demand came back and 70% did not, which is where the 70% midpoint comes from.
- Responder time is mostly sunk (the engineers are paid anyway). Keep it as a line so the number is complete, but do not let it hide the revenue effect.
- Do not add long-term brand damage or churn into the headline number. If you believe it matters, show it as a separate, labelled line with a stated assumption.
How to present it
During an incident: one line, "about $12k per hour right now, plausible range $10k to $15k", plus the band used. It lets responders prioritise (a payment outage in the peak band beats a reporting bug) without arguing about precision. After the incident: show the components and the assumptions, and re-estimate.
Pitfalls
- Using annual revenue divided by 8,760 hours gives a flat $4,566 per hour for a $40M business, which understates the peak and overstates overnight.
- Reporting lost revenue and lost profit as if they were the same number. Choose one for the headline and label it.
- Partial outages: scale by the fraction of users or functions affected. If 40% of users are affected in the evening peak, the revenue-driven lines scale (lost contribution $8,400 x 40% = $3,360; extra support $630 x 40% = $252), while responder time stays $1,200 and the $2,000 of credits is assumed to still apply, giving $1,200 + $2,000 + $3,360 + $252 = $6,812 for the hour.
You need to tell a non-technical CEO that a promised feature will slip because the team must first fix a reliability regression. How would you frame the message in commercial terms?
Sample Answer
Direct answer
Lead with the customer and money, not the bug. Say: "We found a fault that is making some customers fail to pay us. Fixing it first takes about two weeks and moves the new feature by two weeks. Every week we leave it unfixed puts up to $24,000 of orders at risk, against about $15,000 a week of expected feature revenue, and the feature is only delayed, not lost. If many failed customers retry, the true loss is smaller, so I will measure the retry rate and confirm." Then give the new date, the cost of each option, and the one decision you need from them.
How to build the message
- Translate the regression into a business symptom. A "reliability regression" means something that used to work got worse: slower pages, more errors, more downtime. Name the symptom the CEO already cares about: failed checkouts, support tickets, churn risk, a contract with an availability promise.
- Put a number on the cost of leaving it. Use rates and money, not technical measures.
- Put a number on the cost of the delay. Be honest that a delay is revenue deferred, not revenue destroyed, unless a launch window or a customer commitment makes it truly lost.
- Offer a choice, with a recommendation. Two or three options, each with date, cost and risk, and your pick.
- Say what you will report and when, so trust does not depend on the CEO understanding the fix.
Worked example (illustrative numbers)
A shop has 20,000 checkout attempts a week at an average order of $80. After last release, 1.5% of attempts fail with an error.
| Item | Arithmetic | Result |
|---|---|---|
| Failed checkouts per week | 20,000 x 1.5% | 300 |
| Order value at risk per week | 300 x $80 | $24,000 |
| Fix duration | engineering estimate | 2 weeks |
| Feature revenue deferred by the fix | 2 weeks x $15,000 expected per week | $30,000 (pushed later, not lost) |
| Revenue at risk if we ship the feature first and fix later | $24,000 per week, for as long as the fault lives | Open ended |
Some failed customers retry, so $24,000 is an upper bound; say so, and say you will measure the real retry rate. The comparison with the $15,000 a week the feature earns holds only if fewer than about 37.5% of failed checkouts are retried and completed (1 - 15,000 / 24,000 = 37.5%); above that, the fault costs less per week than the feature earns, and the case for fixing first rests on the risk growing and on the feature running on the same checkout. The script in the room:
"A change we shipped last month is making about 1 in 67 checkouts fail. That is roughly $24,000 of orders at risk every week. If we fix it first, the new feature lands two weeks later and about $30,000 of its revenue moves later in the year. If we ship the feature first, the checkout problem stays live and the risk keeps growing, and the feature itself runs on the same checkout. I recommend fixing first. I will send you the failed-checkout rate every Friday until it is back to normal."
(The 1 in 67 is 1 divided by 0.015.)
Trade-offs and pitfalls
- Jargon. Do not say "latency regression", "error budget" or "tech debt" to a CEO. Say "customers waiting", "failures", "money at risk".
- Do not hide the delay inside a vague date. Give a date with a confidence level ("two weeks, and I will confirm by Wednesday once we have traced the cause").
- Do not oversell the fix. If the cause is not yet known, say the range and when you will know.
- What changes the recommendation: if the feature is tied to a dated launch, a contract or a partner event that would be lost by a two-week slip, or if the regression only hits a small, low-value slice, the comparison changes. Say that you checked.
Your service has many small, short incidents every month and none is severe. How would you measure their cumulative business cost, and how would you decide between continuing to patch and investing in a deeper fix?
Sample Answer
Direct answer
Add up the small incidents into one monthly cost (customer-facing downtime plus responder time), classify them by root cause, and then price the fix for each large cause separately. In the illustration below, 30 short incidents cost about $28,500 a month; a broad $96,000 rewrite pays back in about 4.8 months, but fixing only the single biggest cause costs $12,000 and pays back in about 1.4 months. Fix the causes with the shortest payback first (here the second-largest cause pays back slightly faster than the largest), and keep patching the rest.
Step 1: measure the cumulative cost
Illustrative inputs: 30 incidents a month, each affecting customers for 12 minutes on average; each affected hour costs $4,000 in lost profit; each incident takes 2 responders 45 minutes at $100 per hour loaded.
| Component | Calculation | Per month |
|---|---|---|
| Customer-facing downtime | 30 x 12 minutes = 6 hours x $4,000 | $24,000 |
| Responder time | 30 x 2 x 0.75 hour x $100 | $4,500 |
| Total | $28,500 |
Two costs are missing and should be noted, not priced: interruptions and context switching for the engineers, and the erosion of trust. Do not guess them into the total.
Step 2: classify by root cause
Tag every incident with its cause. Illustrative result: cache stampede (many requests rebuilding the same cache entry at once) 9, bad configuration push 7, dependency timeout 5, and 9 spread across 6 other causes, none above 2 incidents each. These sum to 30 (9 + 7 + 5 + 9). The top three causes account for 21 of 30, or 70%; the long tail of small causes is the other 30%, and no single fix touches much of it.
Step 3: price the options
| Option | Cost | Share of incidents removed | Monthly saving | Payback |
|---|---|---|---|---|
| Broad rewrite aimed at the top three causes | 3 engineers x 8 weeks x 40 hours x $100 = $96,000 | 70% | $19,950 | $96,000 / $19,950 = 4.8 months |
| Same deep fix if it removes only 50% | $96,000 | 50% | $14,250 | 6.7 months |
| Same deep fix if it removes 90% | $96,000 | 90% | $25,650 | 3.7 months |
| Fix only the top cause (cache stampede) | 1 engineer x 3 weeks x 40 hours x $100 = $12,000 | 9 of 30 = 30% | $8,550 | 1.4 months |
| Fix the second cause (bad configuration push: validation and staged rollout) | 1 engineer x 2 weeks x 40 hours x $100 = $8,000 | 7 of 30 = 23% | $6,650 | 1.2 months |
| Fix the third cause (dependency timeout: timeouts and fallbacks) | 1 engineer x 4 weeks x 40 hours x $100 = $16,000 | 5 of 30 = 17% | $4,750 | 3.4 months |
| Keep patching | $0 up front | 0% | $0 | none, costs $28,500 a month |
The monthly saving is the share of incidents removed times $28,500, which assumes every incident costs the average $950 ($28,500 / 30); check that the causes you fix are not cheaper or dearer than average. The costs of the second and third fixes are illustrative estimates for the same hourly rate. Doing all three narrow fixes costs $12,000 + $8,000 + $16,000 = $36,000 and removes the same 70% of incidents ($19,950 a month), a payback of $36,000 / $19,950 = 1.8 months, against 4.8 months for the broad rewrite. The rewrite only earns its extra $60,000 if it also removes part of the long tail or prevents new incidents, which is the case to test before funding it.
The decision
Rank the narrow fixes by payback: the second cause (bad configuration push, 1.2 months) and the biggest cause (cache stampede, 1.4 months) come first, and they can run in parallel because each needs one engineer; the third (dependency timeout, 3.4 months) follows, while the numbers keep showing payback inside the window the business accepts. A six-month payback threshold is a common-sense choice, not a standard: set it against what else the same engineers could do. Do not start the broad rewrite until the narrow fixes show the incident count falls as predicted, because the 70% assumption is the one that decides the broad case (at 50% it becomes 6.7 months).
What would change my call
- If the incidents share one underlying architectural cause rather than many small causes, the broad fix becomes the narrow fix.
- If the incident rate is rising rather than steady, favour the deeper fix sooner.
- If the same few engineers carry all the on-call load, weigh the burn-out risk as a reason to fix earlier even when payback is slower.
Pitfalls
- Averaging the costs: ten five-minute incidents and one 50-minute incident are not the same risk to customers, so keep the distribution as well as the sum.
- Pricing the fix without a measured effect. Record how many incidents each completed fix removes and revise the next estimate.
- Counting the responders' time at the same rate as customer downtime. They are different bases, shown separately above.
Your CEO expects 99.99% availability, but the current budget supports roughly 99.95%. How would you quantify the gap in business terms, and how would you negotiate the target and the investment needed?
Sample Answer
Direct answer
99.99% availability allows 52.6 minutes of downtime a year; 99.95% allows 262.8. The gap is 210.2 minutes a year. In the example below, that gap is worth about $84,000 a year in directly lost revenue (on $200 million of online revenue, assuming outages fall in busier hours and block 70% of sales, as set out in the calculation below), against an illustrative $1.2 million a year to close it, so the revenue case alone does not support 99.99%. I would find out what the CEO's target is protecting (contracts, reputation, a competitor claim), price that, and propose a tiered commitment: 99.95% across the board now, 99.99% only for the revenue-critical path, funded when the evidence of what is at stake clears break-even (the point where the benefit equals the cost).
Quantify the gap
Availability is the share of time the service works for customers; an SLO (service-level objective) is the internal target for it, and an SLA (service-level agreement) is a contractual promise with penalties.
MIN_PER_YEAR = 365 * 24 * 60
for target in (0.9999, 0.9995):
print(f"{target:.2%} allows {MIN_PER_YEAR * (1 - target):6.1f} minutes of downtime a year")
gap_minutes = MIN_PER_YEAR * (0.9999 - 0.9995)
print(f"gap: {gap_minutes:.1f} minutes a year")
REVENUE = 200_000_000
per_minute = REVENUE / MIN_PER_YEAR
loss_per_minute = per_minute * 1.5 * 0.7 # peak factor 1.5, 70% of flows blocked
print(f"revenue lost per outage minute: ${loss_per_minute:,.0f}; direct value of closing the gap: ${gap_minutes * loss_per_minute:,.0f} a year")
INVESTMENT = 1_200_000 # illustrative yearly cost of the extra resilience
print(f"break-even needs other benefits of ${INVESTMENT - gap_minutes * loss_per_minute:,.0f} a year")
# availability of a chain multiplies
print(f"app 99.99% x database 99.95% = {0.9999 * 0.9995:.4%}")
print(f"three 99.99% dependencies in series = {0.9999 ** 3:.4%}")
# Option B: 99.99% on the checkout path only (assume checkout carries 60% of revenue at risk)
checkout_value = gap_minutes * loss_per_minute * 0.6
print(f"checkout-only direct value: ${checkout_value:,.0f} a year against an illustrative $350,000 a year")
Output:
99.99% allows 52.6 minutes of downtime a year
99.95% allows 262.8 minutes of downtime a year
gap: 210.2 minutes a year
revenue lost per outage minute: $400; direct value of closing the gap: $84,000 a year
break-even needs other benefits of $1,116,000 a year
app 99.99% x database 99.95% = 99.9400%
three 99.99% dependencies in series = 99.9700%
checkout-only direct value: $50,400 a year against an illustrative $350,000 a year
Reading the output:
- 99.99% leaves 52.6 minutes a year; 99.95% leaves 262.8, so the gap is 210.2 minutes.
- With $200 million of online revenue, a peak factor of 1.5 and 70% of flows blocked, each outage minute loses $400 and the gap is worth $84,000 a year directly.
- If closing the gap costs an illustrative $1,200,000 a year, other benefits of $1,116,000 a year would be needed to break even.
- Availability of components in series multiplies: an app at 99.99% on a database at 99.95% delivers 99.94%, and three 99.99% dependencies in series give 99.97%. A 99.99% target therefore requires every dependency in the chain to be at least that good, which is where much of the cost comes from.
Counting what the direct number omits
The $84,000 is a floor. Add, with real numbers from your business: SLA credits owed under enterprise contracts that promise 99.99%, annual recurring revenue (ARR, the yearly subscription revenue a customer contract brings in) at risk of non-renewal if the promise is missed (ARR at risk x probability of loss; for example a $2,000,000 contract with a 15% chance of leaving after a breach is $300,000 a year of expected loss), brand and sales impact, and the engineering cost of the on-call and failover work. Express it as the break-even test above: the investment is justified when these other benefits exceed $1,116,000 a year in the example.
Negotiating target and investment
In the room, it sounds like this. You: "What is the 99.99% for: a contract, a board statement, or a competitor?" CEO: "A customer asked for it." You: "Then I want to price that. Moving from 99.95% to 99.99% saves about 210 minutes of downtime a year, worth about $84,000 in direct sales, and it costs about $1.2 million a year. It only pays if there is a contract at stake. How much revenue rides on the customer who asked, and what do they do if we say 99.95%?" Then show the options below and ask for a decision on option B.
- Ask what the number is for. A contract clause, a board statement, a competitor's claim and a feeling all imply different answers.
- Show the options in minutes and dollars.
| Option | Downtime allowed | Direct revenue protected vs 99.95% | Illustrative cost |
|---|---|---|---|
| A: stay at 99.95% | 262.8 min/yr | baseline | $0 extra |
| B: 99.99% on checkout only (60% of revenue at risk) | 52.6 min/yr on that path | $50,400 a year | $350,000 a year |
| C: 99.99% everywhere | 52.6 min/yr | $84,000 a year | $1,200,000 a year |
- Recommend B-with-conditions. Commit to 99.95% across the product now. Start the checkout-path design and a firm cost estimate now (weeks of planning, not the $350,000), and release the $350,000 when the exposure test passes: B's direct value is $50,400 against $350,000, so it needs about $299,600 a year of other benefits. One enterprise contract of $2,000,000 ARR with a 15% chance of leaving after a breach is $300,000 a year, which is enough. If the real driver is a contract worth millions of ARR, B and even C pass the break-even test and I would say so.
- Explain what 99.99% demands. One 60-minute outage uses more than the whole year's budget, so recovery must be automatic, not a human paging at 3 a.m. That means health-checked failover (automated probes detect a failing server and traffic switches to a standby copy) with the service running in several zones (separate data centres in one city) or regions (separate geographic areas), so one site failing does not take the service down.
- Measure from the customer's view and agree an error-budget policy. The error budget is the allowed downtime; when it is spent, feature work yields to reliability work. Review quarterly with real incident data.
Pitfalls
- Promising 99.99% in a contract before knowing the architecture can deliver it.
- Measuring availability at the server rather than from where users sit.
- Treating the minutes as evenly spread: a rare 4-hour outage is worse than many 2-minute blips for trust.
Unlock Full Question Bank
Get access to all 6 Business Acumen and Commercial Context interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.