Business Acumen and Commercial Context Questions
Reasoning about how a business earns and spends money, and judging the commercial consequence of a decision from the candidate's own seat, at a generic (company-agnostic) level. Covers translating technical, product or data work into revenue, cost, margin and risk terms: model accuracy, precision and recall choices and decision thresholds priced in dollars; latency, availability and incident downtime turned into revenue at stake; build-versus-buy, vendor, architecture, multi-region, single-tenant and technical-debt decisions weighed commercially. Also covers trade-offs such as growth versus profitability, speed versus cost, engagement versus revenue, and opportunity cost and the cost of delay between competing investments; how an engineering, data or platform function creates and demonstrates business value, including indirect contributions; connecting functional plans and OKRs to company strategy and revenue goals; and explaining the commercial case to non-technical leaders and finance partners, with stated assumptions and uncertainty. Tests whether a candidate can tie day-to-day choices to commercial outcomes and explain that link clearly. Full investment business cases, KPI design, computing unit economics, pricing decisions, case-interview frameworks and researching a specific company are covered elsewhere.
Leadership asks you to estimate what an hour of downtime costs the business, quickly enough that responders can use it to prioritise during an incident. What inputs would you use, how would you handle the uncertainty, and how would you present the number?
Sample Answer
Direct answer
Build the estimate from four components: revenue that is actually lost (not just delayed), the profit on it, contractual credits owed, and the fixed cost of the response. Compute it ahead of time as a small table by time-of-day band so a responder can look it up in seconds, show it as a range with a midpoint, and label it as an estimate. In the illustration below one hour costs about $12,200 in the evening peak, $7,700 in the day and $4,100 overnight, with a peak range of $9,650 to $14,810.
Inputs
Definitions: revenue at risk is what the service would have earned that hour. Unrecovered share is the fraction of that revenue that never comes back (customers who buy later, retry or wait have only delayed their purchase). Contribution margin is revenue minus variable costs, so lost contribution is the real profit lost. SLA credits (service-level agreement credits) are refunds promised to customers when availability drops under the contract.
Illustrative inputs for an online store (replace with your own): AOV $80; orders per hour 600 at weekday-evening peak, 300 in the day, 60 overnight; contribution margin 25%; 12 engineers at $100 per hour loaded (salary plus benefits and overhead, divided by hours worked); 0.25 support tickets per lost order at $6 each; $2,000 of contractual credits per outage hour.
The calculation (peak hour, midpoint 70% unrecovered)
| Component | Calculation | Result |
|---|---|---|
| Revenue at risk | 600 orders x $80 | $48,000 |
| Lost revenue (70% unrecovered) | $48,000 x 70% | $33,600 |
| Lost contribution | $33,600 x 25% | $8,400 |
| Responder time | 12 x $100 | $1,200 |
| Extra support | 600 x 70% x 0.25 x $6 | $630 |
| SLA credits | assumed | $2,000 |
| Total cost of the hour | $12,230 |
The lookup table for responders
| Band | Orders per hour | Revenue at risk | Cost at 50% / 70% / 90% unrecovered |
|---|---|---|---|
| Evening peak | 600 | $48,000 | $9,650 / $12,230 / $14,810 |
| Daytime | 300 | $24,000 | $6,425 / $7,715 / $9,005 |
| Overnight | 60 | $4,800 | $3,845 / $4,103 / $4,361 |
Handling the uncertainty
- Give a range, and move only the inputs that are uncertain. Here the unrecovered share is the main unknown; measure it from past outages by comparing the hours after recovery with a normal day. For example, if an outage removed 600 orders and the three hours after recovery ran 180 orders above the normal level for that time of day, 30% of the demand came back and 70% did not, which is where the 70% midpoint comes from.
- Responder time is mostly sunk (the engineers are paid anyway). Keep it as a line so the number is complete, but do not let it hide the revenue effect.
- Do not add long-term brand damage or churn into the headline number. If you believe it matters, show it as a separate, labelled line with a stated assumption.
How to present it
During an incident: one line, "about $12k per hour right now, plausible range $10k to $15k", plus the band used. It lets responders prioritise (a payment outage in the peak band beats a reporting bug) without arguing about precision. After the incident: show the components and the assumptions, and re-estimate.
Pitfalls
- Using annual revenue divided by 8,760 hours gives a flat $4,566 per hour for a $40M business, which understates the peak and overstates overnight.
- Reporting lost revenue and lost profit as if they were the same number. Choose one for the headline and label it.
- Partial outages: scale by the fraction of users or functions affected. If 40% of users are affected in the evening peak, the revenue-driven lines scale (lost contribution $8,400 x 40% = $3,360; extra support $630 x 40% = $252), while responder time stays $1,200 and the $2,000 of credits is assumed to still apply, giving $1,200 + $2,000 + $3,360 + $252 = $6,812 for the hour.
Some functions, such as platform engineering, security or data, affect revenue and cost only indirectly. How would you quantify and report their contribution to executives, including how you would show uncertainty?
Sample Answer
Direct answer
Translate each function's work into the three things executives already reason about, then show a range instead of a single number. The three are: avoided loss (probability of a bad event times its cost, reduced by the function's work), capacity released (engineer or analyst time freed and what it was redeployed to), and revenue enabled (deals or launches that depended on the function). Report the median and a 10th to 90th percentile range, say which inputs are measured and which are judgement, and show the chance that the function's value falls below its cost.
Step 1: build a causal chain per function
Write one line per stream: activity, then mechanism, then outcome, then dollars. Example for a platform team: "faster, safer deploys, so engineers lose fewer hours to waiting and rework, so more time goes to roadmap work, worth the loaded hourly cost times the share that is really redeployed." Each link is a number someone can challenge. Prefer a counterfactual ("what would have happened without us?") over a raw activity count ("we blocked 4,000 attacks").
Step 2: quantify with ranges and simulate
Illustrative setup for a combined platform and security programme costing $0.4M a year to extend: 120 engineers working 46 weeks a year at a loaded hourly cost of $85 (salary plus benefits and overhead per hour worked); 1 to 3 hours saved per engineer per week; 30% to 70% of that time becomes shipped work; a 12% to 24% yearly chance of a material incident (one serious enough to cost real money or customers) cut by 10% to 30% by new controls; a median incident cost of $1.0M to $2.5M; and 1 to 4 deals a year where a security attestation (for example, a SOC 2 report, an independent audit of security controls) was a stated buyer requirement, each $150k ACV at 80% gross margin, with security credited for 30% to 70% of each.
This is a Monte Carlo simulation: instead of one guess per input, draw a random value from each input's range, compute the result for that one draw, repeat 20,000 times, and read percentiles off the pile of results. One draw by hand for the platform stream: 2 hours saved and 50% redeployed gives 120 x 2 x 46 x $85 x 0.5 = $469,200, or $0.47M. For the security stream, with an 18% incident chance, a 20% relative cut and a $1.75M median loss: losses are skewed, because most incidents cost near the median and a few cost far more. The code models that with a lognormal distribution, a skewed shape whose logarithm is bell-shaped, where sigma (0.8 here) sets how long the tail is. For that shape the mean exceeds the median by the factor exp(sigma^2 / 2) = exp(0.32) = 1.377, so the $1.75M median means a $2.41M mean, and the expected avoided loss (incident chance x relative cut x mean loss) is 0.18 x 0.20 x $2.41M = $0.087M. Each of the 20,000 draws is one such calculation with different random inputs.
import math, random
rng = random.Random(7)
N = 20000
sec, plat, deals = [], [], []
for _ in range(N):
p_before = rng.uniform(0.12, 0.24) # annual chance of a material security incident
rel_cut = rng.uniform(0.10, 0.30) # relative reduction from the controls
median_loss = rng.uniform(1.0, 2.5) # $M, median loss if an incident happens
mean_loss = median_loss * math.exp(0.8 ** 2 / 2) # lognormal with sigma 0.8
sec.append(p_before * rel_cut * mean_loss) # expected avoided loss, $M/yr
hours = rng.uniform(1.0, 3.0) # hours saved per engineer per week
redeploy = rng.uniform(0.3, 0.7) # share of saved time that becomes shipped work
plat.append(120 * hours * 46 * 85 * redeploy / 1e6) # 120 engineers, 46 weeks, $85/hour
n_deals = rng.randint(1, 4) # deals that named a security attestation as a requirement
credit = rng.uniform(0.3, 0.7) # share of each deal credited to security
deals.append(n_deals * 0.150 * 0.8 * credit) # $150k ACV at 80% gross margin, $M
total = [a + b + c for a, b, c in zip(sec, plat, deals)]
def pct(xs, q):
xs = sorted(xs)
return xs[int(q * (len(xs) - 1))]
for name, xs in [("avoided loss", sec), ("platform time", plat), ("deals enabled", deals), ("total", total)]:
print(f"{name:14s} P10 {pct(xs, .1):.2f}M P50 {pct(xs, .5):.2f}M P90 {pct(xs, .9):.2f}M")
print(f"share of draws below the 0.4M programme cost: {sum(x < 0.4 for x in total) / N:.1%}")
Output:
avoided loss P10 0.04M P50 0.08M P90 0.14M
platform time P10 0.26M P50 0.44M P90 0.73M
deals enabled P10 0.05M P50 0.14M P90 0.26M
total P10 0.46M P50 0.69M P90 0.98M
share of draws below the 0.4M programme cost: 3.9%
P10, P50 and P90 are the 10th, 50th and 90th percentiles of the simulated outcomes. Medians do not add exactly (0.08 + 0.44 + 0.14 is 0.66, not 0.69) because the total's percentiles come from summing each draw.
Step 3: what to tell executives
- Headline: "Median value about $0.69M a year against $0.4M cost; 80% range $0.46M to $0.98M; in about 4% of simulated cases it does not cover its cost."
- Where the value comes from: most of it is the platform-time stream, and that is also the softest assumption (the share of saved time that becomes shipped work). Name it as the number to challenge and as the one worth measuring.
- Measured vs assumed: label each input. Deploy lead time and incident counts are measured; incident probability and redeployment share are estimates.
- Leading indicators (early signals that move before the dollars do) you can track quarterly: deploy frequency, mean time to restore service (the average time to recover from an incident), audit findings closed, deals where a security review blocked or delayed signature.
- Update the ranges each quarter as evidence arrives, so uncertainty shrinks visibly.
Trade-offs and pitfalls
- Avoided-loss estimates rest on rare events, so the data cannot confirm them. Say so, and use industry incident data only after verifying the source.
- Do not double count: a deal credited to security should not also be fully credited to sales. The credit share handles that.
- Hours saved are not dollars saved unless the time is redeployed or headcount changes.
- A single point estimate looks authoritative and invites a fight over one number; a range with named drivers moves the discussion to the evidence.
Leadership wants to know whether to break a growing monolith into services or keep scaling it as it is. How would you frame the decision commercially in terms of engineering time, delivery speed, operating risk and cost, and what would you propose?
Sample Answer
Direct answer
Frame it as an investment with a payback, not an architecture preference. Measure what the monolith currently costs the business (engineer time lost to coordination, slow releases, outages), price each option, and compare the payback. My proposal for most organisations in this position: keep the monolith, enforce clear module boundaries inside it (a modular monolith), and extract a service only where there is measured evidence, such as one component that needs to scale or be released independently. A full split pays back in under two years only if services remove nearly all of the coordination cost, which is rarely the case.
Definitions
A monolith is one deployable application. Microservices are many independently deployable services that call each other over a network. Coordination tax is the share of engineering time lost to waiting for other teams, merge conflicts, shared release trains and fixing cross-team breakages. Loaded cost is salary plus benefits and overhead per engineer-year. A modular monolith keeps one deployable but enforces module boundaries in the code. In plain terms: a shop's website written as one program that handles catalogue, cart and payments is a monolith; the same shop as separate catalogue, cart and payments programs that call each other over the network is microservices; the modular monolith is still one program, but the catalogue, cart and payments code may only talk to each other through defined interfaces. A release train is a fixed schedule on which everything ships together, so a team that misses it waits for the next one. Lead time is the elapsed time from a change being merged to it running in production, and deploy frequency is how often a team releases. Blast radius is how much of the system, and how many customers, one failure or bad release can affect. A hot path is the code that handles the highest-volume or most latency-sensitive work. A domain boundary is the line between business areas (catalogue, cart, payments) that can change independently, and a regulatory boundary is a part of the system that must follow stricter rules (for example, payment card data) and so benefits from being isolated. A distributed monolith is a set of services that must still be changed and deployed together, so it pays the costs of both designs.
The commercial frame: four lenses
| Lens | What to measure | Monolith | Full split into services |
|---|---|---|---|
| Engineering time | Share of capacity lost to coordination; migration effort | Pays the coordination tax each year | Large one-off migration, then smaller tax |
| Delivery speed | Lead time from merge to production, deploy frequency | Slow if one release train blocks all teams | Faster per team once boundaries are stable; slower during migration |
| Operating risk | Incident rate and blast radius | One failure can take everything down; simpler to run and debug | Failures are contained but more of them: network calls, partial failures, distributed debugging |
| Cost | Infrastructure and on-call load | Cheaper to run | More running parts, extra platform and tooling cost |
Worked example (all inputs illustrative)
60 engineers at $200,000 loaded cost per engineer-year, a coordination tax measured at 25% of capacity: 60 x 25% x $200,000 = $3.0M a year lost.
| Option | Up-front cost | Extra run cost per year | Share of the tax recovered | Net annual benefit | Payback |
|---|---|---|---|---|---|
| Full split | 15 engineers x 1.5 years x $200,000 = $4.5M | $360,000 | 25% | $390,000 | 11.5 years |
| Full split | $4.5M | $360,000 | 50% | $1,140,000 | 3.9 years |
| Full split | $4.5M | $360,000 | 100% | $2,640,000 | 1.7 years |
| Modular monolith plus one extraction | 5 engineers x 1 year x $200,000 = $1.0M | $60,000 | 20% | $540,000 | 1.9 years |
| Modular monolith plus one extraction | $1.0M | $60,000 | 40% | $1,140,000 | 0.9 years |
Where the illustrative inputs come from, and how a team would measure them. The 25% tax is what a four-week time diary or ticket sample might show: if engineers log about 10 of 40 weekly hours on waiting for other teams, cross-team reviews, merge conflicts and release coordination, that is 25%. The full split is costed at 15 engineers (a quarter of the 60) for 1.5 years because every team has to carve its code and data out of the shared application, and its extra run cost of $360,000 as, for example, one more platform engineer ($200,000) plus $160,000 of extra infrastructure and monitoring. The extraction is costed at 5 engineers for 1 year because it moves one component, with $60,000 of extra run cost. The shares of the tax recovered are hypotheses, not facts. The modular monolith plus one extraction can only remove the part of the tax tied to tangled module boundaries and to the one extracted component, so I assume 20% to 40%. A full split could in principle remove nearly all cross-team waiting, but it swaps some of it for new coordination (versioned interfaces, cross-service changes), so the range is wide, 25% to 100%. The first four weeks of measurement in the plan below replace these guesses with data.
To pay back the full split within 2 years it must earn back the $4.5M up-front cost in 2 years, which is $4.5M / 2 = $2.25M a year, and also cover the $360,000 a year of extra run cost: $2.25M + $0.36M = $2.61M of yearly savings. As a share of the entire $3.0M tax that is $2.61M / $3.0M = 87%. That is the number to put in front of leadership.
What I would propose
- Spend four to six weeks measuring the tax: deploy frequency, lead time, share of pull requests that touch more than one team, incidents caused by cross-team changes.
- Enforce module boundaries in the monolith (ownership per module, dependency rules checked in the build).
- Extract the one or two components with a measured reason: a hot path that scales differently, a team that is blocked by the shared release train, or a regulatory boundary.
- Re-measure after each extraction, with a stop rule: if the measured tax recovered is under half of what the model assumed, pause.
What would flip the call
- A coordination tax that stays large across many teams after module boundaries have been enforced and tried favours a larger split.
- Very different scaling or availability needs between parts of the system favour extraction of those parts.
- A small team usually cannot recover a migration of this size, because the tax it pays is small in absolute dollars.
Pitfalls
- Treating the move as a cost-saving: the run cost rises. The justification is speed and risk, priced.
- Underestimating migration drag: engineers on migration are not shipping features, which is why it is a cost in the table.
- Splitting before the domain boundaries are understood, so services share a database and still deploy together (a distributed monolith).
A proposed reliability project would cut outages by 40%, but customers would never see it. How would you make the case for it to executives, and how would you prove its value during a pilot?
Sample Answer
Direct answer
Make the invisible visible by pricing it: convert "40% fewer outages" into avoided cost per year, add the benefits customers never see but the company does (engineer toil, on-call load, delivery speed), and be honest about payback. Then prove it in a pilot, but not with outage counts alone: outages are too rare for a short pilot to show a 40% change, so the pilot must be judged on leading indicators that occur often enough to measure, with a control group and success criteria agreed beforehand.
Making the case to executives
Frame it as insurance: a known cost now against an expected loss that falls. Illustrative inputs: 10 major outages a year, 50 minutes each, $300 of lost revenue per minute, $20,000 per outage of response time and service credits, and a $220,000 one-off project cost.
OUTAGES_PER_YEAR = 10
MINUTES_PER_OUTAGE = 50
REVENUE_LOST_PER_MINUTE = 300
RESPONSE_AND_CREDITS_PER_OUTAGE = 20_000 # engineer time plus SLA credits, illustrative
PROJECT_COST = 220_000 # one-off, illustrative
REDUCTION = 0.40
per_outage = MINUTES_PER_OUTAGE * REVENUE_LOST_PER_MINUTE + RESPONSE_AND_CREDITS_PER_OUTAGE
annual = OUTAGES_PER_YEAR * per_outage
avoided = annual * REDUCTION
print(f"cost per outage ${per_outage:,}; yearly outage cost ${annual:,}")
print(f"outages avoided per year {OUTAGES_PER_YEAR * REDUCTION:.0f}; value ${avoided:,.0f} a year")
print(f"payback on outage cost alone: {PROJECT_COST / avoided * 12:.1f} months")
Output:
cost per outage $35,000; yearly outage cost $350,000
outages avoided per year 4; value $140,000 a year
payback on outage cost alone: 18.9 months
On outage cost alone, payback is 18.9 months, which is a sober number to show rather than hide. Then add what that number leaves out:
- Tail risk. The average hides the rare bad day; one 4-hour outage costs 4 x 60 x $300 = $72,000 in lost revenue before response costs. Reducing the chance of the long tail is worth more than the average suggests.
- Engineering time. Hours spent in incident response, postmortems and manual recovery, and on-call burnout (attrition costs are real but must be estimated, not asserted).
- Delivery speed. Fewer failed releases means more deploys, which is a feature-velocity benefit.
- Customer-visible proxies. Support tickets, SLA credits and renewal risk for accounts that have had incidents.
Ask for a staged decision: fund the pilot on the two most incident-prone services, with a go or no-go review at a set date. The pitch an executive would hear: "Outages cost us about $350,000 a year. This project removes about 40% of them, roughly $140,000 a year, and also cuts the pages that wake our engineers. It costs $220,000 once, so on outages alone it pays back in about 19 months, sooner once engineering time is counted. I am asking you to fund a pilot on two services, with success criteria we agree today and a review date. If we miss them, we stop."
Proving value in a pilot
Design. Pick pilot services and comparable control services that do not get the project in the same period (same traffic, similar change rate). Agree the metrics, the period and the success threshold before starting.
Why outage counts are not enough. A pilot's expected outages are few. In plain words, the code below answers one question: if the project really does cut outages by 40%, how often would a pilot of a given size show it convincingly, rather than looking like luck? That chance is called power. Suppose the treated and control groups have equal exposure, outage counts are Poisson (random, independent events), and the treated group truly has 40% fewer. The chance that a one-sided exact test at the 5% level detects it (the power) depends on how many outages the control group is expected to have over the pilot:
import math
def pois_pmf(lam, kmax):
p = [math.exp(-lam)]
for k in range(1, kmax + 1):
p.append(p[-1] * lam / k)
return p
def power(control_events, reduction=0.4, alpha=0.05):
"""Expected events in the control group over the pilot; treated group has
(1 - reduction) times as many. Equal exposure. One-sided exact test: given the
total N, the treated count is Binomial(N, 0.5) under no effect."""
lc, lt = control_events, control_events * (1 - reduction)
kmax = int(lc + 12 * math.sqrt(lc) + 20)
pc, pt = pois_pmf(lc, kmax), pois_pmf(lt, kmax)
# largest treated count that is significant, for each total N
cache = {}
def crit(N):
if N not in cache:
cum, best = 0.0, -1
for i in range(N + 1):
cum += math.comb(N, i) / 2 ** N
if cum <= alpha:
best = i
else:
break
cache[N] = best
return cache[N]
total = 0.0
for nt in range(kmax + 1):
for nc in range(kmax + 1):
if nt <= crit(nt + nc):
total += pt[nt] * pc[nc]
return total
for ctrl in (5, 10, 20, 50, 100):
print(f"control events {ctrl:>3}, treated expected {ctrl*0.6:>5.1f}, power {power(ctrl):.2f}")
Output:
control events 5, treated expected 3.0, power 0.09
control events 10, treated expected 6.0, power 0.19
control events 20, treated expected 12.0, power 0.35
control events 50, treated expected 30.0, power 0.69
control events 100, treated expected 60.0, power 0.93
With 5 or 10 expected outages in the control group, power is 9% to 19%; you would need about 100 to be 93% likely to detect the effect. So a quiet pilot proves nothing either way.
How to read the code. The two groups' outage counts are drawn at random from the Poisson probabilities (pois_pmf), which give the chance of 0, 1, 2, ... outages when the average is known. For every pair of counts, the test asks whether the treated group's share of the combined outages is low enough to be surprising. The exact test works like this: if the project did nothing, each outage would be a coin flip between the two groups, so the treated count out of N total outages follows Binomial(N, 0.5). With N = 20 total outages, the treated group having 5 or fewer happens by chance about 2% of the time (6 or fewer is about 6%, which is over the 5% cut-off), so 5 is the critical value for N = 20; the function crit finds that cut-off for each N and caches it. The final double loop adds up the probability of every pair of counts that lands at or below the cut-off. "One-sided" means we only ask whether the treated group is lower, and "5% level" means that when nothing has changed the test wrongly declares success no more than 5% of the time.
Use measures that happen often:
- Change failure rate (the share of deployments causing an incident or rollback).
- Pages (alerts that wake an engineer) per week, and the share that are actionable.
- Faults caught by the new safeguards before customers saw them (near-misses), such as a bad deploy that automatic rollback reverted in two minutes before anyone noticed, or injected failures (a server deliberately switched off in a test) that the system survived.
- Time to detect and time to recover in incidents that do occur.
- Customer-facing minutes of impact, SLA credits (refunds promised in a service-level agreement) and support tickets as lagging confirmation.
A common target is 80% power: an 80% chance of detecting a real 40% effect. With 50 expected events in the control group, power is about 69%, just short; with about 70 it is about 82%, and with 100 it is 93%. So choose indicators that should produce roughly 70 or more events in the control group over the pilot (about 5 a week over a 13-week pilot), and treat anything below that as too thin to read.
Trade-offs and pitfalls
- Choosing the pilot services because they are easy makes the result unrepresentative; choose by incident history and report that choice.
- Regression to the mean (extreme results drift back toward normal on their own): if a service had 8 outages last quarter when it usually has about 4, it will probably have 4 or 5 next quarter with no project at all, which would look like a 40% improvement. A control group chosen the same way shows the drift, which is why it matters.
- Declaring victory on a quiet quarter: absence of outages is the expected outcome in many quarters.
- Agree the kill criteria too: what result would make you stop?
A change cuts page latency from 500ms to 200ms. How would you quantify what that is worth to the business, what do you need to know to do it, and where does such an estimate usually go wrong?
Sample Answer
Direct answer
The value of a latency cut is the extra conversions it produces times what each conversion is worth, minus the cost of delivering the speed-up. You need five inputs: traffic and baseline conversion, contribution per conversion, (most important) a measured relationship between latency and conversion from your own site, which is called the elasticity: the percentage change in conversion for each 100 ms of speed gained, which latency and which users the speed-up applies to, and the cost of delivering it. Such estimates usually go wrong by borrowing someone else's elasticity, extrapolating a straight line, confusing correlation with cause, and forgetting the running cost.
What you need to know
- Traffic and conversion. Sessions per month (10,000,000 in the example) and baseline conversion (2.5%, so 250,000 conversions a month). Conversion here means a session ending in a purchase.
- Value per conversion. Contribution per conversion (revenue less variable cost), $40 in the example. Using revenue overstates the benefit.
- The latency-to-conversion relationship. The relative change in conversion for each 100 ms saved. This is the number everything hangs on and it cannot be assumed.
- Which latency and which users. Median or 95th percentile (p95, the speed that 95% of loads beat)? Which pages (product page versus checkout) and which devices? A 500 ms to 200 ms improvement on the median may be a 3000 ms to 2900 ms change for the slowest users.
- Cost to deliver. A caching layer costs $18,000 a month in infrastructure and upkeep, in this example.
How to get the relationship
Best: a randomised experiment that adds an artificial delay (for example 100 or 300 ms) to a small slice of traffic and measures conversion loss. Next best: a natural experiment (a change nobody planned as a test but that you can compare before and after, such as a deploy that happened to change speed), keeping the device mix the same and adjusting for confounders, meaning other things that changed at the same time, such as a sale or a new ad campaign. Weakest: correlation across pages, because fast pages are usually simple pages and fast users have good devices.
Worked example
SESSIONS_PER_MONTH = 10_000_000
CONV_RATE = 0.025
CONTRIBUTION_PER_CONV = 40.0 # dollars per conversion after variable costs
CACHE_COST_PER_MONTH = 18_000 # infrastructure plus upkeep of the caching layer
MS_SAVED = 300
conversions = SESSIONS_PER_MONTH * CONV_RATE
print(f"baseline conversions per month: {conversions:,.0f}")
print("assumed relative conversion gain per 100 ms saved -> value and net per month")
for per_100ms in (0.0025, 0.005, 0.01, 0.015):
relative_gain = per_100ms * MS_SAVED / 100
value = conversions * relative_gain * CONTRIBUTION_PER_CONV
print(f"{per_100ms:.2%} per 100 ms: +{relative_gain:.2%} = {conversions*relative_gain:>7,.0f} conversions, ${value:>8,.0f}/month, net ${value - CACHE_COST_PER_MONTH:>8,.0f}")
needed_conversions = CACHE_COST_PER_MONTH / CONTRIBUTION_PER_CONV
total_gain = needed_conversions / conversions
print(f"break-even: {needed_conversions:,.0f} extra conversions a month = {total_gain:.2%} total = {total_gain / (MS_SAVED / 100):.2%} per 100 ms")
Output:
baseline conversions per month: 250,000
assumed relative conversion gain per 100 ms saved -> value and net per month
0.25% per 100 ms: +0.75% = 1,875 conversions, $ 75,000/month, net $ 57,000
0.50% per 100 ms: +1.50% = 3,750 conversions, $ 150,000/month, net $ 132,000
1.00% per 100 ms: +3.00% = 7,500 conversions, $ 300,000/month, net $ 282,000
1.50% per 100 ms: +4.50% = 11,250 conversions, $ 450,000/month, net $ 432,000
break-even: 450 extra conversions a month = 0.18% total = 0.06% per 100 ms
- At an assumed 0.5% relative conversion gain per 100 ms, saving 300 ms gives +1.5%, which is 3,750 extra conversions and $150,000 a month; net of the $18,000 cache cost, $132,000 a month, a return on cost of about 7.3 times ($132,000 / $18,000).
- The result is very sensitive to the assumed elasticity: it swings from $57,000 to $432,000 net a month across the range tried, while break-even is only 0.06% per 100 ms. So the decision to build is robust, but the headline benefit number is not and should be quoted as a range.
Where such estimates usually go wrong
- Borrowed statistics. Another company's "100 ms = 1%" is a story about its users and pages.
- Linear extrapolation. Gains tend to shrink below a speed where users stop noticing, and may be larger for slow users; a straight line over 300 ms overstates.
- Correlation as cause. Fast pages convert better partly because of what they are.
- Averages versus tails. The benefit may sit in the p95 population, which the median hides.
- Double counting. Double counting means claiming the same benefit twice, for example counting both higher conversion and lower bounce rate (the share of visitors who leave after one page; it is a symptom of the same effect), or adding SEO gains (better search-engine ranking from faster pages) that the experiment cannot see and that you have not measured.
- Forgetting running cost and opportunity cost. The cache needs maintenance, invalidation bugs (cache invalidation is deciding when stored copies of a page are out of date and must be refreshed) can serve stale prices, and the engineers could be doing something else.
- Pulled-forward demand. Some conversions would have happened anyway, later or through another channel; a faster page may only make a shopper buy today instead of next week, so the net gain is smaller than the count of extra conversions in the window.
What I would present
A range with a stated method ("measured 0.5% per 100 ms from a delay experiment, interval 0.3% to 0.7%"), the cost, the break-even elasticity, and the one thing that would change the call. Spoken version: "We measured that every 100 ms we remove lifts conversion by about 0.5% relative, with a plausible range of 0.3% to 0.7%. Cutting 300 ms is therefore worth about $150,000 a month, between $90,000 and $210,000, against $18,000 a month to run. It stops paying for itself only if the true effect is below 0.06% per 100 ms, far under anything we measured. The one thing that would change my call is if the effect turns out to be concentrated in users the cache does not reach."
Unlock Full Question Bank
Get access to all 14 Business Acumen and Commercial Context interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.