Business Case Development and ROI Analysis Questions
Building and defending the numeric case that an investment or initiative is worth funding. Covers enumerating one-time and recurring cost categories (implementation, licensing, TCO, lifecycle cost, licence-model comparisons), quantifying tangible and intangible benefits and converting efficiency gains, time-to-market, or risk reduction into cash terms, applying ROI, payback, NPV and break-even to a specific proposal (including choosing a defensible discount rate and phasing benefits over time), build-versus-buy and option comparisons, stating and stress-testing assumptions (sensitivity and scenario ranges, optimism bias, auditing a vendor or customer model), packaging the case for an approver (one-page summary, CFO or executive framing, staged funding), deciding between competing initiatives, opportunity cost, and whether to continue, pivot or stop, building customer-facing and vendor-side cases (a competing TCO claim, a bespoke feature or custom integration, a multi-year discount), and tracking realized benefits against the forecast after approval. Accounting treatment and full financial models are outside this topic.
A customer asks whether to adopt a multi-cloud setup for resilience. Build a quantified three-year case that compares it with the alternatives, including doing nothing, and tell me how you would present the decision.
Sample Answer
Direct answer
I would build a three-option, three-year expected-cost comparison: do nothing (single region), multi-region inside one cloud, and multi-cloud, with the customer's own outage cost and outage history as the key inputs. On the illustrative inputs below, doing nothing is cheapest in expected terms, multi-region becomes worthwhile once an outage hour costs more than about $82,000, and multi-cloud beats doing nothing only above about $228,000 an hour (and beats multi-region only above about $1.1 million an hour) or when a rule demands it (for example a regulator's concentration-risk requirement, meaning a rule that you must not depend on a single provider for critical services). Expected means a probability-weighted average, not a forecast of any one year. So my default recommendation is to protect the critical tier (the few workloads whose outage hurts the business most) with multi-region in one cloud, and to treat multi-cloud as a requirement-driven exception, not a resilience upgrade by default.
What I would ask the customer first
- What does an hour of full outage cost (lost revenue, penalties, recovery effort)? This is the most sensitive input.
- How many hours of provider or regional outage have you actually seen in the last 2 to 3 years? For example, outages of 5, 4 and 3 hours over three years give 12 / 3 = 4.0 hours a year, which is where the 4.0 below comes from.
- Which workloads are critical, and what recovery targets (RTO: how long you can be down; RPO: how much data you can lose) do they really need?
- Is there a regulatory or contractual reason to avoid depending on one provider?
The case
Every input is an assumption to be replaced with the customer's own. SLA means service-level agreement. Do nothing = single region, no extra spend. Multi-region = a second region in the same cloud (SLAs and tooling are shared, data replication is simpler). Multi-cloud = a second provider in active/passive mode (the second site is ready but idle). Outage hours per year are the expected hours caused by a provider or region failure, after each option: 4.0, 1.0 and 0.5. The multi-cloud figure is not zero because failover (switching the work to the standby site) itself can fail: configuration drift (the standby slowly stops matching production because changes reached only one side) and untested runbooks (step-by-step recovery instructions nobody has rehearsed).
| Option | Expected outage hours a year | One-time cost | Extra yearly run cost |
|---|---|---|---|
| A. Do nothing | 4.0 | $0 | $0 |
| B. Multi-region, one cloud | 1.0 | $250,000 | $150,000 |
| C. Multi-cloud, active/passive | 0.5 | $900,000 | $450,000 |
The extra cost lines include duplicate infrastructure, data replication and egress fees (what a provider charges to move data out of its cloud), extra tooling and staff skills, failover testing, and the loss of provider-specific managed services you could no longer rely on. Lock-in risk is a benefit of C: I would price it as a possible price increase on the main provider (for example, 5% on $2M a year of spend is $100,000 a year, illustrative). Counting it would add about $258,000 of present value to C's benefit ($100,000 x 2.5771, where 2.5771 is 1/1.08 + 1/1.08^2 + 1/1.08^3, the three-year discount factors at 8%), which does not change the ranking.
# Three-year expected outage cost, USD. Every input is an assumption to be replaced with the customer's own.
COST_PER_HOUR = 30_000 # revenue lost + recovery effort per hour of full outage
options = {
# name: (expected provider/region outage hours per year, one-time cost, extra annual run cost)
"A. Do nothing (single region)": (4.0, 0, 0),
"B. Multi-region, one cloud": (1.0, 250_000, 150_000),
"C. Multi-cloud active/passive": (0.5, 900_000, 450_000),
}
RATE = 0.08
def pv_years(amount, years=3):
return sum(amount / (1 + RATE) ** t for t in range(1, years + 1))
for name, (hours, one_time, run) in options.items():
outage = pv_years(hours * COST_PER_HOUR)
total = one_time + pv_years(run) + outage
print(f"{name:<32} outage PV {outage:>9,.0f} spend PV {one_time + pv_years(run):>9,.0f} total {total:>9,.0f}")
# Break-even: outage cost per hour at which B and C beat doing nothing
for name, (hours, one_time, run) in list(options.items())[1:]:
saved_hours = 4.0 - hours
spend = one_time + pv_years(run)
be = spend / (pv_years(1) * saved_hours)
print(f"{name}: break-even cost per outage hour = {be:,.0f} USD")
# Capital programme vs supplier price cuts vs status quo (3 years, USD millions)
def npv(rate, flows):
return sum(f / (1 + rate) ** t for t, f in enumerate(flows))
for label, saving in (("low", 1.8), ("base", 2.5), ("high", 3.0)):
print(f"6M capex programme, {label:<4} saving {saving}M/yr: NPV {npv(RATE, [-6.0, saving, saving, saving]):6.2f}M")
print(f"Supplier price cut, 1.2M/yr, 0.1M to negotiate: NPV {npv(RATE, [-0.1, 1.2, 1.2, 1.2]):6.2f}M")
print("Status quo: NPV 0.00M (by definition the baseline)")
Output:
A. Do nothing (single region) outage PV 309,252 spend PV 0 total 309,252
B. Multi-region, one cloud outage PV 77,313 spend PV 636,565 total 713,877
C. Multi-cloud active/passive outage PV 38,656 spend PV 2,059,694 total 2,098,350
B. Multi-region, one cloud: break-even cost per outage hour = 82,336 USD
C. Multi-cloud active/passive: break-even cost per outage hour = 228,351 USD
6M capex programme, low saving 1.8M/yr: NPV -1.36M
6M capex programme, base saving 2.5M/yr: NPV 0.44M
6M capex programme, high saving 3.0M/yr: NPV 1.73M
Supplier price cut, 1.2M/yr, 0.1M to negotiate: NPV 2.99M
Status quo: NPV 0.00M (by definition the baseline)
How to read it: over three years at 8%, expected cost is $309,252 for A, $713,877 for B and $2,098,350 for C, with outage cost at $30,000 an hour. Extra spend only pays back if the outage hour is expensive: B breaks even against doing nothing at about $82,336 an hour and C at about $228,351. Those are both measured against doing nothing. Against B, the extra spend of C ($2,059,694 minus $636,565) saves only 0.5 outage hours a year, so C beats B only above (2,059,694 - 636,565) / (2.5771 x 0.5) = about $1,104,000 an hour; at $228,351 an hour B costs about $1.23M over three years and C about $2.35M. That is why multi-region is the default and multi-cloud is a requirement-driven exception.
The three-way comparison for a capital programme (same method, different decision)
A customer deciding how much resilience to buy is usually also deciding how to spend the wider budget, so I would put the resilience options next to the other candidates for the same money. This is the same present-value method on a different decision, unrelated to the multi-cloud numbers above. Here: a $6M capex programme (capex is one-off capital spend) versus a supplier price cut versus the status quo, three years at 8%, in USD millions. The programme saves $2.5M a year in the base case: NPV (net present value: future cash flows converted to today's dollars) +$0.44M; at a low $1.8M saving it is -$1.36M; at a high $3.0M, +$1.73M. A supplier price cut of $1.2M a year that costs $0.1M to negotiate has NPV +$2.99M. Status quo is the zero baseline. On these numbers the price cut beats even the high programme case ($2.99M against $1.73M) because it costs almost nothing, so the programme only wins if the price cut cannot be obtained. The two are not exclusive, but they overlap: if the programme's savings come from cutting spend with that same supplier, a $1.2M a year price cut removes part of the spend the programme would have saved. In the extreme case where all $1.2M overlaps, the programme's saving falls from $2.5M to $1.3M a year and its NPV from +$0.44M to about -$2.65M (-6.0 + 1.3 x 2.5771), so I would negotiate the price first and re-test the programme on the new base.
Presenting the decision
One page, in this order: (1) the recommendation and the decision requested; (2) the three options in a table with three-year cost, expected outage hours and break-even outage cost; (3) the assumptions the customer must confirm, with an owner each; (4) what I did not model, including tail events (rare, extreme outcomes far from the average): an expected value, the probability-weighted average, hides a rare, long outage, so I would show the worst credible case (for example 48 hours at $30,000 an hour is $1.44M) next to the averages, because a customer that cannot survive that event may want protection regardless of the average; (5) the trigger that would change the decision, for instance a measured outage cost above $82,000 an hour moves the critical tier to B.
Trade-offs and pitfalls
- Active/passive multi-cloud adds its own failure modes, so claiming a 100% reduction in outage hours is dishonest.
- Resilience cost ignores people: a team that has never run failover on the second cloud will not manage it well on the day.
- A do-nothing baseline is not free: it carries the expected outage loss. It is simply the cheapest on these inputs.
- Where I would change my call: if the customer's outage cost is above the break-even, if a regulator requires provider diversity, or if a previous outage was longer than the 4 hours assumed.
A customer hands you a spreadsheet that projects a 3x return in two years. How would you check whether it can be trusted, and which errors do you expect to find?
Sample Answer
Direct answer
I would not argue with the 3x; I would rebuild it from its inputs with the author, find where the multiple comes from, and show a bridge from their number to a defensible one. A claim like this is usually a gross figure (benefits divided by cost) with full adoption from day one and a partial cost base. In the illustrative case below, correcting just three things takes the claimed 3.00x down to 0.945x and a negative NPV (net present value, today's value of the cash flows) of -$85,909, which tells me what to expect.
The case under review, in plain words: a customer says a $300,000 spend (licence and build) will return $900,000 over two years, because the benefit is $450,000 a year (say, hours saved and errors avoided) starting at full rate on day one. A bridge is a step-by-step walk from the author's number to the corrected one, showing the effect of each correction. All dollar figures here are illustrative.
Step 0: agree what the claim means
"3x return" is ambiguous. If $300,000 yields $900,000 over two years, that is a 3.0x multiple of benefits to cost, which is an ROI (return on investment, net gain divided by cost) of 200%, not 300%. Ask whether it is gross or net, discounted or not, and over which period.
The eight checks
(If I had thirty minutes I would run 1, 2, 3 and 4 first, because they usually explain most of an inflated multiple; the rest are for a second pass. Terms: a strawman baseline is a deliberately weak alternative that makes the project look good; change management is the work of getting people to actually change how they work; a circular reference is a spreadsheet formula that depends on its own result; go-live is the date the system starts being used; run rate is the benefit per year once fully operating; nominal means undiscounted money.)
- Definition and discounting: multiple or ROI, net or gross, nominal or discounted, which period.
- Adoption: does the model assume 100% of users adopt? Ask for evidence (a pilot, comparable rollouts).
- Ramp and timing: do benefits start on day one at full rate, or after go-live, training and learning?
- Cost completeness: integration, data migration, training, change management, internal staff time, ongoing licence and support. A case with zero training and zero integration effort is a red flag in itself.
- Benefit realism and double counting: hours saved valued at a full loaded rate with no plan to redeploy them; the same saving claimed under two headings.
- Baseline: compared with doing nothing, or with a strawman that no one would actually choose? Does the status quo cost grow in the model?
- Spreadsheet mechanics: hard-coded numbers inside formulas, sums that miss rows, monthly figures treated as annual, sign errors, circular references. Trace each headline cell back to its inputs.
- Stress and ownership: does the case survive inputs 20% worse, and who owns each benefit and its measurement?
Errors I expect, and the bridge from 3.00x
(This section is the main example; the 400% case that follows is a second, separate one.) The code reproduces the claim, then corrects it in three steps using stated, illustrative assumptions (70% adoption instead of 100%; year one at half rate; $90,000 integration, $30,000 training, $40,000 a year run cost).
def npv(rate, flows):
return sum(f / (1 + rate) ** t for t, f in enumerate(flows))
license_build = 300_000 # the whole $300,000 up-front spend the customer quoted: licence plus build
claimed_benefit = 450_000 * 2 # full run-rate from day one, 2 years
print(f"Claim: ${claimed_benefit:,.0f} back on ${license_build:,.0f} = {claimed_benefit / license_build:.2f}x gross, "
f"ROI {(claimed_benefit - license_build) / license_build:.0%}")
step1 = claimed_benefit * 0.70 # 70% of users adopt, not 100%
step2 = 450_000 * 0.70 * (0.5 + 1.0) # year 1 runs at half rate while people ramp up
costs = license_build + 90_000 + 30_000 + 2 * 40_000 # + integration, training, 2 years of run cost
print(f"After adoption 70%: benefit ${step1:,.0f} -> {step1 / license_build:.2f}x on $300,000")
print(f"After year-1 ramp (50%): benefit ${step2:,.0f} -> {step2 / license_build:.2f}x on $300,000")
print(f"After all costs: benefit ${step2:,.0f} vs cost ${costs:,.0f} -> {step2 / costs:.3f}x, net ${step2 - costs:,.0f}")
flows = [-(license_build + 90_000 + 30_000), 450_000 * 0.70 * 0.5 - 40_000, 450_000 * 0.70 - 40_000]
print("Corrected flows:", [f"{f:,.0f}" for f in flows], f"NPV at 10%: ${npv(0.10, flows):,.0f}")
# The 400% year-one ROI claim
cost, run_rate = 120_000, 600_000
print(f"Claim: ROI {(run_rate - cost) / cost:.0%} (benefit {run_rate / cost:.0f}x cost)")
realised = run_rate / 12 * 6 # go-live at month 6 leaves 6 months of benefit
print(f"Only 6 live months: ROI {(realised - cost) / cost:.0%}")
cost_full = cost + 60_000 # add year-one run and maintenance
print(f"Plus $60,000 run cost: ROI {(realised - cost_full) / cost_full:.1%}")
Output:
Claim: $900,000 back on $300,000 = 3.00x gross, ROI 200%
After adoption 70%: benefit $630,000 -> 2.10x on $300,000
After year-1 ramp (50%): benefit $472,500 -> 1.57x on $300,000
After all costs: benefit $472,500 vs cost $500,000 -> 0.945x, net $-27,500
Corrected flows: ['-420,000', '117,500', '275,000'] NPV at 10%: $-85,909
Claim: ROI 400% (benefit 5x cost)
Only 6 live months: ROI 150%
Plus $60,000 run cost: ROI 66.7%
Reading it: adoption removes 30% of the benefit (3.00x to 2.10x), the ramp removes a further quarter of what was left (to exactly 1.575x on the $300,000; the output shows 1.57x because a computer stores 1.575 as slightly less than 1.575 and rounds down), and the missing costs push total cost to $500,000 so the multiple falls to 0.945x, with a net of -$27,500. If the benefit then persists, year three adds $315,000 of benefit less $40,000 of run cost, so the cumulative net after three years would be -$27,500 + $275,000 = $247,500. The deal may still be good; the honest statement is "a two-year payback does not exist, a three-year one does".
A second, separate example: a 400% year-one ROI in a business-intelligence (BI) business case
(Unrelated to the $300,000 case above.) ROI of 400% means benefit of 5x cost within the year. With a $120,000 cost and a $600,000 yearly benefit, the arithmetic does hold, so check the timing. If the dashboard goes live in month 6, only $300,000 of benefit falls in year one, and ROI is 150%. Add $60,000 of year-one run and maintenance and it is 66.7%. The usual error is an annualised run rate (a full year's worth at today's rate) presented as year-one realised benefit.
Correcting someone else's case while keeping credibility
- Go to the author privately first, as a reviewer asking questions, not as an auditor with a verdict ("Can you walk me through where 100% adoption comes from?").
- Keep their structure and their number. Show mine as a bridge, one step at a time, with the source for each change. Their original becomes the upside case.
- Agree on who owns each assumption, and re-issue the case as a new version with a change log, so nobody was "wrong"; the case simply got better evidence.
- When the customer is the author, say what the numbers still support: "The case works from year three. Here is what would bring payback forward."
What is break-even for a proposal, and how would you work it out? Use the choice between a per-user licence and a flat enterprise fee as your example.
Sample Answer
Direct answer
Break-even is the level of a driver at which two options cost the same, or at which an investment's NPV (net present value: the value today of all its cash flows, shrunk by a discount rate, the yearly return used to turn later money into today's money) is exactly zero. The user-count example below uses the first meaning (cost equality); a short example of the second (NPV zero) follows it. For a per-user licence of $18 per user per month against a flat enterprise fee of $120,000 a year, break-even is $120,000 / $216 = 555.6 users: below 556 users per-user pricing is cheaper, from 556 up the flat fee is cheaper.
How to work it out
- Separate fixed from variable cost. The flat fee is fixed: $120,000 a year whatever the user count. The per-user licence is variable: $18 x 12 = $216 per user per year.
- Set the two costs equal and solve for users: 216 x users = 120,000.
- Round in the direction that matters: at 555 users, per-user costs $119,880 (cheaper); at 556 it costs $120,096 (flat is cheaper).
- Draw it (the break-even chart): users on the horizontal axis, annual cost on the vertical. The flat fee is a horizontal line at $120,000, the per-user cost a line rising from zero at $216 per user, and they cross at 556 users. Everything left of the crossing favours per-user, everything right favours flat.
- Overlay a sensitivity: add a second per-user line for a negotiated $15 a month ($180 a year). It crosses the flat line at 666.7 users, so a discount on the per-user price pushes the break-even out by about 111 users.
def npv(rate, flows):
return sum(f / (1 + rate) ** t for t, f in enumerate(flows))
FLAT = 120_000 # enterprise fee, $ per year, any number of users
PER_USER_YEAR = 18 * 12 # $18 per user per month = $216 per user per year
breakeven = FLAT / PER_USER_YEAR
print(f"Break-even: {breakeven:.1f} users -> flat fee cheaper from {int(breakeven) + 1} users")
print(f"At $15 per user per month: {FLAT / (15 * 12):.1f} users")
print("users per-user@$18 per-user@$15 flat")
for u in range(400, 1000, 100):
print(f"{u:>5} {u * 216:>13,} {u * 180:>13,} {FLAT:>9,}")
users_by_year = [500, 650, 800]
extra = [u * PER_USER_YEAR - FLAT for u in users_by_year] # per-user cost minus flat fee
print("Per-user minus flat, years 1-3:", extra, "total", sum(extra))
print(f"NPV at 10% of choosing flat over per-user: ${npv(0.10, [0] + extra):,.0f}")
base_users = 700
saving = base_users * PER_USER_YEAR - FLAT
print(f"At {base_users} users the flat fee saves ${saving:,}/yr; 3-year NPV at 10%: ${npv(0.10, [0] + [saving] * 3):,.0f}")
print(f"Flips if users fall {1 - breakeven / base_users:.1%} (to {breakeven:.0f}) "
f"or the per-user price falls to ${FLAT / (base_users * 12):.2f}/month")
Output:
Break-even: 555.6 users -> flat fee cheaper from 556 users
At $15 per user per month: 666.7 users
users per-user@$18 per-user@$15 flat
400 86,400 72,000 120,000
500 108,000 90,000 120,000
600 129,600 108,000 120,000
700 151,200 126,000 120,000
800 172,800 144,000 120,000
900 194,400 162,000 120,000
Per-user minus flat, years 1-3: [-12000, 20400, 52800] total 61200
NPV at 10% of choosing flat over per-user: $45,620
At 700 users the flat fee saves $31,200/yr; 3-year NPV at 10%: $77,590
Flips if users fall 20.6% (to 556) or the per-user price falls to $14.29/month
Using it for a decision
A plan with 500, 650 and 800 users in years 1 to 3 shows the flat fee costs more in year 1 (per-user is cheaper by $12,000), then wins by $20,400 in year 2 and $52,800 in year 3: $61,200 over three years, or $45,620 in NPV at 10%. The crossing happens during year 2, so the decision rests on the growth forecast, not on the year-one cost. The year-0 flow in the code is 0 because both fees are paid at year end in this simplification, so nothing differs at the start. Traced by hand: -12,000/1.1 + 20,400/1.21 + 52,800/1.331 = -10,909 + 16,860 + 39,669 = 45,620.
The second meaning of break-even, in a short example: a $300,000 project that saves the same amount X a year for 3 years at 10% has NPV zero when X x 2.4869 = $300,000, so X = $120,632 a year (2.4869 is 1/1.1 + 1/1.1^2 + 1/1.1^3). Any saving above that makes NPV positive.
How far an assumption can move before the answer flips
At 700 users the flat fee saves $31,200 a year ($151,200 minus $120,000), a 3-year NPV of $77,590 at 10%. It flips (flat fee no longer cheaper) if users fall 20.6% to 556, or if the negotiated per-user price falls 20.6% to $14.29 a month. The two margins are equal because both enter the cost in the same multiplication. A margin of safety (how far an assumption can move before the decision flips) of about 21% is comfortable, but it tells you which two assumptions to protect: the seat forecast (the expected number of users) and the per-user quote.
Trade-offs and pitfalls
- The flat fee does not shrink if you shrink: after a layoff or a failed rollout you still pay $120,000. Per-user cost flexes down.
- Check the price lock (a contract term fixing the price for several years): a flat fee that jumps 15% at renewal moves the break-even. Check what counts as a "user" (a seat is one licensed account, whether assigned or actually active) and whether there are overage charges (extra fees if usage goes past the plan).
- Recommendation: choose the flat fee if you are confident of at least about 600 users on a sustained basis (a buffer above 556), take a multi-year price lock, and avoid it if headcount is uncertain or the user count could fall. If the supplier offers $15 per user per month, the break-even moves to 667 users, and for the 500, 650 and 800 user plan per-user pricing totals $351,000 against $360,000, so per-user wins.
A sponsor's draft business case for replacing a legacy application with a SaaS platform counts only the vendor's fees. What costs is it likely missing, and how would you estimate each one credibly?
Sample Answer
Direct answer
A draft that counts only vendor fees is missing most of the cost of change: the migration, integration, training and change management, internal staff time, parallel running and decommissioning of the legacy system, contract exit and escalation terms, and probability-weighted risk. These can be as large as the fees; in the illustrative example below they roughly double the three-year total. I estimate each from a source other than the sponsor's optimism, state a confidence level, and show both views: fees-only and all-in.
Likely missing costs and how I estimate each credibly
| Missing cost | How to estimate | Confidence |
|---|---|---|
| Data migration (extract, clean, transfer, testing, cutover (the switch from old to new), rollback rehearsal (practising how to return to the old system if the switch fails), downtime) | Count objects and records; ask the vendor for rates and a reference customer of similar size; engineering estimate for cleaning; add a contingency | Medium |
| Integration with surrounding systems (identity, finance, reporting) | List each interface, estimate by the owning team; vendor connector price if one exists | Medium |
| Training and change management | Users x hours x loaded hourly rate (salary plus benefits and overhead, divided by working hours), plus champions and materials; amortise across business units by headcount | Medium |
| Parallel run (paying for both systems during transition) | Legacy run cost per month x overlap months | High |
| Internal staff and shared overhead (admin, security review, project office, cross-team FTE, meaning full-time-equivalent staff time) | FTE share x loaded annual rate (the full employer cost of a person, not just salary), using finance's allocation method; not zero because they are salaried | Medium to low |
| Contract exit terms (termination fee, data export, SLA credits, price-escalation cap, a limit on how much the vendor can raise fees at renewal) | Read the paper: a service-level agreement (SLA) credit is a remedy, not a benefit to count on | High |
| Decommissioning the legacy system (shutting it down safely: archiving, licences ended, hardware) | Quote from the infrastructure team | Medium |
| Risk, probability-weighted (probability x impact, so a likely small risk and an unlikely large one can be compared) | Probability x impact per risk, with the probability taken from a named source such as the audit team's past findings, the vendor's reference customers, or an industry incident rate | Low |
Worked example (illustrative figures, 3 years)
The sponsor's draft: vendor fees of $240,000 a year, so $720,000. Contract allows a 5% annual price rise: $240,000 + $252,000 + $264,600 = $756,600.
| Item | 3-year cost |
|---|---|
| Fees with 5% uplift | $756,600 |
| Migration (data transfer, testing, cutover, rollback) | $180,000 |
| Integration | $90,000 |
| Training and change management | $70,000 |
| Parallel run (2 months of legacy at $30,000) | $60,000 |
| Internal staff: 0.5 FTE x $150,000 x 3 years | $225,000 |
| Legacy decommission and exit | $25,000 |
| Expected risk cost (probability x impact, derived below: 0.40 x $50,000) | $20,000 |
| All-in | $1,426,600 |
All-in is about 1.98 times the fees-only draft of $720,000. This is a gross cost; for a fair decision I would put it beside the legacy system's avoided run cost, otherwise the comparison is one-sided.
Regulated customer data moving to the cloud: the expected-cost buckets are audit and certification work, legal and data-processing agreements, regional performance or residency fixes, and breach exposure. Example: a 40% chance the audit requires $50,000 of rework gives an expected $20,000 (0.40 x $50,000), the figure above. Show it separately, labelled expected value (the probability-weighted average cost), and put the unweighted worst case in the risk section.
Validate each estimate
Walk the table with the people who own each line (IT, security, the business units, finance, procurement, the buying and contracts team), record who gave the number and a confidence level, and ask the vendor to confirm the migration assumptions in writing. Estimates from the sponsor alone count as low confidence.
Presenting differently
- To procurement: negotiating levers: price-escalation cap, exit fees, SLA credits, migration support in the contract.
- To the executive sponsor: the all-in total, the range, the three largest line items, and the decision gate.
A related use, if you work on the vendor side: the same table shows what a customer faces in switching away from a product (implementation, training, integration, downtime and the opportunity cost of delayed projects). A seller can cite those costs as a retention argument, but they must be shown honestly rather than hidden, because a buyer who finds them later stops trusting the whole case.
Pitfalls
Treating "we already employ those people" as zero cost, counting SLA credits as savings, and presenting one total with no confidence levels.
A programme reports large savings, but teams say service is slipping. How do you tell gross savings from net, and one-time savings from sustainable ones?
Sample Answer
Direct answer
Treat the reported saving as a gross claim and rebuild it as a bridge: gross reported saving, minus what it cost to achieve, minus what it broke, minus what merely moved somewhere else, equals net realised saving. Then split the net into one-time items and a sustainable run-rate (a saving that recurs every year at the same level). Finally, put the service signal next to the money: if teams say service is slipping, part of the "saving" is a cost that has been pushed onto customers or onto the teams, and it will come back as rework, credits or attrition.
The bridge for the example below, in $k:
| Line | $k |
|---|---|
| Gross reported saving | 5,000 |
| Less one-off programme cost | -1,200 |
| Less remediation (SLA credits, rework) | -600 |
| Less leakage (moved elsewhere) | -700 |
| Year-1 net realised saving | 2,500 |
Step by step
- Fix the baseline. A saving is measured against what spending would have been if the programme had not happened (adjusted for volume: if the business now handles 20% more work, last year's $1.0M cloud bill would have been about $1.2M, so that is the baseline, and a $0.9M bill saves $0.3M, not $0.1M), not against last year's bill and not against a budget that was padded.
- Gross to net. Subtract: implementation and programme costs; temporary spend (parallel running, meaning paying for the old and new systems at once during cutover; consultants; contractors); remediation (fixing what the change broke: rework, SLA credits, incident cost, overtime); leakage (spend that reappears in another cost centre, meaning another team's budget line, as a contractor backfill, meaning contractors hired to do the cut roles' work, a shadow tool, meaning software a team buys itself outside the official budget, or cloud cost moved to another team's account). Example of leakage: the platform team's cloud bill falls from $400k to $250k, but another team's tagged bill rises from $100k to $220k because the workload moved, so the real saving is $30k, not $150k.
- One-time versus sustainable. One-time: vendor credits, a contract concession that does not renew, a hiring freeze that will lift, spend deferred to next year (cost postponed, not removed), asset sales. Sustainable: lower unit costs, retired systems, permanently removed work. They are one-time because they happen once or merely move cost to later: a credit is a single refund, deferred maintenance (upkeep skipped now) must be paid eventually, and a freeze ends. Test: "will this appear again next year with no further action?"
- Check service against savings. Put service-level attainment (SLA: the service-level agreement, the promised response or uptime level), ticket backlog, rework rate and attrition beside the savings line. Where service fell, cost the damage.
- Attribute to teams. For platform modernisation savings, tag every resource by team and programme, fix a baseline (usage in the twelve months before migration, normalised for volume, meaning scaled to today's workload), and compute saving as baseline-adjusted-for-volume minus current. For example, a team ran 100 units of work on a $50k monthly bill before migration ($500 per unit); it now runs 120 units on a $54k bill. The adjusted baseline is 120 x $500 = $60k, so the saving is $6k a month, where a raw comparison of $50k with $54k would wrongly show a $4k increase. Without tagging, savings get claimed by everyone or no one.
Headcount cuts that create capacity problems
If the saving came from cutting roles (watch attrition, meaning staff leaving, as well as formal cuts), ask what work the roles did. Three signs of a false saving: overtime or contractor spend rising in the same function, ticket backlog or cycle time worsening, and a rising share of senior time spent on tasks juniors used to do. Those costs belong in the remediation and leakage lines, not in someone else's budget.
Worked example (illustrative, code run)
A programme reports $5.0M of savings. In this example the leakage includes contractor backfill of cut roles and cloud cost that moved to another team's account; the remediation includes SLA credits and rework.
# $k, illustrative programme report for year 1
gross_reported = 5000
one_time_inside_gross = {
"one-off vendor credit": 500,
"deferred maintenance (cost postponed, not removed)": 600,
}
programme_cost = 1200 # one-off: spent once, not every year
remediation = 600 # SLA credits ($250k) and rework and overtime ($350k) (assumed to recur while the cause remains)
leakage = 700 # contractor backfill, cloud cost moved to another team (recurs)
year1_net = gross_reported - programme_cost - remediation - leakage
one_time = sum(one_time_inside_gross.values())
run_rate = gross_reported - one_time - remediation - leakage
print("Year-1 net realised saving $k:", year1_net)
print("One-time items inside the headline $k:", one_time)
print("Sustainable run-rate saving $k per year:", run_rate)
print("Run-rate as share of headline:", round(100 * run_rate / gross_reported), "%")
# Reconcile: year-1 net = run-rate + one-time items - one-off programme cost
print("Check:", run_rate + one_time - programme_cost == year1_net)
Output:
Year-1 net realised saving $k: 2500
One-time items inside the headline $k: 1100
Sustainable run-rate saving $k per year: 2600
Run-rate as share of headline: 52 %
Check: True
Two different numbers come out, and they answer different questions. The year-one net is $2.5M ($5.0M less $1.2M programme cost, $0.6M remediation and $0.7M leakage). The sustainable run-rate is $2.6M a year: the headline less the $1.1M of one-time items ($0.5M vendor credit plus $0.6M of deferred maintenance that will come back), less the $0.6M and $0.7M that keep recurring while the causes remain. That is 52% of the headline. The reconciliation line shows the two tie together: run-rate $2.6M plus one-time $1.1M minus one-off programme cost $1.2M equals the $2.5M. The number to defend in year two is the $2.6M, and it can rise toward $3.9M (the headline less the one-time items) only if the remediation cause is fixed and the leakage is closed, which is the follow-up action.
Trade-offs and pitfalls
- Over-netting is also an error: do not subtract a one-off programme cost from every future year of a recurring saving. Show both the first-year net and the run-rate, as above.
- A one-time saving is not worthless (it may fund the programme), but it must not be annualised.
- Leakage is hard to see without a cost-centre-level view; ask finance for a before-and-after spend by category for the affected teams, not only the programme's own report.
- What would change the call: if service metrics are flat and the recurring items are validated by finance, the gap between gross and net shrinks and I would accept more of the headline.
Unlock Full Question Bank
Get access to all Business Case Development and ROI Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.