Cloud Cost Optimization and FinOps Questions
Controlling and optimizing cloud spend: cost modeling and forecasting, rightsizing, reserved capacity and savings plans, autoscaling for cost, tagging and chargeback, and the FinOps operating model. Covers building the business justification for infrastructure spend and continuously driving efficiency at scale without sacrificing reliability. Cost as a first-class architectural concern.
A team is considering a 3-year commitment for compute and database reservations. How would you model and hedge the risk that the workload shrinks, the vendor changes pricing, or the underlying technology becomes obsolete before the term is up? What contract structures or operational strategies would reduce that downside?
Sample Answer
Direct answer
I'd quantify the exposure as a probability-weighted range of outcomes rather than a single number, then hedge it with a mix of contract-level flexibility and operational elasticity, because the uncomfortable fact under all of this is that a reservation hedges price risk (you lock in a discount) but does very little to hedge demand risk (you're still on the hook for the committed dollars if the workload shrinks). The compute and database halves of a "compute and database" commitment are not equally hedgeable either, which matters for how you structure the deal.
Structured elaboration
Modeling the risk
- Build a 3-year scenario model: base case, and workload-shrink cases (for example -30%, -50%), each with a probability.
- Layer in vendor price-change scenarios and a technology-obsolescence case (cost of migrating off before the term ends).
- Key inputs to pin: baseline steady-state usage, the discount rate the commitment buys, the migration/obsolescence cost, and your shrink-probability estimates. Use a full Monte Carlo simulation only if the number of interacting variables genuinely justifies it; for most commitment decisions a three- or four-scenario expected-value model, shown below, is transparent enough to defend in a budget review and doesn't hide its assumptions inside a simulation nobody can re-derive by hand.
Contractual hedges, and where compute and database reservations diverge
This is the part worth being precise about, because the two halves of the commitment behave differently:
- Compute: AWS EC2 offers two reservation shapes. A Standard Reserved Instance (RI) is fixed. A Convertible RI can be exchanged for a different instance family, operating system, or tenancy, but AWS requires the new configuration's value to be equal to or greater than the remaining value of the original, so an exchange can reshape the commitment, it cannot shrink the dollar amount. Only Standard RIs, not Convertible RIs, can be resold on the EC2 Reserved Instance Marketplace, and only after being active at least 30 days, capped at $50,000 and 5,000 instances over the lifetime of the account, with AWS taking a 12% fee on the sale price. AWS Savings Plans (SP) are more flexible day to day, a Compute Savings Plan applies regardless of instance family, size, operating system, or region, but they have no general early-exit path: AWS documents only a narrow return window (commitments of $100/hour or less, purchased in the past 7 days, same calendar month), not an ongoing cancellation or resale mechanism.
- Database: Amazon RDS Reserved (database) Instances are structurally less hedgeable than either compute option. They cannot be cancelled, full stop, you're billed for the committed term whether you use the capacity or not. They cannot be resold on any marketplace (RDS reservations are explicitly excluded from the EC2 Reserved Instance Marketplace). The only flexibility is size changes within the same instance class type, same region, and same database engine. Practically: if a "compute and database" 3-year deal is negotiated as one symmetric package, the database portion is the part that actually can't flex if the workload shrinks, and that needs to be sized more conservatively than the compute portion, not identically to it.
- Azure Reservations, for comparison, currently allow exchanging within the same product family (compute-for-compute, SQL-for-SQL) as long as the new reservation's value is equal to or greater than the remaining commitment, and allow outright cancellation/refund up to $50,000 per rolling 12-month window per billing profile, with no early-termination fee charged today (Microsoft's own documentation flags that a fee may be introduced later). One caveat worth flagging as time-sensitive rather than permanent: Azure's compute-reservation instance/region exchange flexibility is in a documented wind-down "grace period" in favor of Azure Savings Plan for compute, so the specific exchange terms should be re-checked against current Microsoft documentation before being relied on in a contract negotiated today.
- Google Cloud offers two committed-use discount (CUD) shapes: spend-based CUDs, which apply across eligible usage in any project linked to the billing account, and resource-based CUDs, tied to a specific region and project (with terms up to six years for Compute Engine). I did not find a documented early-cancellation or exchange path for either GCP CUD type in Google's own documentation, so I'm not asserting one either way here; verify current GCP terms directly before relying on cancellability as a hedge.
Operational hedges
- Architect for elasticity: containerization, autoscaling, or serverless where it fits, to shrink the footprint that actually needs a commitment.
- Right-sizing cadence: automated telemetry plus quarterly review to reduce reserved-but-unused capacity.
- Keep the reserved floor genuinely conservative, cover only the predictable steady-state baseline and put variable load on-demand or spot.
- A multi-cloud or portability fallback is a real hedge against obsolescence and lock-in, but it's expensive to maintain continuously; reserve it for workloads where lock-in risk is a strategic, not incidental, concern.
Financial tactics and governance
- Reserve only the predictable floor (illustrated below), not the full observed peak.
- Use internal chargeback or showback to keep utilization visible and catch drift early.
- Run real-time dashboards against the committed floor, with alerts when utilization falls below a set threshold, and revisit the model annually against actuals.
Worked example
Pinned inputs: the team's assessed steady-state floor for compute and database combined is $40,000/month at on-demand-equivalent pricing. They commit that floor via a blended 3-year deal (Convertible RI for compute, RDS Reserved Instance for database) at a 35% blended discount, billed at $26,000/month for 36 months regardless of actual usage, a fixed 3-year bill of $936,000.
3-year committed bill: C=26,000×36=$936,000
Three scenarios, each producing the 3-year on-demand-equivalent value of what was actually needed, compared against the fixed $936,000 bill:
| Scenario | Probability | Need/month | 3-yr on-demand value | Billed | Net vs. billed |
|---|---|---|---|---|---|
| Base (flat) | 0.5 | $40,000 | $1,440,000 | $936,000 | +$504,000 |
| Shrink -30% | 0.3 | $28,000 | $1,008,000 | $936,000 | +$72,000 |
| Shrink -50% | 0.2 | $20,000 | $720,000 | $936,000 | -$216,000 |
EV=0.5(1,440,000−936,000)+0.3(1,008,000−936,000)+0.2(720,000−936,000)
EV=0.5(504,000)+0.3(72,000)+0.2(−216,000)=252,000+21,600−43,200=$230,400
Reading this: the commitment is expected-value positive ($230,400 over 3 years) even accounting for real shrink probability, but the downside tail (the -50% case, at 20% probability) produces a concrete $216,000 loss versus a no-commitment counterfactual, because the fixed bill doesn't shrink with usage. That tail is exactly what the reservation doesn't hedge, and it's the number to bring to a risk conversation, not the expected value alone. Because the database half of this commitment can't be resold or cancelled at all, the -50% tail is a real, uncushioned exposure specifically on the database portion; the compute portion at least has a Convertible RI exchange or marketplace-resale path (for Standard RIs) to partially recover value.
Trade-offs and pitfalls
- The core conceptual pitfall: teams model reservation risk as if it were symmetric with the discount, "we get 35% off, worst case we're even," when in fact the commitment is a fixed bill and the downside is real dollars, not just a foregone discount.
- Treating "the reservation" as one homogeneous instrument when compute and database reservations have materially different exit paths is the specific trap this question is testing for; size the database portion more conservatively than the compute portion for exactly this reason.
- Convertible RI exchanges reshape, they don't shrink, the commitment; relying on "we can always exchange it down" without checking the equal-or-greater-value rule is a common and avoidable mistake.
- Marketplace resale is a real but narrow safety valve: it exists only for EC2 Standard RIs, is capped, and costs a 12% fee, it is not a general escape hatch for a database commitment or for a Convertible RI.
- Vendor flexibility terms are not permanent contract features, they're current policy that vendors change (the Azure exchange wind-down cited above is a live example); re-verify the specific mechanism against current vendor documentation before signing, not against what was true when the deal was last negotiated.
How would you forecast and size reserved capacity or savings plans for a workload with seasonal peaks? What inputs would you need, how would you build in a margin for under- or over-commitment, and how would you present a conservative option versus an aggressive one to finance?
Sample Answer
Direct answer
Size the commitment against a demand curve across the year, not a single number: forecast the full range from trough to peak, then choose what percentile of that curve to commit against. Committing near the P50 (the level demand is at or above about half the time) is conservative and safe but leaves savings on the table in the low-demand months, while committing near P80-P90 captures more savings but raises the odds of paying for capacity you don't use in the trough. Present both to finance as a genuine trade-off, not just an upside number, and for a fast-growing or uncertain business, weigh term length (1-year vs. 3-year) at least as heavily as the percentile, since breakage risk from locking into a forecast that turns out wrong usually costs more than the extra discount from a longer term is worth.
Structured elaboration
Required inputs
- At least 12-24 months of historical instance-hour usage, ideally annotated with known seasonality drivers (a marketing campaign, end-of-quarter usage, a seasonal sales event).
- A business growth forecast (expected growth rate, planned migrations or new features that would shift the baseline).
- Current on-demand and spot usage patterns, so the forecast isn't built purely from historical reserved usage.
- Financial constraints: maximum budget, and how much risk of an unused commitment the business is actually willing to carry.
Building in a margin for under- or over-commitment
- Stagger commitments in smaller tranches (quarterly or monthly phasing) rather than one large annual purchase, so a wrong forecast is a smaller mistake.
- Keep a buffer of on-demand or spot capacity sized to cover the gap between the committed level and the historical peak, so peak demand isn't dependent on the commitment alone.
- Prefer convertible or flexible commitments over rigid ones where the discount difference is small, since flexibility is itself a form of margin.
- Reassess on a fixed cadence (every one to two quarters) rather than locking in a forecast and revisiting only at renewal.
Presenting conservative versus aggressive to finance
| Conservative (commit near P50) | Aggressive (commit near P80-P90) | |
|---|---|---|
| Expected savings | Lower | Higher |
| Risk in trough months | Low; commitment rarely exceeds actual need | Higher; commitment can exceed actual need, meaning you pay for unused capacity |
| Best suited to | Uncertain or fast-changing workloads, early-stage forecasting | Well-understood, historically consistent seasonal patterns |
Always show finance the downside explicitly, not just the projected savings: what does the aggressive option cost in the single worst (lowest-demand) month, compared to having made no commitment at all that month. That's usually a more persuasive number for a risk-averse finance stakeholder than an annualized savings percentage.
Term length versus growth uncertainty
A 3-year term typically carries a deeper discount than 1-year, but for a company expecting to grow quickly or change its infrastructure shape, that's often the wrong trade: breakage risk (being locked into a specific instance family, region, or size the growth trajectory outgrows) erodes the discount faster than the discount itself is worth. In that situation, favor 1-year or convertible commitments even at a smaller headline discount, and resize every couple of quarters rather than committing three years out against a forecast likely to be wrong well before the term ends. This is also a cash-flow question, not just a discount question: an all-upfront 3-year payment ties up cash a fast-growing, still-cash-constrained company may need for hiring or infrastructure elsewhere, so a smaller upfront (or no-upfront, amortized monthly) 1-year commitment can be the right call even when the 3-year option is cheaper on paper.
Worked example
Suppose historical analysis shows monthly demand ranging from a trough of 4,000 instance-hours in the low season to a peak of 10,000 in the high season, with a P50 "typical month" of 6,000 hours. On-demand costs $0.10/hour; a 1-year reserved commitment amortizes to $0.065/hour (a 35% discount).
Conservative option, commit at P50 (6,000 hrs/month):
Trough month, the committed hours cost the same whether used or not, compared against what pure on-demand would have cost for the 4,000 hours actually needed:
Committed:6,000×$0.065=$390vs.On-demand-only:4,000×$0.10=$400
Even with 2,000 hours of committed capacity going unused that month, the commitment is still slightly cheaper than not having committed at all.
Peak month, the committed 6,000 hours plus 4,000 hours of on-demand for the remainder, compared against full on-demand:
Mix:(6,000×$0.065)+(4,000×$0.10)=$390+$400=$790vs.Full on-demand:10,000×$0.10=$1,000
A $210 saving in the peak month.
Aggressive option, commit at P90 (9,000 hrs/month):
Trough month:
Committed:9,000×$0.065=$585vs.On-demand-only:4,000×$0.10=$400
This is the concrete downside: in the trough month, the aggressive commitment costs $185 more than making no commitment at all.
Peak month:
Mix:(9,000×$0.065)+(1,000×$0.10)=$585+$100=$685vs.Full on-demand:10,000×$0.10=$1,000
A $315 saving in the peak month.
The conservative option never costs more than doing nothing, in any month; the aggressive option saves more in the peak month but genuinely costs more than doing nothing in the trough month. That's the exact number to put in front of finance: not "aggressive saves more on average" but "aggressive costs $185 more than no commitment at all in our worst month."
Trade-offs and pitfalls
- Showing only the annualized savings number hides the specific downside month. Finance needs to see what the aggressive option costs in the worst month, not just the average across the year.
- A longer term amplifies both the discount and the breakage risk together, so match term length to forecast confidence, not just to whichever term has the deepest headline discount.
- Ignoring correlated risk across workloads understates the true worst case. If multiple teams' demand curves are all tied to the same seasonal driver (a shared marketing calendar, a shared fiscal quarter-end), you can't diversify that risk away by spreading the commitment across teams.
Explain the difference between showback and chargeback as cloud cost allocation models. What operational and behavioral impacts does each have on engineering teams, and in what situation would you recommend one over the other?
Sample Answer
Direct answer
Showback reports each team's cloud costs for visibility without moving any money: nobody's budget is actually debited. Chargeback goes further and allocates real costs to a team's budget, typically through an internal invoice or a direct debit against their cost center. The mechanics of allocation (tagging, cost pools) are identical between the two models; what differs is whether the number is informational or binding, and that single difference changes team behavior more than almost any other FinOps decision.
Structured elaboration
Operational requirements
- Showback needs accurate tagging and a reporting pipeline (dashboards built on the provider's billing export), but no accounting integration. It is comparatively cheap to stand up.
- Chargeback needs everything showback needs, plus allocation rules for shared and hard-to-attribute costs (a shared database, a platform team's infrastructure), an internal billing or budget-debit mechanism, and usually a dispute process for when a team contests its bill. It is meaningfully more operational overhead.
Behavioral impacts
- Showback creates awareness but relies on a team choosing to act on it. It works well when the goal is building cost literacy and trust in the data, and it fails quietly: a team can see an inflated bill for months and simply not prioritize fixing it, because nothing forces the issue.
- Chargeback creates direct, budget-line accountability, which reliably produces the fastest optimization response. It also produces predictable second-order effects: teams start negotiating over shared-cost allocation formulas, and some teams under-provision or avoid experimentation because the cost is now visibly theirs. Badly designed chargeback (especially unfair shared-cost splits) actively damages trust in the whole program.
When to recommend which
- Recommend showback when tagging discipline and cost data are still immature, when the organization is early in FinOps adoption and needs cultural buy-in before it can survive a contentious billing dispute, or when the goal this quarter is visibility, not enforcement.
- Recommend chargeback once allocation is trustworthy, budget owners are clearly defined, and leadership needs teams to make trade-offs against a real budget constraint (a business unit that must self-fund its cloud spend, for example).
- In practice the strongest programs run a hybrid: chargeback for costs that are cleanly attributable to a single team (dedicated compute, a service's own database), and showback for genuinely shared infrastructure (a shared Kubernetes cluster, a platform team's networking spend) where a clean per-team split would be arbitrary and would just generate disputes instead of better decisions. This avoids forcing a false precision onto costs that are structurally shared.
Worked example
A platform team's shared cluster costs $40,000 a month and hosts workloads for three product teams, roughly split 50/30/20 by measured resource requests. Under showback, all three teams see "$20,000 / $12,000 / $8,000, informational" on a dashboard, and it is up to each team whether to act on their share. Under chargeback, those same three figures are debited from each team's budget as an internal invoice line, and a team now has to justify that $20,000 (or reduce it) the same way it justifies any other budget line. A hybrid design would chargeback the dedicated services each team also runs outside the shared cluster (fully attributable, no allocation dispute possible) while keeping the shared cluster on showback, because a resource-request-based 50/30/20 split is an estimate, not a precise cost, and billing teams against an estimate they can contest is a common source of program-trust failure.
Trade-offs and pitfalls
The biggest pitfall is skipping straight to chargeback before tagging and allocation are trustworthy: teams will contest a bill they believe is wrong, and if the underlying data really is wrong, the program loses credibility fast and is hard to recover. A second pitfall is chargeback without a clear owner for genuinely shared costs, which pushes teams toward proportional formulas nobody fully agrees with and creates ongoing friction that has nothing to do with actual waste. A third, subtler failure is showback with no organizational follow-through: if visibility never translates into any consequence, teams learn to ignore the dashboard, and the "awareness" goal quietly fails too.
A legacy on-premises application currently costs $600,000 a year to run. Moving it to managed cloud services is estimated at $200,000 a year with a one-time $150,000 migration fee. Build the 3-year TCO, ROI, and payback period for this migration, and call out the assumptions and sensitivities you'd want to flag to stakeholders before they sign off.
Sample Answer
Direct answer
Over three years this migration looks strong on paper: about $1.05 million in net savings, a 140% return on the cloud investment, and a payback period of roughly 4.5 months on the one-time migration fee. The number that actually matters for sign-off isn't the base case though, it's how fast that case degrades under a higher-than-planned cloud bill or a bigger-than-planned migration effort, since both are common ways this kind of estimate goes wrong in practice.
Structured elaboration
Building the comparison: the on-premises cost is a flat annual run-rate; the cloud cost is a one-time migration fee plus a lower annual run-rate. Three-year total cost of ownership (TCO) for each side:
TCOon-prem=600,000×3=1,800,000 TCOcloud=150,000+200,000×3=750,000Net benefit is the difference, and return on investment (ROI) expresses that benefit as a percentage of what was actually invested (the cloud spend, since that's the money being committed to get the savings):
ROI=750,0001,800,000−750,000=750,0001,050,000=1.40=140%Payback period asks a different question: how long until the migration fee is recovered from the ongoing run-rate saving alone (not the full three-year benefit)?
Payback=400,000150,000=0.375 years≈4.5 monthswhere $400,000 is the annual run-rate saving ($600,000 minus $200,000).
Assumptions to state explicitly before anyone signs off:
- The $600,000 on-prem figure is the fully-loaded cost (hardware refresh, facilities, and the operations labor to run it), not just the visible infrastructure line.
- The $200,000 cloud figure covers equivalent capacity, licensing, monitoring, and support at the same service level, not a narrower slice of what the on-prem number included.
- The $150,000 migration fee covers discovery, execution, testing, and cutover, with no material re-architecture beyond a lift-and-shift-plus-managed-services move.
- Usage and traffic stay roughly flat over the three years; this is a cost comparison, not a growth forecast.
Sensitivities to flag to stakeholders, ranked by how often they actually bite:
- Cloud run-rate coming in above plan. If actual managed-service cost lands at $240,000 a year (a 20% miss) instead of $200,000, the annual saving drops to $360,000. This is the single most common way these estimates go wrong, because early estimates rarely capture the full data-transfer and support-tier costs until the workload is actually running in production.
- Migration cost overrun. If the one-time fee comes in at $300,000 instead of $150,000 (a common outcome when discovery underestimates integration complexity), payback stretches to 9 months. Still fast, but worth stating as a range rather than a single number.
- Hidden costs not in either baseline: license portability terms, compliance or data-residency controls that require extra configuration, and the egress cost of anything that still needs to talk back to on-prem systems during a phased cutover.
- Time value of money. All the figures above are undiscounted. For a rigorous board-level comparison I'd also compute net present value (NPV) using the company's discount rate, since $400,000 saved in year three is worth less today than $400,000 saved in year one, and a purely undiscounted payback period can make a slow-starting case look better than it is.
Worked example
The same methodology extends directly to a narrower, more technical version of this question, and over a different time horizon: comparing a self-managed database against a managed equivalent over five years instead of three, for instance running PostgreSQL on owned hardware versus a managed offering like Amazon RDS (Relational Database Service) or Aurora (a managed, cloud-native relational database service). The mechanics are identical, a flat legacy run-rate against a lower managed run-rate plus a one-time migration effort, just with database-specific line items on each side and one more year of run-rate in the TCO sum: on-prem includes patching and backup labor and license costs, managed includes the service's own pricing tier plus a smaller migration effort (schema and data migration, connection cutover) instead of a full application re-platform. The same TCO, ROI, and payback formulas apply unchanged over five years; only the inputs and the time horizon differ.
When presenting either version of this case to a non-technical audience, lead with the plain-language headline (the number of months to break even, and the multi-year dollar total), show the sensitivity range as a small table rather than a wall of formulas, and hold the underlying spreadsheet in reserve for anyone who wants to check the math.
Trade-offs and pitfalls
- Reporting a single-point ROI without the sensitivity range is the most common way this kind of business case loses credibility later: if the actual cloud bill lands 20% high (a routine outcome, not an edge case), a board that was shown only the base case will remember the miss, not the caveat.
- Undiscounted payback is easy to compute and easy to explain, which is exactly why it's tempting to present as the whole story. It ignores the time value of money and can make a large, slow-arriving benefit look better than a smaller, faster one; pair it with NPV for anything above a routine sign-off.
- A common wrong turn is treating the migration fee as the only one-time cost. Parallel-running both environments during cutover, temporary double licensing, and staff retraining are real one-time costs that belong in the migration-cost line, not left as an unstated risk.
- Comparing "cloud run-rate" against "on-prem run-rate" without normalizing for what's actually included on each side (does on-prem's number include the ops labor? does cloud's number include support?) is the single easiest way to make either side look artificially better than it is.
You need to compare the cost of running the same workload on two different cloud providers. What would you actually measure to make that an apples-to-apples comparison rather than just comparing list prices, and how would you account for the hidden costs of running multi-cloud at all?
Sample Answer
Direct answer
To make it apples-to-apples you must normalize workload, environment, and consumption before comparing bills, not compare list prices: same compute-equivalent capacity, same region cost tier, same commitment level (on-demand vs on-demand, reserved vs reserved), and the same service-level agreement (SLA), then express cost per unit of useful work (for example, cost per 1,000 successful requests at a fixed 95th-percentile (p95) latency target). Then you have to add back the "membership fee" of running multi-cloud at all: the cross-cloud egress, duplicated tooling, and extra engineering time that a single-provider comparison never has to pay, which routinely erases 10-20% or more of any per-unit savings.
Structured elaboration
1. What to normalize before comparing sticker price
- Compute-equivalent sizing: map each provider's instance tiers to comparable throughput, not to matching names.
- Same commitment tier: comparing provider A's on-demand price to provider B's multi-year reserved price is the single most common apples-to-oranges mistake in these comparisons.
- Same region cost tier and same redundancy posture (single availability zone vs multi-zone).
- Same SLA and support tier.
- Fully loaded egress and storage costs, not compute alone.
- Express the result as cost per unit of useful work at a fixed quality bar (cost per 1,000 successful requests at a fixed p95 latency target), so a cheaper-but-slower or cheaper-but-flakier option isn't silently favored.
2. The hidden costs of running multi-cloud at all (this is the part a list-price comparison misses entirely)
- Cross-cloud egress: moving data between the two providers for replication, backup, or a shared data layer is billed by both sides and is usually the largest hidden line item.
- Duplicated tooling and control planes: separate identity and access management (IAM) models, separate monitoring stacks or an abstraction layer over both, separate continuous integration/continuous delivery (CI/CD) targets.
- Lost volume discounts: splitting spend across two providers means neither crosses the threshold for the best committed-use tier, so the cheaper per-unit list price may not be the price you actually pay at your split volume.
- The engineering and on-call cost of maintaining two operational runbooks, two areas of deep expertise, and two incident-response paths.
- Two separate billing, tagging, and chargeback pipelines to reconcile for governance.
3. When multi-cloud is worth the hidden cost anyway
- Regulatory or data-residency requirements that no single provider satisfies alone.
- A genuine best-of-breed need, where one provider's managed service is meaningfully ahead and the capability gap outweighs the duplication tax.
- Negotiating leverage: a credible second provider improves your position in an enterprise discount program (EDP) renewal, even if you never route meaningful production traffic to it.
- Contrast this with the common, and usually weaker, justification: "avoid vendor lock-in" by itself rarely pays for the duplication tax on its own. Lock-in risk is real, but permanently carrying a 10-20% operational overhead to hedge a switching cost you may never pay is often worse than negotiating portability into how you build (standard formats, containerized workloads, avoiding the most proprietary managed services for your riskiest components) while staying single-cloud. Multi-cloud reduces total cost of ownership (TCO) only when the duplication tax is smaller than a concrete, already-materializing cost; it increases TCO whenever it is adopted as insurance against a hypothetical future lock-in cost rather than a priced one.
Worked example
Workload: a steady 15,000,000 requests/month, needing 4 vCPU / 8 GB of sustained capacity, generating 500 GB of egress/month.
Provider A: $0.68/hour for the equivalent instance, 730 hours/month, 100 GB free egress then $0.09/GB.
computeA=$0.68/hr×730hr=$496.40
egressA=(500−100)GB×$0.09/GB=$36.00
totalA=$496.40+$36.00=$532.40⇒$0.0355 per 1,000 requests
Provider B: $0.55/hour for the equivalent instance, 200 GB free egress then $0.12/GB.
totalB=$401.50+$36.00=$437.50⇒$0.0292 per 1,000 requests
| Provider A | Provider B | |
|---|---|---|
| Compute (730 hr) | $496.40 | $401.50 |
| Egress (500 GB, after free tier) | $36.00 | $36.00 |
| Total/month | $532.40 | $437.50 |
| Cost per 1,000 requests | $0.0355 | $0.0292 |
At list price alone, B looks about 18% cheaper per unit of work ($0.0292 vs $0.0355). But if this were a genuine multi-cloud deployment rather than a straight migrate-to-the-cheaper-one decision, add the duplication tax: roughly 50 GB/month of cross-cloud replication egress at $0.08/GB ($4/month), plus an amortized 0.1 full-time equivalent (FTE) of extra ops time to run both stacks at $150,000/year fully loaded (about $1,250/month). That is roughly $1,254/month of multi-cloud-specific cost against a base bill of $437-532/month, meaning the hidden cost is two to three times the entire compute-plus-egress bill it was supposed to optimize. That is why "which provider is cheaper for this workload" and "should we run both" are two separate calculations.
Trade-offs and pitfalls
- Comparing on-demand pricing to reserved pricing: always match commitment tier, or normalize both to a blended effective rate, before drawing a conclusion.
- Ignoring egress asymmetry: providers price egress very differently, and a client-facing, egress-heavy workload can flip the ranking entirely once egress is included.
- Chasing multi-cloud for lock-in avoidance without pricing the duplication tax: as shown above, the ops and cross-cloud egress cost is often the dominant term, not a rounding error.
- Benchmarking once instead of over a representative traffic window: load shape, time-of-day, and seasonal effects change the answer.
- Comparing raw infrastructure cost instead of cost per unit of useful work at a fixed SLA: an option that is 20% cheaper but only meets your latency target 90% of the time is not actually cheaper once the SLA risk is priced in.
Unlock Full Question Bank
Get access to all 18 Cloud Cost Optimization and FinOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.