Cloud Cost Optimization and FinOps Questions
Controlling and optimizing cloud spend: cost modeling and forecasting, rightsizing, reserved capacity and savings plans, autoscaling for cost, tagging and chargeback, and the FinOps operating model. Covers building the business justification for infrastructure spend and continuously driving efficiency at scale without sacrificing reliability. Cost as a first-class architectural concern.
How would you evaluate whether a given workload is actually a good candidate for spot or interruptible instances? Compare a stateful workload against a stateless one, and describe what would have to be true operationally before you'd recommend running the stateful one on spot. What savings would you expect, and what's the main risk?
Sample Answer
Direct answer
Spot (interruptible) capacity is a good fit when a workload is stateless or can checkpoint cheaply, tolerates being killed with little warning, and runs across enough parallel replicas that losing a few doesn't threaten the whole job. A stateless web tier behind a load balancer is close to the ideal case. A stateful workload (a database, a single-writer queue consumer, a long-running training job holding state in memory) can still go on spot, but only after you've engineered around the interruption, not by default. Expect roughly 50-90% off on-demand pricing depending on instance family, region, and how flexible you can be across types, with the main risk being correlated capacity reclaims: interruptions cluster by instance type and availability zone (AZ, a physically isolated data center location within a cloud region), so "one interruption" is often really "many at once."
Structured elaboration
Suitability criteria, in the order I'd check them:
- State locality: does losing the instance lose data that isn't durably stored elsewhere? If yes, that's the central risk to solve before anything else.
- Interruption notice window: most providers give a short warning (commonly around two minutes) before reclaiming capacity. Can the workload act on that window (flush buffers, deregister, checkpoint)?
- Restart cost: how expensive is it to lose progress and restart from the last checkpoint? A five-minute batch job restarting is nothing; a six-hour training run restarting from scratch is real money and time.
- Diversification headroom: can the workload run across multiple instance types, sizes, and AZs so a reclaim in one pool doesn't take out the whole fleet at once?
- Dependency coupling: does it hold external state (open DB connections, distributed locks, licensing seats) that a hard kill would corrupt or orphan?
| Dimension | Stateless (e.g. web/API worker) | Stateful (e.g. primary database, single-writer consumer) |
|---|---|---|
| Data loss on kill | None: request just retries elsewhere | Real, unless checkpointed or replicated |
| Recovery | New instance joins the pool immediately | Requires restore, replay, or failover |
| Diversification | Trivial (any instance meeting the spec works) | Constrained by attached storage, licensing, or leader state |
| Default fit for spot | Strong default | Conditional, only with the safeguards below |
What has to be true before I'd put a stateful workload on spot:
- State is externalized to a durable, non-spot service (a managed database, object storage, a managed cache) so the compute layer itself is disposable, or the workload checkpoints to durable storage frequently enough that replay cost after a kill is acceptable.
- There's a real handler for the interruption notice: on receiving it, the process drains in-flight work, writes a checkpoint, and deregisters cleanly, rather than being hard-killed.
- Leadership or single-writer roles are held by an on-demand (non-spot) instance, or use leader election with a persistent lock so a new leader can take over automatically if the current one disappears.
- The workload is spread across multiple instance types and AZs, because spot capacity pools and interruption risk are correlated within a single type/AZ.
- There's a fallback to on-demand if spot capacity is unavailable, so the workload degrades gracefully instead of stalling.
Worked example
Take a nightly batch job that re-trains a recommendation model, currently running on a single on-demand large instance for 6 hours. It's a reasonable spot candidate if:
- It checkpoints model state every 15 minutes to object storage.
- On receiving the interruption notice it saves a checkpoint immediately rather than waiting for the next scheduled interval.
- It's launched across 3-4 similarly-sized instance types so the scheduler can fall back to another pool instead of waiting for one specific type.
With that in place, an interruption mid-run costs at most ~15 minutes of recompute, not the full 6 hours, and the job still finishes inside its overnight window on most runs. Contrast that with the same team's primary transactional database: even with replication, putting the primary on spot risks a write-path outage measured in the failover time, not just lost compute, so it stays on-demand while read replicas or batch-analytics replicas (which are disposable) are strong spot candidates.
Trade-offs and pitfalls
- The headline 50-90% number is an upper bound; realized savings shrink once you account for the extra on-demand capacity you keep as fallback and the engineering time spent hardening the workload.
- Interruption rate is not uniform: newer or less-flexible instance types in a small number of AZs get reclaimed far more often than a diversified, older-generation footprint.
- A common wrong turn is treating "we added retries" as sufficient hardening. Retries handle transient failure; they don't handle a workload that loses in-memory state on every retry and therefore never makes forward progress under sustained interruption pressure.
- Licensing and data-locality constraints (per-core software licenses, data residency requirements tied to a specific AZ) can rule out spot for a workload that otherwise looks stateless.
What does 'unit economics' mean for a cloud service, and how would you measure cost per request and cost per customer for a multi-tier application? What data sources would you use, and what are the common pitfalls in attributing shared costs?
Sample Answer
Direct answer
Unit economics ties infrastructure cost to a business-meaningful unit, cost per request or cost per customer, so spend can be judged against value delivered instead of judged in isolation. You compute it by pulling total attributable cost for a service over a time window from the billing export, dividing by a volume metric (requests, active customers) for that same window from application telemetry, and the entire exercise lives or dies on how honestly you handle costs that don't belong to a single request or customer, which is the hard part.
Structured elaboration
Computing cost per request
Pull infrastructure cost for the service (compute, storage, networking, and an amortized share of any reserved capacity) from the billing export or Cost and Usage Report (CUR) for a fixed window, and total request count for the identical window from application performance monitoring (APM) telemetry or load balancer logs. Use an hourly or daily window for a service with volatile traffic, since averaging over a month can hide the fact that off-peak requests are effectively free (fixed capacity, low traffic) while peak requests are expensive (the capacity that gets added specifically to handle them).
Computing cost per customer
Aggregate the same billing data, plus any per-tenant resources (a dedicated database shard, a customer-specific storage bucket), by customer ID over a monthly window, since that aligns with billing cycles and typical churn reporting. Divide by active customers in the same window, not total signed-up customers, or a slow month for actual usage will make the metric look artificially good.
Data sources
- Billing export / CUR for raw dollar cost by service, region, and resource.
- Resource tags to attribute shared infrastructure to the right service.
- APM or request-log telemetry for volume (requests, active users).
- Container or orchestration metrics (CPU/memory requests) when a service shares a cluster with others, to split shared compute proportionally.
Attribution pitfalls
- Shared infrastructure (a load balancer, a shared cache, a shared database) has no natural single owner. Splitting it by proportional resource usage (CPU-seconds, request share) is the practical compromise, but it is an estimate, not a fact, and should be labeled as one in any report.
- Reserved capacity and committed discounts are paid for whether or not they're fully used in a given window; amortizing the commitment evenly across the term (rather than crediting it entirely to whichever week happened to use it) avoids a misleading cost-per-request spike in a quiet week.
- Caching materially changes the picture: a cache hit costs close to nothing at the origin, so blending cache hits and misses into one average cost-per-request understates the true marginal cost of a cache miss. Track them separately when the cache hit rate is high enough to matter.
- A handful of very large customers can swing a mean cost-per-customer figure enough to mislead a business conversation; report the distribution (median and a high percentile) alongside the mean, not the mean alone.
Worked example
A service handled 2,400,000 requests last month and its fully attributed infrastructure cost (direct compute plus its proportional share of a shared load balancer and database) was $19,200 for the month.
cost per request=2,400,000$19,200=$0.008Of that $19,200, $15,000 is directly attributable compute for this service alone, and $4,200 is this service's proportional share (based on measured request volume through the shared load balancer) of a $12,000 shared load balancer and cache bill split across three services. If a second service using that same shared infrastructure grows its traffic share next month, this service's $4,200 allocated portion drops even though its own direct compute cost didn't change, which is exactly the kind of shift a report needs to call out explicitly rather than let it read as an unexplained cost swing.
For cost per customer, if the service serves 8,000 active customers that same month:
cost per customer=8,000$19,200=$2.40If 50 of those 8,000 customers are enterprise accounts driving disproportionate request volume, the median customer's actual cost is well below $2.40 and the top-percentile customers are well above it, so reporting only the $2.40 mean to a pricing conversation would understate what the largest accounts actually cost to serve.
Trade-offs and pitfalls
The most common mistake is treating an allocation formula for shared costs as precise when it is an estimate, and letting a stakeholder make a pricing or roadmap decision on false precision. A second is misaligning the telemetry window and the billing window (comparing an hourly request count against a monthly bill), which produces numbers that look wrong even when the underlying data is fine. A third is reporting only a mean cost-per-customer in a business with a skewed customer-size distribution, which hides the accounts that are actually unprofitable to serve at current pricing.
A SaaS business is paying about $2 million a year in network egress. Propose architectural and operational changes that could realistically cut that by 30 to 50 percent, and explain how you'd model the ROI and payback time on the engineering investment needed to get there.
Sample Answer
Direct answer
I'd treat this as a portfolio of tactics, not one big fix: content delivery network (CDN, a network of edge servers that caches content closer to users) caching, compression, and delta-based transfer for large recurring payloads together plausibly reach the 30-50% target, and the business case is strong because the required engineering investment is small relative to a $2 million annual bill, typically paying back in a handful of months even under a pessimistic estimate of the savings.
Structured elaboration
The tactics, roughly ordered by effort-to-impact ratio:
- CDN adoption and caching. Serve static and cacheable content from edge locations instead of the origin. Impact scales with what fraction of total bytes are cacheable and how well cache headers are tuned; typically the single largest lever for a content- or API-heavy service.
- Compression. Enable modern compression (Brotli or gzip for text and JSON, modern image codecs like WebP or AVIF for images) on anything not already compressed. Cheap to implement, moderate impact, and the main cost is a small increase in origin CPU load.
- Delta or diff-based transfer. For large payloads that change incrementally between requests (client sync flows, binary updates), send only the changed portion instead of the full object. High impact where applicable, but only applies to specific traffic patterns, not general-purpose traffic.
- Direct network peering or a committed connectivity arrangement. For high-volume, predictable flows (bulk data transfer to a known partner, or a specific customer with heavy traffic), a direct connection or negotiated committed-volume pricing can cut the per-gigabyte rate materially, though it requires predictable volume to justify the setup cost.
- Reducing cross-region and cross-cloud replication for hot paths. Prefer region-local processing and asynchronous, lower-frequency replication over synchronous cross-region traffic on the hot path, since egress between regions or providers is often the most expensive category per gigabyte.
Modeling the ROI and payback:
- Classify current traffic by content type and destination to establish a baseline: how many bytes, at what cost, going where.
- Estimate each tactic's savings as a percentage of the baseline it actually applies to (compression only affects compressible content, delta transfer only affects flows with incremental change), not as a percentage of total spend.
- Sum the tactics' savings for a total expected reduction.
- Estimate the engineering cost as engineer-months at a fully-loaded cost rate.
- Payback period is the engineering investment divided by the annual savings, expressed in months.
Worked example
Baseline: $2,000,000 a year in egress. A conservative, tactic-by-tactic estimate: CDN and caching cuts 20% of total egress ($400,000), compression cuts 8% ($160,000), and delta transfer on the highest-volume sync flows cuts 12% ($240,000), for a combined 40% reduction, $800,000 a year in savings.
0.20+0.08+0.12=0.40,2,000,000×0.40=800,000Engineering cost: a first phase (CDN rollout and enabling compression) at 2 engineers for 2 months is 4 engineer-months, and a second phase (delta transfer implementation and peering setup) at 3 engineers for 4 months is 12 engineer-months, for 16 engineer-months total. At a fully-loaded cost of $15,000 per engineer-month, that's $240,000 in engineering investment.
Payback=800,000240,000=0.30 years=3.6 monthsSensitivity: if the tactics only achieve the low end of the target range, 30% instead of 40% ($600,000 a year), payback stretches to 4.8 months. At the high end, 50% ($1,000,000 a year), payback shortens to about 2.9 months. Even the conservative end of the range clears payback well inside a year, which is the strength of this specific business case: the engineering cost is small and mostly one-time, while the savings are large and recur every year afterward.
Trade-offs and pitfalls
- Estimating a tactic's savings as a percentage of total egress instead of the fraction of traffic it actually applies to is the most common way this kind of estimate goes wrong; compression does nothing for already-compressed video, and delta transfer does nothing for traffic that doesn't have a stable base to diff against.
- CDN caching introduces its own operational cost: cache invalidation complexity and the risk of serving stale content if cache lifetimes (time-to-live, TTL) aren't tuned correctly. That's a real, ongoing cost, not just a one-time setup cost.
- Compression trades bandwidth for CPU time; if origin servers are already CPU-constrained, the "savings" partially show up as a new compute cost instead of a pure win, and that trade-off needs to be measured, not assumed away.
- Multi-cloud or multi-region replication done for redundancy can quietly increase egress if it's not designed with cost in mind; prefer region-local processing and asynchronous, lower-frequency replication over synchronous cross-region traffic wherever the workload can tolerate it.
- The engineer-month cost estimate is only as good as the fully-loaded rate used; using a rate that's too low (ignoring benefits, overhead, and management time) makes every proposed initiative look more attractive than it actually is.
Define Total Cost of Ownership for a cloud migration. What cost components would you include when comparing an on-premises deployment to public cloud, including the one-time, ongoing, and easy-to-forget hidden costs, and how does the shift from capex to opex change how this gets reported to finance?
Sample Answer
Direct answer
Total cost of ownership (TCO) for a cloud migration is the full cost of moving to and running on the cloud over a fixed horizon, typically 3 to 5 years, not just the sticker price of the compute and storage you provision. It has to include one-time migration costs, ongoing operational costs, and a set of hidden costs that are easy to leave out of the first draft. Underneath the cost model sits an accounting shift that changes how finance evaluates the whole decision: on-premises infrastructure is largely capital expenditure (capex), paid upfront and depreciated over years on the balance sheet, while cloud spend is operating expenditure (opex), an ongoing monthly cost that hits the income statement as it's incurred, and that shift changes who approves the spend and how it's reported, independent of whether the total dollar amount is higher or lower.
Structured elaboration
One-time migration costs
Assessment and planning, re-architecting or refactoring applications that don't lift-and-shift cleanly, the data migration itself (transfer tooling, bandwidth, and often a real egress charge for the initial bulk move), landing-zone setup (networking, identity, security baseline), and the cutover and validation window, including the cost of any parallel-running period or rollback plan if the migration doesn't go cleanly.
Ongoing operational costs
Compute, storage, and networking at whatever mix of on-demand and committed pricing the workload ends up using; managed services (databases, messaging, CDN) priced per-use rather than as a fixed asset; the operations and SRE (site reliability engineering) staffing needed to run the environment, including on-call load, which doesn't disappear just because infrastructure moved to a managed platform; backup, disaster recovery, and monitoring tooling; and provider support-tier contracts.
Hidden and easy-to-forget costs
Software licensing terms often change under cloud deployment (per-core on-prem licensing doesn't always map cleanly to per-vCPU cloud pricing, and some vendors charge a premium for cloud deployment specifically). Team training and the productivity dip while staff ramp on new tooling is real money even though it never appears on a cloud invoice. Data egress charges recur beyond the initial migration if the architecture routinely moves data out of the cloud provider's network. And the performance-tuning and cost-optimization cycles that follow a migration (the first few months of "why is this more expensive than we modeled" work) are themselves an ongoing cost, not a one-time true-up.
The capex to opex shift, and why it changes the finance conversation
On-premises hardware is typically capitalized: bought upfront, placed on the balance sheet as an asset, and depreciated over its useful life, so the income-statement impact in any given year is just that year's depreciation, not the full purchase price. Cloud spend is usually opex: an ongoing operating cost that lowers reported profit in the period it's incurred, in full, the same way a utility bill does. This changes three things finance cares about independent of whether the total spend is higher or lower: the approval process (a large capex purchase typically needs a one-time capital-budget approval, while opex is a recurring line item reviewed every budget cycle, which can mean more frequent scrutiny), the reported financial metrics (heavy capex improves near-term reported profit relative to opex of the same economic cost, because depreciation spreads the hit over years), and the predictability of the number finance has to plan around (a capex purchase is a known fixed cost for its depreciation life; cloud opex scales with usage and can vary month to month, which is exactly why a TCO model with clear assumptions matters to the finance stakeholders reading it).
Assumptions to document alongside the number
State the time horizon and any discount rate used, the utilization and growth assumptions behind the ongoing-cost estimate, how each on-prem resource maps to a cloud instance type or managed-service tier, the pricing model assumed (on-demand vs. committed), and a confidence level per line item, since the migration and hidden-cost categories are usually far less certain than the ongoing compute estimate.
Worked example
A company is comparing a 3-year on-prem refresh against a 3-year cloud migration for one application. On-prem: a $600,000 hardware refresh (capex, depreciated straight-line over 3 years, so $200,000/year hits the income statement) plus $90,000/year in colocation, power, and ops staffing (opex), for a 3-year total cash outlay of $600,000 + $270,000 = $870,000, though the year-1 income-statement impact is only $200,000 (depreciation) + $90,000 (opex) = $290,000. Cloud: $40,000 in one-time migration cost, plus $220,000/year in compute, storage, managed database, and reduced ops staffing (opex throughout), for a 3-year total of $40,000 + $660,000 = $700,000, all recognized as opex in the year incurred.
TCOon-prem, 3yr=$600,000+(3×$90,000)=$870,000 TCOcloud, 3yr=$40,000+(3×$220,000)=$700,000Cloud is $170,000 cheaper over 3 years on total cost, but a finance team focused only on year-1 reported profit sees on-prem hit the income statement for $290,000 in year 1 versus cloud's $260,000 (the $40,000 migration cost plus $220,000 opex), a smaller gap than the 3-year total suggests, and by year 3 on-prem's income-statement hit is $200,000 (depreciation) + $90,000 (opex) = $290,000 again while cloud stays at $220,000. Presenting only the 3-year total TCO to a finance stakeholder who evaluates budgets year by year misses that the two options have a genuinely different shape over time, not just a different total.
Trade-offs and pitfalls
The most common mistake is comparing cloud opex against only the visible on-prem opex (power, colocation) while forgetting that the on-prem hardware's capex has a real, if deferred, cost through depreciation, which understates the true on-prem TCO. The second is presenting a single TCO number to finance without the capex-versus-opex framing, which leads to a confusing conversation when the "cheaper" option somehow needs a harder budget approval, because it's asking for a new recurring opex line instead of a one-time capital purchase finance may already have approved. The third is treating the migration-cost and hidden-cost categories with the same confidence as the ongoing-cost estimate; they are consistently the most underestimated part of a cloud TCO model, and should be flagged with a lower confidence level and a sensitivity range, not presented as a single precise figure.
What does 'rightsizing' mean for cloud compute, and what monitoring signals would tell you an instance is overprovisioned? Name one signal that looks convincing but can actually be misleading, and explain why.
Sample Answer
Direct answer
Rightsizing means matching an instance's provisioned CPU, memory, and I/O capacity to what the workload actually uses, so you're neither overpaying for idle headroom nor risking a service-level breach from underprovisioning. The signal that looks convincing but is actually misleading is a low average CPU utilization number by itself: averaged over a wide window, it can hide a workload with a sharp, business-critical peak that genuinely needs the capacity the average makes look wasted.
Structured elaboration
Signals that indicate genuine overprovisioning
- Sustained low CPU utilization (for example, consistently under 15%) across weeks, not just a quiet day, with request latency and error rates staying flat, meaning the extra CPU headroom isn't being called on even under normal load variation.
- Low memory usage with no swap activity and no out-of-memory events, indicating the instance could run on a smaller memory tier without risking a crash under load.
- Low disk I/O and network throughput relative to the instance's provisioned limits, with no I/O wait or throttling observed, indicating storage and network capacity are also oversized.
The misleading signal, and why
A low average CPU utilization computed over a long window (a week or a month) can average away a real, recurring peak: a batch job that runs at 90% CPU for two hours every night looks like "8% average utilization, clearly oversized" if you only look at the mean. Downsizing based on that average would cause the instance to fail or badly degrade during the exact two hours it matters most. The fix is to look at the P95 or P99 (95th or 99th percentile) utilization over the same window, not just the mean, and to check for scheduled or bursty jobs explicitly before resizing anything with a suspiciously low average.
How to act on the signals safely
Downsize in one step at a time, not straight to the theoretical minimum, and monitor for a full peak cycle (including any weekly or monthly batch jobs) before taking the next step down. Keep enough headroom above the P95, not the mean, to absorb normal variance and the occasional unplanned spike, and treat anything customer-facing or on the request path more conservatively than an internal batch worker, since the cost of underprovisioning a batch job (it runs slower) is much cheaper than the cost of underprovisioning a user-facing service (it errors out).
Worked example
An instance shows a 4-week average CPU utilization of 9%, memory averaging 22%, and no I/O throttling, which reads as a clear rightsizing candidate on the average alone. Pulling the P95 utilization for the same window shows CPU at 78%, driven by a nightly reconciliation job that runs for roughly 90 minutes each night at high CPU. The 9% average was real, but it was averaging 22.5 hours a day of near-idle time against 1.5 hours a day of near-saturation, and a downsize based on the average alone would have made that nightly job fail or run far past its window. The correct rightsizing move here is not "downsize," it's "keep capacity sized to the P95 peak, and consider whether the nightly job should run on separate, right-sized capacity instead of sharing the always-on instance," which is a different fix than the average suggested.
Trade-offs and pitfalls
The main pitfall is exactly the one this question is testing: resizing off a mean utilization number without checking for a P95/P99 tail or a scheduled workload hiding inside the average. A second is resizing too aggressively in one step and finding out only under the next real traffic spike that there's no headroom left, which turns a cost optimization into an incident. A third is treating rightsizing as a one-time cleanup instead of an ongoing signal to monitor, since workload shape changes over time (a service that grows its user base or adds a new scheduled job needs its rightsizing baseline revisited, not assumed permanent).
Unlock Full Question Bank
Get access to all 44 Cloud Cost Optimization and FinOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.