Cloud Cost Optimization and FinOps Questions
Controlling and optimizing cloud spend: cost modeling and forecasting, rightsizing, reserved capacity and savings plans, autoscaling for cost, tagging and chargeback, and the FinOps operating model. Covers building the business justification for infrastructure spend and continuously driving efficiency at scale without sacrificing reliability. Cost as a first-class architectural concern.
For a throughput-oriented service that's moderately stateful, how would you decide between covering it with reserved instances or savings plans versus mixing in spot instances with on-demand? What would you need to assume about utilization and interruption rates, and how would you validate the chosen mix safely before committing to it at scale?
Sample Answer
Direct answer
The decision comes down to how expensive an interruption actually is for this specific service. A reserved instance or Savings Plan (SP) covers a guaranteed baseline at a discount with zero interruption risk; spot buys a deeper discount in exchange for eviction risk. For a moderately stateful, throughput-oriented service, the right structure is usually a reserved or Savings Plan floor sized to the steady minimum load, so core capacity never carries interruption risk, plus spot covering the elastic portion above that floor, with on-demand as the fallback when spot capacity isn't available, validated incrementally rather than committed to at full scale on day one.
Structured elaboration
What you need to know before deciding
- Baseline and peak utilization, so you know how much load is genuinely steady versus how much is elastic.
- Historical interruption rate for the specific instance family and region you'd run spot on; this is workload-specific and should be measured, not assumed from a generic industry figure.
- Mean time to recovery (MTTR): how long it takes the service to recover from an interruption, and what that recovery actually costs, in replayed work, extra network or storage I/O (input/output), or a brief latency hit.
- The specific nature of the statefulness: session affinity, whether writes go through a write-ahead log, how frequently the service checkpoints. "Moderately stateful" is doing a lot of work in this question; a service that checkpoints every few seconds tolerates interruption very differently from one that holds long-lived in-memory session state.
Portfolio design
Monthlymix=FloorHours×ReservedRate+ElasticHours×(1+RetryOverhead)×SpotRate
Size the floor to the steady minimum load the service never drops below, covered by reserved capacity or a Savings Plan so it's never at interruption risk. Size the elastic band to the variable load above that floor, covered by spot, with on-demand as a last-resort fallback when spot capacity isn't available in the target instance family or region. This mirrors the same floor-plus-elastic-band logic that applies to a large batch-job fleet: there, the floor is whatever backlog has to clear even in a slow week, and the elastic band is burst capacity, which is a much easier target for spot than a live stateful service because batch jobs tolerate restarts far more cheaply.
Validating the mix before committing at scale
- Canary a small percentage of production traffic onto the mixed fleet first, with autoscaling and an on-demand fallback path already wired up, rather than assuming the mix works and finding out otherwise in production.
- Run controlled interruption tests (deliberately terminating spot capacity in the canary) to validate the MTTR and recovery-cost assumptions against reality, not just the historical interruption-rate figure.
- Track cost per month, tail latency (P99, 99th percentile), error rate, and any lost or replayed work as the canary scales up, and only widen the spot percentage once those metrics hold at each step.
Worked example
Suppose the service needs a steady floor of 600 instance-hours/month plus an elastic band averaging 400 instance-hours/month on top, for 1,000 hours/month total. On-demand costs $0.10/hour, a 1-year reserved commitment (amortized) costs $0.06/hour, and spot costs $0.02/hour. Because this service is stateful (checkpoint and replay cost money, unlike a stateless batch job), assume a higher interruption-driven retry overhead than a purely stateless workload, 15%, based on the higher end of a typical measured range for this kind of workload.
Option A, all reserved:
1,000 hrs×$0.06=$60.00 per month
Zero interruption risk, but paying for the full 1,000-hour floor at all times even though only 600 of it is steady load.
Option B, floor plus elastic mix:
Floor:600×$0.06=$36.00
Elastic band with 15% retry overhead:400×1.15=460 effective hours,460×$0.02=$9.20
Total:$36.00+$9.20=$45.20 per month
Savings:
$60.00−$45.20=$14.80 per month,60.0014.80≈24.7% cheaper than covering the whole workload with reserved capacity
That saving comes in exchange for accepting eviction risk on the 400-hour elastic band, plus whatever operational cost comes from handling those interruptions gracefully, which isn't captured in this dollar figure and has to be validated separately through the canary process above.
Trade-offs and pitfalls
- Sizing the floor too small exposes steady-state load to interruption risk it shouldn't have to carry. This is the specific mistake the "moderately stateful" framing in the question is pointing at: a stateful service often can't just retry cheaply, so the floor needs to protect whatever load genuinely can't tolerate an interruption.
- Sizing the floor too large gives up savings a fault-tolerant elastic band could have captured, effectively turning the whole workload back into option A without admitting it.
- Skipping the incremental validation step means finding out the MTTR assumption was wrong in production, at full scale, instead of in a controlled canary test where the blast radius is small.
- Reusing a generic industry interruption-rate figure instead of measuring your own misprices the whole decision, since interruption rates vary meaningfully by instance family, region, and time.
Explain the difference between showback and chargeback as cloud cost allocation models. What operational and behavioral impacts does each have on engineering teams, and in what situation would you recommend one over the other?
Sample Answer
Direct answer
Showback reports each team's cloud costs for visibility without moving any money: nobody's budget is actually debited. Chargeback goes further and allocates real costs to a team's budget, typically through an internal invoice or a direct debit against their cost center. The mechanics of allocation (tagging, cost pools) are identical between the two models; what differs is whether the number is informational or binding, and that single difference changes team behavior more than almost any other FinOps decision.
Structured elaboration
Operational requirements
- Showback needs accurate tagging and a reporting pipeline (dashboards built on the provider's billing export), but no accounting integration. It is comparatively cheap to stand up.
- Chargeback needs everything showback needs, plus allocation rules for shared and hard-to-attribute costs (a shared database, a platform team's infrastructure), an internal billing or budget-debit mechanism, and usually a dispute process for when a team contests its bill. It is meaningfully more operational overhead.
Behavioral impacts
- Showback creates awareness but relies on a team choosing to act on it. It works well when the goal is building cost literacy and trust in the data, and it fails quietly: a team can see an inflated bill for months and simply not prioritize fixing it, because nothing forces the issue.
- Chargeback creates direct, budget-line accountability, which reliably produces the fastest optimization response. It also produces predictable second-order effects: teams start negotiating over shared-cost allocation formulas, and some teams under-provision or avoid experimentation because the cost is now visibly theirs. Badly designed chargeback (especially unfair shared-cost splits) actively damages trust in the whole program.
When to recommend which
- Recommend showback when tagging discipline and cost data are still immature, when the organization is early in FinOps adoption and needs cultural buy-in before it can survive a contentious billing dispute, or when the goal this quarter is visibility, not enforcement.
- Recommend chargeback once allocation is trustworthy, budget owners are clearly defined, and leadership needs teams to make trade-offs against a real budget constraint (a business unit that must self-fund its cloud spend, for example).
- In practice the strongest programs run a hybrid: chargeback for costs that are cleanly attributable to a single team (dedicated compute, a service's own database), and showback for genuinely shared infrastructure (a shared Kubernetes cluster, a platform team's networking spend) where a clean per-team split would be arbitrary and would just generate disputes instead of better decisions. This avoids forcing a false precision onto costs that are structurally shared.
Worked example
A platform team's shared cluster costs $40,000 a month and hosts workloads for three product teams, roughly split 50/30/20 by measured resource requests. Under showback, all three teams see "$20,000 / $12,000 / $8,000, informational" on a dashboard, and it is up to each team whether to act on their share. Under chargeback, those same three figures are debited from each team's budget as an internal invoice line, and a team now has to justify that $20,000 (or reduce it) the same way it justifies any other budget line. A hybrid design would chargeback the dedicated services each team also runs outside the shared cluster (fully attributable, no allocation dispute possible) while keeping the shared cluster on showback, because a resource-request-based 50/30/20 split is an estimate, not a precise cost, and billing teams against an estimate they can contest is a common source of program-trust failure.
Trade-offs and pitfalls
The biggest pitfall is skipping straight to chargeback before tagging and allocation are trustworthy: teams will contest a bill they believe is wrong, and if the underlying data really is wrong, the program loses credibility fast and is hard to recover. A second pitfall is chargeback without a clear owner for genuinely shared costs, which pushes teams toward proportional formulas nobody fully agrees with and creates ongoing friction that has nothing to do with actual waste. A third, subtler failure is showback with no organizational follow-through: if visibility never translates into any consequence, teams learn to ignore the dashboard, and the "awareness" goal quietly fails too.
What does 'rightsizing' mean for cloud compute, and what monitoring signals would tell you an instance is overprovisioned? Name one signal that looks convincing but can actually be misleading, and explain why.
Sample Answer
Direct answer
Rightsizing means matching an instance's provisioned CPU, memory, and I/O capacity to what the workload actually uses, so you're neither overpaying for idle headroom nor risking a service-level breach from underprovisioning. The signal that looks convincing but is actually misleading is a low average CPU utilization number by itself: averaged over a wide window, it can hide a workload with a sharp, business-critical peak that genuinely needs the capacity the average makes look wasted.
Structured elaboration
Signals that indicate genuine overprovisioning
- Sustained low CPU utilization (for example, consistently under 15%) across weeks, not just a quiet day, with request latency and error rates staying flat, meaning the extra CPU headroom isn't being called on even under normal load variation.
- Low memory usage with no swap activity and no out-of-memory events, indicating the instance could run on a smaller memory tier without risking a crash under load.
- Low disk I/O and network throughput relative to the instance's provisioned limits, with no I/O wait or throttling observed, indicating storage and network capacity are also oversized.
The misleading signal, and why
A low average CPU utilization computed over a long window (a week or a month) can average away a real, recurring peak: a batch job that runs at 90% CPU for two hours every night looks like "8% average utilization, clearly oversized" if you only look at the mean. Downsizing based on that average would cause the instance to fail or badly degrade during the exact two hours it matters most. The fix is to look at the P95 or P99 (95th or 99th percentile) utilization over the same window, not just the mean, and to check for scheduled or bursty jobs explicitly before resizing anything with a suspiciously low average.
How to act on the signals safely
Downsize in one step at a time, not straight to the theoretical minimum, and monitor for a full peak cycle (including any weekly or monthly batch jobs) before taking the next step down. Keep enough headroom above the P95, not the mean, to absorb normal variance and the occasional unplanned spike, and treat anything customer-facing or on the request path more conservatively than an internal batch worker, since the cost of underprovisioning a batch job (it runs slower) is much cheaper than the cost of underprovisioning a user-facing service (it errors out).
Worked example
An instance shows a 4-week average CPU utilization of 9%, memory averaging 22%, and no I/O throttling, which reads as a clear rightsizing candidate on the average alone. Pulling the P95 utilization for the same window shows CPU at 78%, driven by a nightly reconciliation job that runs for roughly 90 minutes each night at high CPU. The 9% average was real, but it was averaging 22.5 hours a day of near-idle time against 1.5 hours a day of near-saturation, and a downsize based on the average alone would have made that nightly job fail or run far past its window. The correct rightsizing move here is not "downsize," it's "keep capacity sized to the P95 peak, and consider whether the nightly job should run on separate, right-sized capacity instead of sharing the always-on instance," which is a different fix than the average suggested.
Trade-offs and pitfalls
The main pitfall is exactly the one this question is testing: resizing off a mean utilization number without checking for a P95/P99 tail or a scheduled workload hiding inside the average. A second is resizing too aggressively in one step and finding out only under the next real traffic spike that there's no headroom left, which turns a cost optimization into an incident. A third is treating rightsizing as a one-time cleanup instead of an ongoing signal to monitor, since workload shape changes over time (a service that grows its user base or adds a new scheduled job needs its rightsizing baseline revisited, not assumed permanent).
Compare reserved instances, savings plans, and committed-use discounts across the major cloud providers. What is the mechanical difference between them in commitment scope, term, and flexibility across instance types, and how would you decide what percentage of a steady-state workload's capacity to commit?
Sample Answer
Direct answer
All three mechanisms trade a usage commitment for a lower price, but they commit to different things: AWS Reserved Instances (RIs) commit to a specific instance configuration, AWS Savings Plans commit to a dollar-per-hour spend level that flexes across instance types, Google Cloud committed-use discounts (CUDs) commit to either a resource quantity or a dollar-per-hour spend depending on which CUD type you buy, and Azure Reservations commit to a specific VM configuration similar to AWS RIs. The general pattern across every provider is the same trade-off: the more precisely you commit to a specific instance shape, the bigger the discount; the more flexibility you keep, the smaller the discount but the lower your risk if the workload changes shape.
Structured elaboration
Mechanism comparison
| Mechanism | Provider | Commits to | Term | Flexibility |
|---|---|---|---|---|
| Standard Reserved Instance | AWS | Specific instance family, size, region | 1 or 3 yr | Least flexible: can change availability zone and, within limits, instance size in the same family, but not family or OS |
| Convertible Reserved Instance | AWS | Instance family (exchangeable) | 1 or 3 yr | Can exchange for a different family, size, or OS during the term, at a lower discount than Standard |
| Compute Savings Plan | AWS | Dollar-per-hour compute spend | 1 or 3 yr | Most flexible: applies across instance family, size, OS, tenancy, and region, and across EC2, Fargate, and Lambda |
| EC2 Instance Savings Plan | AWS | Dollar-per-hour spend, locked to one instance family and region | 1 or 3 yr | Flexible on size and OS within that family and region only; typically a larger discount than Compute Savings Plans for the same term because it's narrower |
| Resource-based CUD | Google Cloud | A quantity of vCPUs, memory, GPUs, or similar, on Compute Engine | Typically 1 or 3 yr | Locked to the committed resource type and quantity; scope can be a single project or shared across a billing account |
| Flexible (spend-based) CUD | Google Cloud | Dollar-per-hour spend | 1 or 3 yr | Pools eligible spend across only three services, Compute Engine, Google Kubernetes Engine (GKE), and Cloud Run, similar in spirit to an AWS Compute Savings Plan in that the discount follows a dollar-per-hour spend level rather than a specific SKU. BigQuery and Cloud SQL are NOT part of this pool: each has its own separate, service-specific spend-based commitment, purchased and applied independently |
| Reserved VM Instance | Azure | Specific VM series, size, and region | 1 or 3 yr | Instance-size flexibility within the same VM size-flexibility group; can be rescoped after purchase to a subscription, resource group, shared billing scope, or management group without a new commercial transaction |
All four providers offer some form of upfront, partial-upfront, or no-upfront (pay monthly) payment on these commitments at the same total cost, so the payment option is a cash-flow decision, not a discount-size decision on most of these products.
Why the scope difference matters in practice
A resource-level commitment (Standard RI, resource-based CUD, Azure Reservation) only pays off if the workload keeps needing that exact shape for the whole term; if the team migrates to a different instance family six months in, the commitment sits partially wasted (though AWS and Azure both allow some exchange or resale mechanisms to recover part of that). A spend-based commitment (Compute Savings Plan, flexible CUD) survives an instance-family change automatically, because the discount is applied to dollars spent on eligible usage, not to a specific SKU, at the cost of a somewhat smaller discount than the narrowest resource-level option.
Deciding what percentage of steady-state capacity to commit
Start from the floor, not the average: pull 3 to 6 months of utilization history for the workload, and find the usage level that held true on the worst week, not the typical week. That floor, not the mean, is the safe commitment baseline, because a commitment above the actual steady floor pays for idle capacity on every low-usage day. From there, the commitment size is a risk trade-off, not a fixed rule: a stable, mature workload with a long recent history of holding above that floor supports committing close to the full floor, while a workload still changing shape (recent re-architecture, aggressive growth, planned migration) justifies leaving more of the floor on-demand or covering it with a flexible, spend-based commitment instead of a rigid resource-level one, specifically because the risk being managed is "commitment outlives the workload's actual shape," not "commitment size in the abstract."
Worked example
A team's steady-state EC2 fleet held at a minimum of 40 instances of a given family over the last 4 months, with normal weekday peaks around 55 and occasional bursts to 70. The 40-instance floor is the commitment candidate, not the 55-instance average and not the 70-instance peak: committing at 55 would mean paying the commitment rate for capacity that isn't reliably used on quieter days, and any spike above 40 (up to and including the 70-instance bursts) is served by on-demand or spot capacity regardless of the commitment size. If this workload is expected to stay on the same instance family for the full term, an EC2 Instance Savings Plan or Standard RI sized to 40 instances captures the largest discount available on that stable floor; if a re-platforming project is likely to change instance family within the year, a Compute Savings Plan sized to the equivalent dollar-per-hour spend protects the same floor's discount while surviving the family change.
Trade-offs and pitfalls
The most common mistake is committing to the peak or the average instead of the floor, which either overpays for capacity that isn't reliably used or, worse, sizes a "safe" commitment so conservatively it captures almost none of the available discount. The second is choosing the narrowest, highest-discount resource-level commitment for a workload that's still changing shape, and then discovering the commitment doesn't match the new instance family, wasting real money for the rest of the term. The third, specific to the flexible/spend-based products, is assuming "flexible" means "no attention needed": a spend-based commitment still needs the underlying usage to stay above the committed dollar level, or the unused portion is still paid for and simply not applied to any usage.
You join a company as Cloud Architect and inherit a large account with no tagging discipline and no real cost visibility. Walk through your first 90 days retrofitting cost governance: what you tag first, how you handle the backlog of untagged resources, and how you get product teams to keep tagging going forward.
Sample Answer
Direct answer
The first 90 days should move in three stages: get a minimal mandatory tag schema enforced on everything new, triage the untagged backlog by cost rather than by resource count so the highest-spend unknowns get remediated first, and build the habit of tagging into how teams already provision, rather than adding it as a separate compliance chore that fades once the initial push ends.
Structured elaboration
Days 1-30: define and enforce for new resources
- Publish a minimum mandatory tag schema: owner (team or individual), product, environment (dev/stage/prod), and a cost-center code finance can map to the general ledger. Keep it short; a long mandatory list gets worked around.
- Enforce it at provisioning time going forward, through infrastructure-as-code templates that fail to apply without the required tags, and cloud-native policy guardrails as a backstop. The actual enforcement code and CI (continuous integration) wiring is platform/infrastructure work; what FinOps owns here is defining the schema and the policy, not writing the policy-as-code itself.
- Get an executive sponsor and a stated deadline, since a 90-day retrofit without visible leadership backing tends to lose priority against feature work by week three.
Days 31-60: triage the existing backlog by cost, not count
- Run an asset-discovery inventory to find every untagged resource, then sort by cost, not by how many resources are untagged. A thousand small untagged resources at $2/month each matter far less than a handful of large untagged compute clusters.
- Auto-tag what can be inferred confidently (resource creator identity, naming convention, VPC or account membership), but treat auto-tagged results as provisional, especially before they ever feed a chargeback dispute, since a wrong inference there just relocates the trust problem.
- For the remainder, run a tagging campaign: assign the highest-cost untagged resources to specific owners with a deadline, and apply a visible "unknown" tag with an owner notification for anything that misses it.
Days 61-90: make it stick
- Launch a showback dashboard so teams can see their own tagged spend, which is what turns tagging from a mandate into something teams actually want kept current.
- Set a follow-up cadence (monthly for the next two quarters) to tighten enforcement gradually, for example moving from "warn on missing tags" to "block deploy on missing tags" once teams have had time to adjust.
Worked example
Imagine three weeks into the role, a monitoring alert flags an unexpected $50,000 charge for the month. Digging in, it traces to a set of large storage volumes and a couple of oversized compute instances that were spun up for a since-abandoned proof of concept over a year earlier, never tagged, and therefore invisible to every dashboard the previous team relied on. This is exactly the failure mode the cost-first triage in days 31-60 is designed to catch: because the backlog is sorted by cost rather than resource count, a handful of large untagged resources like this surface immediately instead of getting lost among thousands of small, cheap, untagged items. The incident also becomes the concrete justification for the executive sponsorship asked for in the first 30 days, since "here is $50,000 a year we were spending on nothing, invisible because nothing was tagged" is a much easier case to make to leadership than an abstract governance pitch.
Trade-offs and pitfalls
- Over-enforcing on day one blocks legitimate work. Making every tag mandatory and deploy-blocking before teams have had time to adjust their workflows builds resentment and invites workarounds (a placeholder tag value just to get past the check) that defeat the purpose.
- Under-enforcing means it never actually happens. A tagging policy with no enforcement mechanism and no deadline stays a wiki page nobody reads.
- Auto-tagging heuristics are provisional, not authoritative. Inferring an owner from a resource's creator or naming pattern is a reasonable starting point for visibility, but treating it as final, especially once it starts driving a chargeback bill, will produce disputes you can't defend.
- Sorting the backlog by resource count instead of cost wastes the 90 days. The $50,000 surprise above sat in a handful of resources; spending the sprint tagging thousands of near-zero-cost items first would have missed it entirely.
Unlock Full Question Bank
Get access to all 23 Cloud Cost Optimization and FinOps interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.