SLIs, SLOs, SLAs, and Error Budgets Questions
Defining and operating reliability targets. Covers choosing service level indicators, setting service level objectives and agreements, computing and spending error budgets, and using them to drive engineering decisions. Includes negotiating reliability targets with stakeholders.
Describe the key success metrics and SLOs you defined for a production data service you built. Include quantitative thresholds (throughput, p95/p99 latency, data freshness, error rate), how you chose the thresholds based on business needs, and how you validated those SLOs during launch and operation.
Sample Answer
Choosing SLOs for a data service you actually built starts from the specific business need the service exists to serve, then works down to the quantitative thresholds, not the other way around.
Structured elaboration
Success metrics worth defining: throughput (requests or records served per unit time, confirming the service can sustain expected load), p95/p99 latency (tail experience, since a data service's consumers often run bursty batch queries where the tail matters as much as, or more than, the median), data freshness (how current the served data is relative to its source), and error rate (proportion of requests failing outright). Thresholds should be chosen against what the CONSUMING use case actually requires: a dashboard refreshed hourly can tolerate looser freshness than a real-time alerting pipeline consuming the same underlying data service.
Worked example
For a production data service feeding both an hourly-refreshed executive dashboard and a near-real-time fraud-detection pipeline: p95 latency < 200ms (chosen based on what the fraud pipeline's own SLA requires, the tighter of the two consumers), freshness < 5 minutes (again driven by the tighter consumer, even though the dashboard alone would have tolerated much looser freshness), error rate < 0.1%. Validation during launch: a soft-launch period against a subset of real consumer traffic, checking that the chosen thresholds are both ACHIEVABLE (the service can actually sustain them under real load) and SUFFICIENT (the tightest consumer's actual needs are genuinely met, confirmed by that consumer's own downstream metrics staying healthy).
Trade-offs and pitfalls
Setting one blended threshold to satisfy the AVERAGE of all consumers' needs, rather than the TIGHTEST consumer's actual requirement, is a common mistake: it can look reasonable on paper while quietly under-serving the one consumer (often the most business-critical one) whose needs were actually the tightest. It's also worth revisiting these thresholds as new consumers are added over the service's life, since a threshold validated against the original set of consumers may become insufficient once a new, more demanding consumer starts depending on the same service without anyone re-checking whether the existing SLO still covers their needs.
You want to detect regressions from a canary using statistical hypothesis testing rather than fixed thresholds. Describe a test design to detect a significant change in latency or error rate during canary rollout that balances false positives and time-to-detection. Include choice of test, sample sizes, significance level, and how you would action on a positive result.
Sample Answer
Fixed-threshold canary checks either miss a real but modest regression or false-alarm on normal noise; a proper statistical test explicitly balances those two error types and gives you a principled way to decide how much data you need before acting.
Structured elaboration
Test design: a two-sample statistical test comparing the canary's observed metric (error rate, or latency) against the stable control's observed metric over the same time window, using a test appropriate to the metric's actual distribution: a two-proportion z-test (or Fisher's exact test at low sample sizes) for a rate metric like error rate, and a test appropriate for skewed distributions (e.g. a Mann-Whitney U test, since latency data is rarely normally distributed) rather than a naive t-test for a continuous metric like latency. Sample size and significance level: choose a significance level (e.g. alpha = 0.01, a stricter-than-typical bar, since a false positive here triggers an unnecessary rollback with real cost) and compute the minimum sample size needed to detect a MEANINGFUL effect size (e.g. a 20% relative increase in error rate) at that significance level with reasonable statistical power (e.g. 80%), rather than testing continuously on tiny early samples that have no real power to detect anything reliably yet.
Worked example
Balancing false positives against time-to-detection: at a stricter significance level (lower alpha), you need MORE samples before the test can confidently flag a real regression, which means slower detection; a looser significance level detects faster but risks more false-positive rollbacks. The practical compromise many teams use is a SEQUENTIAL testing approach (checking at multiple points as data accumulates, with a significance level adjusted for the repeated testing to avoid inflating the overall false-positive rate from checking too many times), rather than either a single test at a fixed, possibly-too-early sample size or waiting for one enormous sample before ever checking at all.
Trade-offs and pitfalls
The single most common statistical mistake here is repeatedly peeking at a test's p-value as data accumulates and stopping as soon as it first crosses the significance threshold, without correcting for the fact that repeated peeking inflates the TRUE false-positive rate well above the nominal alpha (a well-known problem sometimes called "peeking" or "optional stopping"); a genuinely valid sequential design (like a group sequential test, or a fixed pre-registered set of peek points with an adjusted significance level at each) is needed to avoid this, not an ad hoc "keep checking until it looks significant" approach. Actioning a positive result: a confirmed statistically significant regression should trigger an automatic rollback, but the specific effect size detected (not just "significant," which says nothing about MAGNITUDE) should inform how urgently to act, since a statistically significant but practically tiny effect size may not warrant the same urgency as a large one, even at the same significance level.
Implement (or outline) a service that rate-limits approval of feature releases based on global and per-service error budget constraints. Inputs: service_id, requested_approvals_per_hour; outputs: allowed_approvals and suggested_cooldown. Describe data model, algorithms, and how to handle burst approvals and manual overrides.
Sample Answer
This is a rate-limiter, but the thing being limited is release APPROVALS rather than requests, and the limit itself should shrink as either the global or the service-level error budget burns down.
Structured elaboration
Data model: a per-service record tracks {service_id, hard_cap_per_hour, budget_remaining_pct, rolling_approval_count, window_start_timestamp}, refreshed from the same error-budget system that computes SLO burn; a single global record tracks the analogous {global_hard_cap_per_hour, global_budget_remaining_pct, rolling_global_approval_count} shared across all services; and every approval decision (including overrides) is written to an append-only approval-log table {service_id, timestamp, requested, allowed, cooldown, manual_override, justification} so the audit trail exists independent of the current-state records, which only hold the latest rolling counts. Two caps apply simultaneously: a hard per-service cap (protects that service from too-frequent changes) and a hard global cap (protects shared deploy infrastructure and blast radius across all services); the function must always honor the tighter of the two. On top of the hard caps, each cap is scaled down as its corresponding budget burns: above 50% remaining budget the full hard cap applies, and below 50% it ramps linearly to zero at 0% remaining, so approvals get progressively harder to obtain exactly when the service or the whole fleet is already unhealthy. A manual override exists for genuine emergencies (a critical security patch), but it must still respect the hard caps (never grant more than the absolute ceiling) so a panicked override can't itself become the cause of an outage.
Worked example (executed; python3, verified against 5 scenarios)
import math
def approve_release(service_id, requested_approvals_per_hour, global_budget_remaining_pct,
service_budget_remaining_pct, global_cap_per_hour=20,
service_cap_per_hour=5, manual_override=False):
def budget_scaled_cap(hard_cap, remaining_pct):
if remaining_pct >= 50:
return hard_cap
return math.floor(hard_cap * (remaining_pct / 50.0))
if manual_override:
return min(requested_approvals_per_hour, global_cap_per_hour, service_cap_per_hour), 0
g_cap = budget_scaled_cap(global_cap_per_hour, global_budget_remaining_pct)
s_cap = budget_scaled_cap(service_cap_per_hour, service_budget_remaining_pct)
allowed = max(0, min(requested_approvals_per_hour, g_cap, s_cap))
binding = min(g_cap, s_cap)
if binding <= 0:
cooldown = 60
elif requested_approvals_per_hour > binding:
cooldown = math.ceil(60 / max(binding, 1))
else:
cooldown = 0
return allowed, cooldown
Confirmed behavior: a request for 3 approvals/hour with both budgets healthy passes through unchanged (allowed=3, cooldown=0). A burst request of 10/hour against a healthy 5/hour service cap is capped at 5 with a 12-minute suggested cooldown between bursts. When the service budget alone drops to 10% remaining, the effective service cap shrinks to 1/hour even though the request and the global budget are unchanged, correctly localizing the throttle to the unhealthy service. A manual override at 5%/5% budget still caps at 5 (the hard service cap), not the requested 10, so the override cannot exceed the absolute ceiling.
Trade-offs and pitfalls
A pure per-hour cap can be gamed by bursting right at the top of every hour boundary; a production version should use a sliding window (or a token-bucket refill) rather than a fixed hourly reset. The linear budget-scaling ramp is a simplification; teams often prefer a steeper (e.g. quadratic) ramp near zero so the last sliver of budget is protected much more aggressively than the middle of the range. Logging every override with its justification is not optional: a bypass mechanism with no audit trail is exactly the kind of control that gets silently abused under deadline pressure.
You are designing SLIs and SLOs for a multi-tenant SaaS product that serves 50k tenants and handles 100k requests/s. Propose candidate SLIs (e.g., p95 latency, success rate, background job completion), the aggregation strategy (per-tenant vs global), reasonable measurement windows, and SLO targets. Explain trade-offs when choosing per-tenant SLOs vs a single global SLO.
Sample Answer
At 50k tenants and 100k requests/second, the central design decision is whether SLOs live per-tenant, globally, or in tiers in between, and that choice has to be made deliberately rather than defaulting to whichever is easiest to implement.
Structured elaboration
Candidate SLIs: p95 latency, request success rate, and background-job completion rate (since a SaaS product typically has both synchronous request paths and asynchronous background work, and the two need separate SLIs). On aggregation: a single GLOBAL SLO is cheap to compute and dashboard but can hide a small number of severely-affected tenants inside a healthy-looking average; a strict PER-TENANT SLO gives precise visibility but is expensive to compute and alert on at 50,000 tenants, and most tenants don't generate enough traffic for a per-tenant SLI to be statistically meaningful anyway. The practical middle ground most teams land on is TIERED: free/paid/enterprise (or premium/lower-priority-internal) classes each get their own SLO target and error-budget allocation, with high-volume individual tenants (usually enterprise) getting a dedicated per-tenant SLO where their traffic supports it, and everyone else rolled into their tier's aggregate. On measurement windows: alerting should run on a short rolling window, 5 minutes to 1 hour, so an active incident is caught quickly, while the SLO target itself should be evaluated over a longer rolling window, 28 or 30 days, so ordinary day-to-day noise doesn't make the target itself unstable. At 100k requests/second, even a 5-minute window carries tens of millions of samples for the global aggregate and typically thousands to millions for a mid-size tenant tier, so the window choice here is really about alert responsiveness versus target stability, not about needing more raw samples to be statistically meaningful.
Worked example
Enterprise tier: dedicated per-tenant SLOs, 99.95% success over a rolling 30-day window (since these tenants have both enough traffic to make the number meaningful and enough business importance to justify the operational cost of tracking them individually), with a 5-minute rolling window feeding burn-rate alerting for fast detection. Paid tier: a shared aggregate SLO across the tier, 99.9% over the same 30-day window, with automated alerting if any single paid tenant's own error rate spikes well above the tier average over a 1-hour window, without maintaining a full per-tenant SLO for every one of them. Free tier: a looser aggregate SLO, 99.5% over 30 days, reflecting lower business criticality and correspondingly lower operational investment, with alerting on a longer window (e.g. daily) since page-worthy urgency is lower for this tier.
Trade-offs and pitfalls
Shared infrastructure creates a real capacity-conflict risk in a tiered model: if enterprise and free-tier traffic share underlying resources, a burst of low-priority traffic can consume capacity that degrades a high-value tenant's SLO, which argues for either resource isolation (dedicated capacity per tier) or explicit prioritization/throttling of lower tiers under load, backed by a policy for what automatically happens (e.g. throttling or feature restriction) when a given tier's shared error budget is exhausted. Tracking only a global SLO with no tenant-tier breakdown at all is the single biggest risk at this scale: a handful of severely-impacted tenants (even a fraction of a percent of 50,000) is easily invisible in an aggregate that's 99.9% healthy overall.
Describe how to integrate error budget checks into progressive feature-flag rollouts (e.g., incrementing percentage of users). Include the gating rules, automated rollback triggers, rollback thresholds relative to burn rate, and how to scale rollout speed when budget is healthy.
Sample Answer
Integrating error-budget checks into a progressive feature-flag rollout means the rollout's SPEED, not just its final go/no-go outcome, should respond continuously to the target service's current budget health.
Structured elaboration
Gating rules: before advancing a flag's rollout percentage to the next stage, check both the FEATURE's own observed SLI (is the new code path itself healthy) and the underlying SERVICE's overall error budget (is there enough margin to safely absorb the feature's risk at all, even if the feature's own metrics look fine so far). Automated rollback triggers: if the feature-specific error rate or latency at the current rollout percentage crosses a threshold relative to the pre-rollout baseline, automatically revert the flag to 0% (or the last known-good percentage) rather than merely pausing progression. Rollback thresholds relative to burn rate: the threshold for triggering a rollback should scale with the SERVICE's overall budget health, not stay fixed; if the service's overall budget is already healthy, tolerate a slightly larger feature-specific regression before rolling back (more margin to absorb it); if the service's overall budget is already low, roll back on a much smaller feature-specific regression, since there's little room left to absorb ANY additional risk.
Worked example
With the service's overall budget healthy (>70% remaining), the flag can advance from 5% to 25% to 100% on a standard cadence (e.g. every 30 minutes if metrics stay clean), and the feature-specific rollback threshold tolerates up to a 2x relative regression before reverting. With the service's overall budget already at 20% remaining, the SAME flag's rollout automatically slows (longer soak time at each stage, smaller percentage increments) and the feature-specific rollback threshold tightens to a 1.2x relative regression, since the service can't afford to absorb much additional risk right now regardless of how good the feature itself looks.
Trade-offs and pitfalls
A rollout gate that only checks the feature's OWN metrics, ignoring the underlying service's overall budget health, can approve a feature rollout that looks individually fine while the SERVICE as a whole is already in a fragile state, compounding an existing problem rather than being appropriately cautious about it. It's also worth deciding explicitly how multiple SIMULTANEOUS feature rollouts on the same service interact with a shared budget: two features each individually well within their own thresholds could together exhaust the service's overall budget if evaluated independently without accounting for concurrent risk from other in-flight rollouts.
Unlock Full Question Bank
Get access to all SLIs, SLOs, SLAs, and Error Budgets interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.