Site Reliability Engineering Principles Questions
The core SRE practice model: service-level objectives and indicators, error budgets, toil reduction, and reliability as an engineering discipline. Covers the principles and trade-offs behind treating operations as a software problem and balancing reliability against feature velocity. The conceptual foundation questions specific to SRE-style roles.
Your service consists of five serially-dependent services (A → B → C → D → E) and the end-to-end availability SLO is 99.95%. Propose how to apportion availability targets across these services, show the math converting per-service availability to an end-to-end SLO, and explain how you would detect which service contributes most to end-to-end failures and how to set per-service error budgets.
Sample Answer
Apportioning a 99.95% end-to-end target across five serial services means solving for a per-service availability that, multiplied five times, reaches the target, then building the attribution machinery to find which service is actually the biggest contributor when it's missed.
Structured elaboration
The EQUAL-apportionment starting point: if all five services split the target evenly, each needs 0.99951/5 availability. Computed: 0.99950.2≈0.99990, i.e. roughly 99.990% per service. In practice, equal apportionment is a reasonable DEFAULT but not always the right final answer: a service that's inherently harder to make reliable (a stateful data store versus a stateless compute layer) might reasonably get a looser individual target while a cheaper-to-harden service absorbs a tighter one, as long as the composite still clears 99.95%.
Worked example (computed and verified)
0.99951/5=0.99950.2≈0.999900. Verified: raising 0.9999 to the 5th power gives back approximately 0.9995, confirming the apportionment is self-consistent. Detecting which service contributes most to failures requires REQUEST-ID trace propagation: every request carries a shared trace ID through all five hops, so when an end-to-end failure occurs, the trace reveals exactly which hop(s) failed, letting failure attribution be computed directly from real traces rather than inferred indirectly. Per-service error budgets are then set from each service's own apportioned target, and consumption is attributed to whichever service the trace shows actually failed.
Trade-offs and pitfalls
A critical subtlety when composing a FRONTEND SLO from BACKEND SLOs across a hierarchy (not just this flat five-service chain) is avoiding DOUBLE-COUNTING: if the frontend's own SLO is computed independently from user-facing telemetry AND each backend's SLO is tracked separately, a single failure can appear to "count" against multiple budgets simultaneously unless ownership and counting rules are made explicit (e.g. the frontend's SLO reflects the FULL user experience including backend failures, while each backend's own SLO is tracked and attributed separately for THAT team's accountability, without either double-counting the same failure as two independent incidents in aggregate reporting). This general attribution approach also needs to handle the case where the chain isn't purely serial: multiple DIFFERENT user journeys sometimes share a single common piece of infrastructure (not a simple A-through-E chain but a shared dependency several paths route through), and responsibility separation there is different from a serial chain, since a single shared-infrastructure failure can simultaneously consume budget across several otherwise-independent journeys, requiring the attribution logic to recognize "one root cause, multiple affected journeys" rather than attributing the failure five separate times as if it were five independent incidents.
Cascading failures also complicate simple per-hop attribution: if B's failure causes C to time out waiting on B, and C's timeout in turn cascades to D, a trace-based attribution needs to identify B as the ROOT cause, not falsely spread blame (and budget consumption) evenly across B, C, and D as though each failed independently.
Given Prometheus counters: http_requests_total{service="images",code="200"} and http_requests_total{service="images"}, write a PromQL expression that computes the 5-minute availability SLI (successful requests / total requests) for service 'images' and a rule that fires an alert if availability < 99.95% for 15 minutes. Explain your use of functions and evaluation interval.
Sample Answer
A simple availability alert compares a directly-computed success ratio against a fixed threshold and requires that condition to persist for a sustained window, which is the right first alert to build before adding burn-rate sophistication.
Structured elaboration
The ratio itself comes from two counters (successful requests over total requests for the service), aggregated with sum(rate(...)) over a short window (5 minutes here) so the ratio reflects RECENT behavior rather than an all-time average; a for: clause on the alert then requires the low-availability condition to hold continuously for the stated duration before paging, which filters out single-scrape noise.
Worked example (executed; promtool unit test against a real Prometheus binary)
- record: job:availability:5m
expr: sum(rate(http_requests_total{service="images",code="200"}[5m]))
/ sum(rate(http_requests_total{service="images"}[5m]))
- alert: LowAvailability
expr: job:availability:5m < 0.9995
for: 15m
labels:
severity: page
Verified against a synthetic counter pair (994 good : 6 bad requests per interval, i.e. an observed availability of 99.4%): the alert correctly fired after the 15-minute persistence window against a 99.95% threshold.
Trade-offs and pitfalls
A short 5-minute evaluation window is sensitive to real changes quickly but is also the window most vulnerable to a low-traffic service producing a wildly noisy ratio from just one or two errors; if this service handles only a handful of requests per 5 minutes, widen the window or add a minimum-request-count guard so the alert doesn't fire on statistically meaningless samples. This alert alone also does not distinguish a slow, sustained low-grade degradation from a sudden severe outage, both look identical once the 15-minute persistence has elapsed; that distinction is exactly what a second, burn-rate-based alert (see the companion burn-rate design) is for.
Explain the differences between SLI, SLO, and SLA. Provide a concrete example for a customer-facing REST API: specify one SLI (metric and units), an SLO target (with measurement window), and a sample SLA clause suitable for a contract. Describe how you would operationalize the SLO (measurement, dashboards, alerting) and how error budget policies would influence release velocity and incident response.
Sample Answer
An SLI measures what is actually happening, an SLO is the internal target you hold yourself to, and an SLA is the external promise with consequences attached; the three form a strict hierarchy of increasing formality and decreasing strictness.
Structured elaboration
For a customer-facing REST API, a concrete SLI might be "proportion of HTTP requests to /orders that return within 300ms and a non-5xx status, measured over a 5-minute rolling window, numerator = good requests, denominator = total requests." The SLO built on that SLI would be a target like "99.9% of requests meet the SLI over a rolling 30-day window," which is an internal engineering commitment, not a contract. The SLA is the customer-facing document, typically looser than the SLO (e.g. 99.5% monthly uptime, with defined service credits for breaches), because the internal target needs headroom to absorb normal operational noise before ever risking the external promise. Ownership typically follows this same layering: engineering defines and owns the SLI's instrumentation, SRE/engineering leadership own the SLO target, and legal/business own the SLA's contractual language, though all three should be negotiated together, not handed off in sequence. Operationalizing the SLO requires measurement, dashboards, and alerting working together: measurement is the SLI's live computation described above; the dashboard should show the rolling SLI value, the remaining error-budget percentage, and the current burn rate side by side, not just a single pass/fail indicator against 99.9%, since a snapshot alone hides whether the trend is improving or actively degrading; and alerting should be tiered by burn rate (a fast, short-window check for a sudden severe spike, a slower, longer-window check for a sustained low-grade degradation), feeding both the on-call rotation and the release-management process that consults the same dashboard before approving a risky deploy.
Worked example
A team ships an SLO of 99.9% success rate over 30 days and signs an SLA of 99.5% monthly uptime with a service credit for anything below that. The 0.4 percentage-point buffer between the two is the team's safety margin: at 99.9% they are inside their own target and comfortable; if reliability degrades toward 99.6-99.7%, the SLO is already breached and internal error-budget policy should be triggering a release freeze, well before the customer-facing SLA is ever at risk. A common pitfall is mapping SLIs straight into an SLA with no such buffer (setting the SLA equal to, or tighter than, the SLO); that removes the internal early-warning function entirely, since an SLO breach and an SLA breach then happen simultaneously. In practice, this team's dashboard shows the 30-day SLI trending at 99.92%, remaining error budget at 60%, and a 1-hour burn rate of 1.2x, which is exactly the picture that lets an engineer distinguish "healthy and stable" from "healthy today but trending toward trouble" at a glance.
Trade-offs and pitfalls
Retries and asynchronous callbacks are the classic edge case in how the SLI's numerator/denominator get computed: does a request that fails once but succeeds on an automatic client retry count as "good" (from the end user's perspective, yes) or does it still reflect a real backend problem worth tracking separately? Most teams track both a client-observed success rate (post-retry, what the SLA should reference) and a raw backend error rate (pre-retry, what triggers internal alerting), because collapsing them into one number hides whether retries are quietly masking a growing problem. Operationally, when error budget starts burning, release velocity should slow (fewer risky deploys, more canary time) well before an SLA-level incident response is triggered; the SLO's own error-budget policy is what should catch problems early enough that the SLA-level contractual machinery never needs to engage.
Write a PromQL expression to compute a rolling 1-hour burn rate for an SLO based on a 'success_rate' metric, and write an alert that fires if burn rate > 2 for 30 minutes. Explain how burn rate is calculated from SLI time-series and justify the window lengths and threshold chosen.
Sample Answer
A rolling 1-hour burn-rate alert needs the burn rate computed from the SAME success-rate series that the SLO is measured against, and the alert should page only once that elevated rate has PERSISTED, not on a single noisy evaluation.
Structured elaboration
Burn rate is (1−observed success rate)/(1−SLO): a value of 1 means the service is consuming budget exactly at the sustainable rate, and a value of 2 means it is burning twice as fast as the budget can tolerate long-term. The critical, easy-to-get-wrong detail is what success_rate actually IS in your metric system: if it is already a computed ratio gauge (0 to 1), it must be averaged over the window with avg_over_time, never passed through rate() (rate() is only valid on monotonically increasing counters; applied to a gauge it silently returns something close to zero for a roughly-constant value, which then maxes out the burn-rate formula and pages on EVERY evaluation regardless of true health, a genuine bug this answer caught by executing the query rather than just reasoning about it).
Worked example (executed; promtool unit test against a real Prometheus binary)
- record: job:burn_rate:1h
expr: (1 - avg_over_time(success_rate[1h])) / (1 - 0.999)
- alert: BurnRateHigh
expr: job:burn_rate:1h > 2
for: 30m
labels:
severity: page
First version of this rule used rate(success_rate[1h]) instead of avg_over_time; tested against a synthetic success_rate gauge held constant at 0.9999 (well above the 99.9% SLO), the buggy version STILL fired the alert on every evaluation, because rate() on a non-monotonic, roughly-flat gauge returns ~0, making (1-0)/0.001 = 1000. After switching to avg_over_time, the same 0.9999-constant series produces no alert, and a constant 0.99 series (genuinely 10x over budget) correctly fires BurnRateHigh after persisting 30 minutes. Both cases confirmed with promtool test rules.
Trade-offs and pitfalls
A single 1-hour window, even correctly computed, will page on a real but short-lived blip that resolves before 30 minutes are up only if the blip itself is long enough to sustain a >2x average over the full hour; pairing this with a much shorter window (e.g. 5 minutes) at a higher burn-rate threshold catches fast, severe outages sooner, which is the standard multi-window pattern. The for: 30m clause is doing real work here (requiring the condition to hold for the whole 30 minutes before firing) and is a separate lever from the window size in the avg_over_time range selector; conflating the two is a common source of confusion when tuning alert sensitivity.
How do you test and validate that an SLI truly reflects customer experience and is not biased by instrumentation or sampling? Provide a validation plan including synthetic tests, correlating real-user telemetry (RUM), shadow traffic and acceptance criteria.
Sample Answer
An SLI is only as trustworthy as its correlation with real user experience, so validating it means actively looking for evidence the metric could be biased by exactly HOW it's collected, not just trusting that a well-designed-sounding SLI must be accurate.
Structured elaboration
Validation plan: (1) synthetic tests, deliberately injecting a known failure or degradation and confirming the SLI actually moves as expected, in the expected direction and magnitude, a direct causal check that the instrumentation actually detects what it claims to; (2) correlating against RUM (real-user monitoring), checking that periods where the SLI shows good health also show healthy real-user-observed experience, and vice versa, since a persistent disagreement between the two is exactly the kind of bias signal worth investigating (e.g. a server-side SLI staying green while RUM shows real degradation, suggesting the SLI is measuring something too narrow, like server-side-only latency, missing a real client-side or network-level problem); (3) shadow traffic, routing a copy of real production traffic through an alternate measurement path and comparing its SLI reading against the primary path's, which can reveal instrumentation-specific bias (e.g. if the shadow path samples differently or measures at a different point in the request lifecycle and gets a meaningfully different number, that difference itself is diagnostic).
Worked example
Acceptance criteria: the synthetic-failure-injection test should show the SLI degrading by an expected, roughly-predictable amount for a known-severity injected failure (not necessarily exact, but in the right direction and rough magnitude); the RUM correlation should show no SUSTAINED divergence beyond an agreed tolerance over a representative period (occasional brief divergence is expected noise, a persistent gap is a validation failure); shadow-traffic comparison should show the two measurement paths agreeing within an agreed tolerance band, with any persistent, unexplained gap treated as an open validation defect requiring investigation before the SLI is trusted for a real SLO commitment.
Trade-offs and pitfalls
Sampling bias is a specific, easy-to-miss failure mode worth checking explicitly: if the SLI's underlying data collection systematically excludes a subset of real traffic (e.g. requests from users on unstable connections failing to successfully report their own telemetry, a pattern discussed elsewhere), the SLI will look healthier than reality specifically for the population it's failing to measure, and neither a synthetic test (which doesn't naturally exercise this bias) nor a simple RUM-correlation check (if RUM has the SAME sampling bias) would necessarily catch it; shadow traffic with a DELIBERATELY different sampling or measurement mechanism is specifically valuable here because it's less likely to share the exact same bias as the primary path. This validation shouldn't be a one-time exercise either: as the underlying system and its traffic patterns evolve, a periodic re-validation (not just a one-off at initial SLI design time) is what catches bias that develops later, rather than assuming a metric validated once stays valid forever.
Unlock Full Question Bank
Get access to all 24 Site Reliability Engineering Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.