Site Reliability Engineering Principles Questions
The core SRE practice model: service-level objectives and indicators, error budgets, toil reduction, and reliability as an engineering discipline. Covers the principles and trade-offs behind treating operations as a software problem and balancing reliability against feature velocity. The conceptual foundation questions specific to SRE-style roles.
You are leading an initiative to raise SLO reliability targets across the organization, but teams resist because of the perceived additional work. How do you lead this change: how do you make the case, negotiate with resistant teams, and sequence the rollout?
Sample Answer
Leading an org-wide push to raise reliability targets against real resistance is a change-management problem as much as a technical one: teams resist not because they disagree reliability matters, but because higher targets translate into concrete near-term costs (slower releases, more testing, on-call burden) that they, not the person setting the target, will absorb. The approach that works is building the case with data specific to each team rather than a blanket mandate, sequencing the rollout so early wins build credibility before asking for the hardest changes, and giving teams real say in how they meet the new bar rather than dictating the mechanism.
Making the case
- Start from cost of the status quo, not the desired future state. "Raise your service-level objective (SLO) to 99.95%" is an abstract ask; "your service's current reliability cost the team an estimated 40 engineer-hours in incident response last quarter, and a 99.95% target would have prevented roughly two of those three incidents" is concrete and specific to the team you're talking to.
- Show, don't just tell, what the new bar requires. Bring a worked example from a team that already operates at the target level, including what it actually cost them to get there, so the ask isn't abstract.
Negotiating with resistant teams
- Ask what's actually driving the resistance before countering it; it's often a legitimate, specific constraint (a fragile legacy dependency, a team already stretched thin on headcount) rather than disagreement with the goal itself, and the right response differs depending on which it is.
- Offer a real say in the mechanism. Teams that get to choose how they close the gap (more automated testing versus a slower release cadence versus additional on-call staffing) resist less than teams handed a single prescribed solution.
- Make the new bar's cost visible and funded, not an unstated expectation added on top of existing commitments; a target raised without any corresponding capacity is a target that quietly gets ignored under deadline pressure.
Sequencing the rollout
- Start with one or two teams who are closest to the target already, or most receptive, to produce an early, visible success story.
- Use that story, with real numbers, to make the case to the next wave rather than repeating the abstract pitch from scratch.
- Reserve the hardest, most resistant teams for last, once there's a track record and, ideally, tooling or process improvements from the earlier waves that reduce their cost of adoption.
Worked example
An organization wants every customer-facing service to meet a 99.9% availability target within a year; three teams are well below it. Rather than mandating the change everywhere at once, the rollout starts with the team closest to the bar already (a two-month push gets them there, at a documented cost of roughly one sprint of focused work), uses that team's before-and-after incident data to make the case to the next two teams, and specifically asks the most resistant of the three what's blocking them: it turns out to be a shared, fragile authentication dependency, not disagreement with the target. The fix becomes a shared investment (fixing the dependency) that unblocks that team and de-risks two others at the same time, rather than three teams independently working around the same underlying problem.
Trade-offs and pitfalls
A pure top-down mandate is faster to announce but slower to actually land, since teams that feel a target was imposed without their input tend to comply minimally or find ways to game the metric rather than genuinely improve; a fully bottom-up, consensus-driven approach respects that but can stall indefinitely on the hardest, most resistant teams. The sequencing approach here is a deliberate middle path, and it costs time up front (the early wins take real effort to produce) in exchange for the later waves moving faster once there's credibility and, often, shared infrastructure fixes to build on.
Draft an SLA negotiation template to use with enterprise customers. Include measurable components (availability, latency, throughput), measurement windows, exclusions and blackout windows, monitoring sources for disputes, remedy/credit structure, dispute resolution and escalation clauses, and how to align the negotiated SLA to internal SLOs and capacity planning.
Sample Answer
An SLA negotiation template needs to name every component a real dispute would eventually turn on, so it functions as a genuine starting point for negotiation rather than a document that looks complete but has gaps a sophisticated enterprise counterparty will immediately probe.
Structured elaboration
Measurable components: availability, latency (with explicit percentiles, e.g. p95 and p99, not just an average), and throughput, each defined precisely enough that both sides agree in advance what "meeting the target" means. Measurement windows: state both the evaluation period (typically monthly) and, separately, any shorter windows used for real-time internal alerting versus the contractual reporting period, since these can legitimately differ. Exclusions and blackout windows: scheduled maintenance (with a defined maximum monthly/quarterly allowance and advance-notice requirement) and force-majeure categories, named specifically rather than left vague. Monitoring sources for disputes: state whose data is authoritative (provider logs, a named third-party monitor, or a reconciliation process between both parties' data) to avoid the classic "our numbers don't match yours" stalemate. Remedy/credit structure: tiered service credits scaled to severity, with a stated cap. Dispute resolution: a defined escalation path (technical review, then a management-level conversation, then, if unresolved, a named arbitration or mediation process) before any resort to litigation. Alignment to internal SLOs and capacity planning: internally, confirm the negotiated SLA sits comfortably inside the existing internal SLO's safety margin before signing, since committing externally to a number your internal SLO can't sustainably support creates immediate structural risk.
Worked example
A template clause: "Availability shall be measured as [defined ratio] over each calendar month, using Provider's production monitoring system, cross-checked quarterly against [named third-party monitor]. Scheduled maintenance, not exceeding 4 hours per quarter with 72 hours' advance notice, is excluded. Should monthly availability fall below 99.9%, Customer is entitled to a service credit of 10% of that month's fees; below 99.5%, 25%; three consecutive months below 99.9% entitles Customer to terminate for cause without penalty. Disputes regarding measured availability shall first be reviewed jointly by both parties' technical teams within 10 business days, escalating to management review if unresolved, and to [named] mediation before any litigation."
Trade-offs and pitfalls
Negotiating a tighter SLA than your internal SLO can sustainably support, purely to win the deal, creates a structural mismatch that surfaces painfully later, either as chronic credit payouts or as constant internal pressure to hit a number the architecture wasn't built for; the internal-alignment check should happen BEFORE the negotiation, not be discovered as a problem after signing. It's equally risky to leave the measurement-source question vague ("availability shall be measured appropriately"), since that vagueness is exactly what a real dispute exploits; naming the authoritative source and a reconciliation process up front removes the single most common point of later conflict.
Given a PostgreSQL table http_logs(id PK, service varchar, status_code int, occurred_at timestamptz), write a SQL query that computes the daily error rate (percentage of requests with status_code >= 500) for service = 'checkout' over the last 30 days. Ensure your query returns days with zero requests as 0% error.
Sample Answer
Reporting a daily error rate that must show 0% on days with zero traffic requires generating the full calendar of days first and left-joining actual data onto it, rather than grouping only the rows that exist.
Structured elaboration
A plain GROUP BY date(occurred_at) silently omits any day with no matching rows, which is exactly the case this question calls out as a requirement to handle correctly. The fix is a recursive date-spine (or an equivalent calendar table) covering the full 30-day range, LEFT JOINed to the aggregated daily counts, with COALESCE converting missing aggregates to zero before computing the percentage (and guarding the percentage calculation itself against a divide-by-zero when total is zero).
Worked example (executed; sqlite3, verified against a dataset with an intentional zero-request day)
WITH RECURSIVE days(d) AS (
SELECT date('2026-07-01')
UNION ALL
SELECT date(d, '+1 day') FROM days WHERE d < date('2026-07-01', '+29 day')
),
daily AS (
SELECT date(occurred_at) AS d, COUNT(*) AS total,
SUM(CASE WHEN status_code >= 500 THEN 1 ELSE 0 END) AS bad
FROM http_logs
WHERE service = 'checkout'
AND occurred_at >= datetime('2026-07-01') AND occurred_at < datetime('2026-07-01','+30 day')
GROUP BY date(occurred_at)
)
SELECT days.d,
COALESCE(daily.total, 0) AS total_requests,
COALESCE(daily.bad, 0) AS errors,
ROUND(CASE WHEN COALESCE(daily.total,0) = 0 THEN 0.0
ELSE daily.bad * 100.0 / daily.total END, 2) AS error_pct
FROM days LEFT JOIN daily ON days.d = daily.d
ORDER BY days.d;
Against a dataset with requests on July 1st and 3rd but NONE on July 2nd, the query correctly returned 50.0% error for July 1st (1 of 2 requests failed), and exactly 0.0% (not a missing row, not NULL) for the zero-traffic July 2nd, confirming the calendar-spine approach satisfies the stated requirement.
Trade-offs and pitfalls
A date-spine built with a recursive CTE is portable but can be slow at very large day-ranges in some engines; PostgreSQL's generate_series is the more idiomatic and typically faster equivalent for the same purpose. It's also worth being explicit about WHY a zero-traffic day reports 0% rather than being excluded or shown as NULL: for an SLO dashboard, a silent gap (day missing from the output entirely) is easy to misread as "no problem that day" when it might actually mean the service was down so hard it received no traffic at all, which a naive GROUP BY would hide.
Explain the differences between SLI, SLO, and SLA. Provide a concrete example for a customer-facing REST API: specify one SLI (metric and units), an SLO target (with measurement window), and a sample SLA clause suitable for a contract. Describe how you would operationalize the SLO (measurement, dashboards, alerting) and how error budget policies would influence release velocity and incident response.
Sample Answer
An SLI measures what is actually happening, an SLO is the internal target you hold yourself to, and an SLA is the external promise with consequences attached; the three form a strict hierarchy of increasing formality and decreasing strictness.
Structured elaboration
For a customer-facing REST API, a concrete SLI might be "proportion of HTTP requests to /orders that return within 300ms and a non-5xx status, measured over a 5-minute rolling window, numerator = good requests, denominator = total requests." The SLO built on that SLI would be a target like "99.9% of requests meet the SLI over a rolling 30-day window," which is an internal engineering commitment, not a contract. The SLA is the customer-facing document, typically looser than the SLO (e.g. 99.5% monthly uptime, with defined service credits for breaches), because the internal target needs headroom to absorb normal operational noise before ever risking the external promise. Ownership typically follows this same layering: engineering defines and owns the SLI's instrumentation, SRE/engineering leadership own the SLO target, and legal/business own the SLA's contractual language, though all three should be negotiated together, not handed off in sequence. Operationalizing the SLO requires measurement, dashboards, and alerting working together: measurement is the SLI's live computation described above; the dashboard should show the rolling SLI value, the remaining error-budget percentage, and the current burn rate side by side, not just a single pass/fail indicator against 99.9%, since a snapshot alone hides whether the trend is improving or actively degrading; and alerting should be tiered by burn rate (a fast, short-window check for a sudden severe spike, a slower, longer-window check for a sustained low-grade degradation), feeding both the on-call rotation and the release-management process that consults the same dashboard before approving a risky deploy.
Worked example
A team ships an SLO of 99.9% success rate over 30 days and signs an SLA of 99.5% monthly uptime with a service credit for anything below that. The 0.4 percentage-point buffer between the two is the team's safety margin: at 99.9% they are inside their own target and comfortable; if reliability degrades toward 99.6-99.7%, the SLO is already breached and internal error-budget policy should be triggering a release freeze, well before the customer-facing SLA is ever at risk. A common pitfall is mapping SLIs straight into an SLA with no such buffer (setting the SLA equal to, or tighter than, the SLO); that removes the internal early-warning function entirely, since an SLO breach and an SLA breach then happen simultaneously. In practice, this team's dashboard shows the 30-day SLI trending at 99.92%, remaining error budget at 60%, and a 1-hour burn rate of 1.2x, which is exactly the picture that lets an engineer distinguish "healthy and stable" from "healthy today but trending toward trouble" at a glance.
Trade-offs and pitfalls
Retries and asynchronous callbacks are the classic edge case in how the SLI's numerator/denominator get computed: does a request that fails once but succeeds on an automatic client retry count as "good" (from the end user's perspective, yes) or does it still reflect a real backend problem worth tracking separately? Most teams track both a client-observed success rate (post-retry, what the SLA should reference) and a raw backend error rate (pre-retry, what triggers internal alerting), because collapsing them into one number hides whether retries are quietly masking a growing problem. Operationally, when error budget starts burning, release velocity should slow (fewer risky deploys, more canary time) well before an SLA-level incident response is triggered; the SLO's own error-budget policy is what should catch problems early enough that the SLA-level contractual machinery never needs to engage.
Your SLA requires 99.9% freshness for derived metrics used on dashboards. Define 4 SLIs and an SLO you would recommend for services that compute these metrics and describe how you'd measure and report them.
Sample Answer
A 99.9% freshness SLA on derived dashboard metrics needs to be decomposed into SLIs that each capture a different stage where freshness could be lost, since a single blended "is it fresh" number hides WHERE in the pipeline a delay is actually occurring.
Structured elaboration
Four SLIs: (1) source-to-ingestion lag - time from the source event occurring to it landing in the raw ingestion layer; (2) ingestion-to-transform lag - time from raw ingestion to the derived-metric computation completing; (3) end-to-end freshness - the combined total, source event to dashboard-visible metric, which is the number that actually maps to the 99.9% SLA commitment; (4) computation success rate - proportion of scheduled metric-computation runs that complete successfully at all, since a freshness number computed only from SUCCESSFUL runs silently ignores runs that failed entirely, which is itself a freshness (and correctness) problem.
Worked example
SLO: 99.9% of end-to-end freshness measurements land under a defined threshold (e.g. under 10 minutes) over a 30-day rolling window, with the three component SLIs (source-to-ingestion, ingestion-to-transform, computation success rate) tracked as diagnostic breakdowns rather than separately-committed SLAs. If the end-to-end SLI degrades, the three component SLIs let you immediately localize whether the delay is happening at ingestion (a source-system problem) or at transform (a compute-pipeline problem), without which you'd only know "it's slow" with no actionable next step.
Trade-offs and pitfalls
Reporting only the end-to-end number without the component breakdown is fine for the customer-facing SLA report but nearly useless for diagnosing an actual regression internally; both views need to exist, aimed at different audiences. It's also worth being careful about how the computation-success-rate SLI interacts with the freshness SLI: a completely FAILED computation run has no freshness reading at all (there's no metric to measure the lag of), so it must be explicitly counted against the SLO as a worst-case freshness violation, not silently excluded from the freshness average simply because it produced no valid data point to measure.
Unlock Full Question Bank
Get access to all 40 Site Reliability Engineering Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.