Site Reliability Engineering Principles Questions
The core SRE practice model: service-level objectives and indicators, error budgets, toil reduction, and reliability as an engineering discipline. Covers the principles and trade-offs behind treating operations as a software problem and balancing reliability against feature velocity. The conceptual foundation questions specific to SRE-style roles.
How do you test and validate that an SLI truly reflects customer experience and is not biased by instrumentation or sampling? Provide a validation plan including synthetic tests, correlating real-user telemetry (RUM), shadow traffic and acceptance criteria.
Sample Answer
An SLI is only as trustworthy as its correlation with real user experience, so validating it means actively looking for evidence the metric could be biased by exactly HOW it's collected, not just trusting that a well-designed-sounding SLI must be accurate.
Structured elaboration
Validation plan: (1) synthetic tests, deliberately injecting a known failure or degradation and confirming the SLI actually moves as expected, in the expected direction and magnitude, a direct causal check that the instrumentation actually detects what it claims to; (2) correlating against RUM (real-user monitoring), checking that periods where the SLI shows good health also show healthy real-user-observed experience, and vice versa, since a persistent disagreement between the two is exactly the kind of bias signal worth investigating (e.g. a server-side SLI staying green while RUM shows real degradation, suggesting the SLI is measuring something too narrow, like server-side-only latency, missing a real client-side or network-level problem); (3) shadow traffic, routing a copy of real production traffic through an alternate measurement path and comparing its SLI reading against the primary path's, which can reveal instrumentation-specific bias (e.g. if the shadow path samples differently or measures at a different point in the request lifecycle and gets a meaningfully different number, that difference itself is diagnostic).
Worked example
Acceptance criteria: the synthetic-failure-injection test should show the SLI degrading by an expected, roughly-predictable amount for a known-severity injected failure (not necessarily exact, but in the right direction and rough magnitude); the RUM correlation should show no SUSTAINED divergence beyond an agreed tolerance over a representative period (occasional brief divergence is expected noise, a persistent gap is a validation failure); shadow-traffic comparison should show the two measurement paths agreeing within an agreed tolerance band, with any persistent, unexplained gap treated as an open validation defect requiring investigation before the SLI is trusted for a real SLO commitment.
Trade-offs and pitfalls
Sampling bias is a specific, easy-to-miss failure mode worth checking explicitly: if the SLI's underlying data collection systematically excludes a subset of real traffic (e.g. requests from users on unstable connections failing to successfully report their own telemetry, a pattern discussed elsewhere), the SLI will look healthier than reality specifically for the population it's failing to measure, and neither a synthetic test (which doesn't naturally exercise this bias) nor a simple RUM-correlation check (if RUM has the SAME sampling bias) would necessarily catch it; shadow traffic with a DELIBERATELY different sampling or measurement mechanism is specifically valuable here because it's less likely to share the exact same bias as the primary path. This validation shouldn't be a one-time exercise either: as the underlying system and its traffic patterns evolve, a periodic re-validation (not just a one-off at initial SLI design time) is what catches bias that develops later, rather than assuming a metric validated once stays valid forever.
Design SLIs/SLOs for a machine-learning prediction service where ground-truth labels arrive hours or days later. Propose proxy SLIs for near-real-time monitoring (e.g., prediction distribution drift, model confidence), explain how you would backfill true correctness when labels arrive, and how error budgets could be applied to model rollbacks.
Sample Answer
When ground truth arrives hours or days late, the SLI strategy needs a fast, imperfect PROXY for real-time monitoring and a slower, authoritative BACKFILL process for eventually confirming (or correcting) what the proxy suggested, with error-budget policy built around both.
Structured elaboration
Proxy SLIs for near-real-time monitoring: prediction-distribution drift (comparing the live distribution of model outputs against a known-healthy baseline distribution, detectable instantly without waiting for ground truth), and model confidence (if the model exposes a confidence/probability score, a sustained drop in average confidence can be an early, if imperfect, signal of degrading performance even before ground truth confirms it). Backfilling true correctness: once labels arrive (hours or days later), retroactively compute the TRUE accuracy for that historical window and reconcile it against what the proxy signals suggested at the time, both to correct the historical record and to continuously calibrate how well the proxy signals actually predict eventual ground-truth accuracy.
Worked example
Error budget applied to model rollbacks: if drift and confidence signals both cross a concerning threshold in near-real-time, that's enough evidence to trigger a PRECAUTIONARY rollback to the last validated model version, even before ground truth confirms the degradation, since waiting days for confirmation while a genuinely degraded model keeps serving live traffic is its own real cost; the rollback decision here is explicitly a risk-based bet on the proxy's reliability, not a certainty. When ground truth eventually arrives and CONFIRMS the degradation, that validates the rollback decision and should count against the model's error budget for the period the degraded model was actually live; if ground truth instead shows the proxy signal was a false alarm (the model was actually fine), that's valuable calibration data suggesting the proxy's threshold may need adjustment, not necessarily evidence the rollback itself was the wrong call given what was known at the time.
Trade-offs and pitfalls
The central tension here is that the proxy signals are, by design, imperfect early warnings, so acting on them (a precautionary rollback) always carries some risk of an unnecessary rollback on a model that would have turned out fine; the alternative, waiting for ground truth before ever acting, guarantees a real, if unquantified, exposure window where a genuinely bad model continues serving live traffic. The right calibration comes from tracking, over time, how often the proxy signals turn out (once ground truth confirms) to have been right versus false alarms, and tuning the proxy threshold based on that track record rather than treating the initial threshold choice as permanently correct.
You are responsible for defining SLA tiers for a portfolio of services with varying business impact. Propose a framework to map business impact to SLA tiers, include example SLA/SLO targets per tier, measurement and monitoring requirements, alerting implications, and how you would price or justify premium SLA tiers to stakeholders.
Sample Answer
Mapping business impact to SLA tiers means grounding the tier BOUNDARIES in something measurable (revenue exposure, customer count, regulatory requirement) rather than a subjective "feels important" judgment, since the tiering decision needs to survive scrutiny from stakeholders who didn't make the original call.
Structured elaboration
Framework: score each service's business impact along a few concrete dimensions (direct revenue dependency, number of customers or user-facing surface affected, any regulatory/compliance requirement, and reputational/brand exposure), then map the resulting composite impact score to a small number of discrete tiers (e.g. Tier 1 "mission-critical," Tier 2 "important," Tier 3 "best-effort"), each with its own SLA/SLO target, measurement rigor, and alerting posture. Example targets per tier: Tier 1, 99.99% availability, p95 latency tightly bounded, real-time monitoring with immediate paging; Tier 2, 99.9% availability, monitored continuously but with a slightly more relaxed paging threshold; Tier 3, 99% availability, monitored but primarily via periodic review rather than real-time alerting.
Worked example
Pricing/justifying premium tiers to stakeholders: frame the tier's target and its associated engineering investment cost TOGETHER, e.g. "Tier 1 status for this service requires multi-region redundancy and a dedicated on-call rotation, at an estimated ongoing cost of $X/year; given this service represents $Y in direct monthly revenue and serves Z% of our customer base, that investment is justified," making the cost-benefit explicit rather than treating tier assignment as a status symbol services compete for informally. Monitoring requirements scale with tier: Tier 1 services get the fullest measurement rigor (multiple SLIs, real-time dashboards, dedicated on-call); Tier 3 services get a lighter-weight, less resource-intensive monitoring setup proportional to their lower business stakes.
Trade-offs and pitfalls
The most common failure in a tiering framework like this is TIER INFLATION: without a disciplined, evidence-based scoring process, teams reasonably want their OWN service classified at the highest tier (for prestige, resourcing, or genuine belief in its importance), and a framework that doesn't push back with concrete criteria eventually ends up with most services claiming Tier 1 status, which defeats the entire purpose of tiering (differentiated investment proportional to actual business impact). A periodic re-scoring process is also necessary, since a service's business impact genuinely changes over time (a service that started as a minor internal tool might become business-critical after an unplanned dependency grows around it), and a tier assigned once at launch and never revisited will eventually misrepresent the service's actual current importance.
Given the table schema service_metrics(service_name TEXT, ts TIMESTAMP, success BOOLEAN), write an SQL query (any dialect) to compute the 30-day error budget burn rate per service where SLO = 99.95%. Output: service_name, total_requests, observed_errors, allowed_errors, burn_rate. Explain assumptions (time window alignment, nulls) and how you would run this efficiently on large datasets.
Sample Answer
A per-service burn-rate report needs the allowed-errors figure computed from each service's OWN request volume, not a single fleet-wide number, since services with different traffic naturally have different absolute error budgets even at the same SLO percentage.
Structured elaboration
For each service, total requests and observed errors are aggregated over the 30-day window, and the allowed-errors figure is simply the SLO's allowed-bad-fraction multiplied by that service's own total request count; burn rate is then observed errors divided by allowed errors, giving a comparable, unitless number across services regardless of how much traffic each one handles.
Worked example (executed; sqlite3, two services with different traffic and error counts)
WITH agg AS (
SELECT service_name,
COUNT(*) AS total_requests,
SUM(CASE WHEN success = 0 THEN 1 ELSE 0 END) AS observed_errors
FROM service_metrics
WHERE ts >= date('2026-07-01') AND ts < date('2026-07-01', '+30 day')
GROUP BY service_name
)
SELECT service_name, total_requests, observed_errors,
total_requests * 0.0005 AS allowed_errors,
ROUND(observed_errors * 1.0 / NULLIF(total_requests * 0.0005, 0), 2) AS burn_rate
FROM agg;
Against a synthetic dataset with checkout (2 requests, 0 errors) and images (4 requests, 1 error) at a 99.95% SLO: checkout correctly returned a burn rate of 0.0, and images returned 500.0 (1 observed error against an allowed-errors figure of 4 * 0.0005 = 0.002, meaning even a single error on a very small sample massively exceeds the allowed rate), correctly illustrating why burn rate on a low-traffic service needs a minimum sample-size guard before it's trusted.
Trade-offs and pitfalls
The NULLIF guard prevents a divide-by-zero for a service with zero requests in the window, but it does not solve the more subtle problem the images result illustrates: at very small request counts, a tiny absolute number of errors produces an enormous, statistically unstable burn-rate multiple that looks alarming but reflects noise more than a real trend; a production version should suppress or flag burn-rate values for services below a minimum request-count threshold rather than reporting them at face value.
Design a lightweight SLA enforcement mechanism that prevents teams from repeatedly violating an enterprise-wide data ingestion SLO. Include detection, enforcement actions, and incentive-aligned remediation steps.
Sample Answer
A lightweight enforcement mechanism for a shared, enterprise-wide data-ingestion SLO needs detection that's cheap to run continuously, an escalating consequence that's proportional to the severity of repeated violation, and incentives that make compliance the path of least resistance rather than a burden teams route around.
Structured elaboration
Detection: a simple, automated periodic check (e.g. daily) comparing each team's ingestion pipeline against the SLO's defined freshness/completeness thresholds, flagging any miss and, critically, tracking the PATTERN of misses over time (a single miss is normal noise; a team missing repeatedly across several consecutive periods is the actual enforcement target). Enforcement actions, escalating: a first miss generates an automated notification to the owning team, no further consequence; a SECOND consecutive miss triggers a required written remediation plan from the team (a low-cost but real accountability step); a THIRD consecutive miss (a genuinely repeated pattern) escalates to that team's engineering leadership and may trigger a more structural consequence (e.g. that team's ingestion pipeline being deprioritized for new feature requests until compliance is restored). Incentive alignment: rather than a purely punitive mechanism, pair enforcement with a positive incentive (e.g. teams with a strong compliance record get priority access to shared platform support or infrastructure investment), so the mechanism isn't purely a stick.
Worked example
A team's ingestion pipeline misses its freshness SLO on Monday (first miss): automated Slack notification to the team, no other action. It misses again Tuesday and Wednesday (second and third consecutive miss): the team is required to submit a one-page remediation plan within 48 hours, reviewed by a lightweight cross-team governance group. If the pattern continues into a fourth consecutive week despite a submitted remediation plan, the team's other feature requests to the shared platform team are deprioritized until the ingestion issue is demonstrably resolved, a concrete, felt consequence that creates real urgency without requiring a heavyweight enforcement bureaucracy.
Trade-offs and pitfalls
A purely punitive mechanism with no positive incentive tends to produce resentment and gaming (teams finding ways to technically avoid triggering the detection without genuinely fixing the underlying problem) rather than genuine compliance; pairing enforcement with a real, valued incentive for good behavior (priority support, infrastructure investment) changes the calculus toward genuine cooperation. It's also worth keeping the detection and escalation genuinely LIGHTWEIGHT as intended: a heavyweight review process for every single miss would itself become a source of friction and resistance, undermining the "lightweight" design goal the mechanism is explicitly meant to satisfy.
Unlock Full Question Bank
Get access to all 24 Site Reliability Engineering Principles interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.