Site Reliability Engineering Principles Questions

The core SRE practice model: service-level objectives and indicators, error budgets, toil reduction, and reliability as an engineering discipline. Covers the principles and trade-offs behind treating operations as a software problem and balancing reliability against feature velocity. The conceptual foundation questions specific to SRE-style roles.

HardTechnical
81 practiced

Your service consists of five serially-dependent services (A → B → C → D → E) and the end-to-end availability SLO is 99.95%. Propose how to apportion availability targets across these services, show the math converting per-service availability to an end-to-end SLO, and explain how you would detect which service contributes most to end-to-end failures and how to set per-service error budgets.

MediumTechnical
119 practiced

Given Prometheus counters: http_requests_total{service="images",code="200"} and http_requests_total{service="images"}, write a PromQL expression that computes the 5-minute availability SLI (successful requests / total requests) for service 'images' and a rule that fires an alert if availability < 99.95% for 15 minutes. Explain your use of functions and evaluation interval.

EasyTechnical
138 practiced

Explain the differences between SLI, SLO, and SLA. Provide a concrete example for a customer-facing REST API: specify one SLI (metric and units), an SLO target (with measurement window), and a sample SLA clause suitable for a contract. Describe how you would operationalize the SLO (measurement, dashboards, alerting) and how error budget policies would influence release velocity and incident response.

HardTechnical
87 practiced

Write a PromQL expression to compute a rolling 1-hour burn rate for an SLO based on a 'success_rate' metric, and write an alert that fires if burn rate > 2 for 30 minutes. Explain how burn rate is calculated from SLI time-series and justify the window lengths and threshold chosen.

HardTechnical
137 practiced

How do you test and validate that an SLI truly reflects customer experience and is not biased by instrumentation or sampling? Provide a validation plan including synthetic tests, correlating real-user telemetry (RUM), shadow traffic and acceptance criteria.

Unlock Full Question Bank

Get access to all 24 Site Reliability Engineering Principles interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.