SLIs, SLOs, SLAs, and Error Budgets Questions
Defining and operating reliability targets. Covers choosing service level indicators, setting service level objectives and agreements, computing and spending error budgets, and using them to drive engineering decisions. Includes negotiating reliability targets with stakeholders.
Draft an error budget borrowing and repayment policy between teams, including allowed borrow amounts, repayment time windows, visibility requirements, and penalties or incentives. Provide examples of compliant and non-compliant borrowing scenarios.
Sample Answer
Error-budget borrowing lets a team spend AHEAD of its normal allocation for a specific, justified reason (a planned risky migration, a time-sensitive launch), as long as it's repaid, and the policy's whole job is making sure "borrowing" doesn't quietly become "just not tracking budget properly."
Structured elaboration
Allowed borrow amounts: capped at a modest fraction of the team's NEXT period's allocation (e.g. up to 20% of next month's budget can be borrowed against this month), preventing a team from borrowing so much that they're effectively operating with zero real margin for an extended stretch. Repayment time windows: a defined, short horizon (e.g. borrowed budget must be repaid, meaning the team operates with EXTRA caution and reduced risk-taking, within the following month), not an open-ended IOU that can persist indefinitely. Visibility requirements: every borrow event is logged in the same shared, auditable registry used for other governance actions, visible to any team or stakeholder who might reasonably want to know a team is currently operating on borrowed budget. Penalties or incentives: a team that fails to repay within the agreed window loses borrowing privileges for a defined cooldown period (e.g. two subsequent months), a real but proportionate consequence.
Worked example
Compliant scenario: Team A wants to ship a planned, moderately-risky database migration next week that they've assessed will likely consume more budget than their current month's remaining allocation supports; they formally request to borrow 15% of next month's budget, documented with the specific justification and expected repayment plan, approved via the standard lightweight sign-off, and visible to other teams sharing any affected infrastructure. Non-compliant scenario: Team B has been quietly running over budget for three consecutive months without ever formally requesting to borrow, simply treating each month's overage as if it will "sort itself out," with no visibility to anyone else and no explicit repayment plan; this is exactly the undisciplined pattern the formal borrowing policy exists to replace, since an UNTRACKED overage gives none of the planning benefit a formally-documented borrow arrangement provides.
Trade-offs and pitfalls
Without a real, felt consequence for failing to repay, "borrowing" degrades into a polite fiction that lets a team perpetually operate over budget while calling it something more palatable; the cooldown-period consequence needs to be genuinely enforced, not waived informally when a team has a good excuse each time. It's also worth deciding whether TWO teams should be allowed to simultaneously borrow against a SHARED downstream dependency's stability, since two individually-approved, modest borrow requests could combine to create a much larger combined risk than either looks like in isolation; a genuinely lightweight policy still needs at least a basic check for this kind of compounding risk across simultaneous borrow requests.
Define SLOs, SLIs and an error-budget policy for a core payment-processing API. Explain how these should map to CI quality gates, what metrics to collect and alert on, how to detect gradual regressions, and what automated or manual actions to take when the error budget is exhausted.
Sample Answer
A payment-processing API needs SLOs anchored on user-visible correctness and latency, an error-budget policy that ties directly into automated release gates, and detection tuned for slow drift, not just hard failures.
Structured elaboration
Core SLIs: request success rate (non-5xx and no payment-specific soft-failure code), p95/p99 latency, and a THIRD signal specific to payments: the rate of transactions that time out ambiguously (neither confirmed success nor clean failure), since that state is the one that actually causes customer and support harm. SLO targets might be 99.95% success over 30 days, p95 < 300ms, ambiguous-timeout rate < 0.01%. The error-budget policy maps burn-rate thresholds to CI/CD gate actions: healthy budget allows normal deploy velocity; moderate burn (2-4x sustainable rate) requires an approved canary-only rollout; budget exhaustion blocks all non-critical releases automatically via the same pipeline check used for any other release gate. Gradual regressions (a slow latency creep rather than a sudden spike) are the hardest to catch with a single-window alert, so a longer trailing-window comparison (this week's p95 vs the prior 4-week baseline) should run alongside the fast burn-rate alert.
Worked example
As a concrete illustration with different numeric targets: three SLIs for a checkout endpoint might be set at 99.9% success over 30 days, p95 latency < 250ms, and a burn-rate alert firing at 2x sustained over 1 hour with a page at 10x sustained over 5 minutes. When burn rate crosses 2x, the CI/CD pipeline automatically restricts merges to hotfix-only until an SRE manually clears the gate; at complete budget exhaustion, only security patches are permitted through, mirroring the same policy pattern used for lower-throughput internal-only platforms and even a simpler image-serving API, just with different absolute numbers appropriate to each service's criticality and traffic (an image-serving API might need only two SLIs, availability and latency, where a payment API needs the third, ambiguous-outcome signal).
Trade-offs and pitfalls
Mapping error budget directly into an automated CI gate is powerful but risky if the gate has no override path for a genuine emergency (a security patch that must ship regardless of budget state); every automated gate needs a logged, audited manual override. On staffing implications: a service with an aggressive SLO and a strict error-budget gate needs on-call coverage sized to actually respond within the alerting windows chosen, not just a policy on paper; setting a 5-minute page threshold with no one actually able to respond within that window is a policy that looks rigorous but isn't enforceable.
Your company must produce auditable SLA reports to customers and regulators. Design the data retention, immutability, and reporting pipeline so that SLA measurements are tamper-evident and reproducible. Include backup, timezone handling, and legal considerations.
Sample Answer
Auditable SLA reporting needs the underlying data pipeline treated with the same rigor as a financial reporting system: immutable, reproducible, and legally defensible, since a regulator or a disputing customer will eventually ask exactly how a number was produced.
Structured elaboration
Data retention and immutability: raw measurement data (not just the computed monthly percentage) should be retained for the full period any SLA dispute could reasonably be raised (often driven by contractual or regulatory requirements, commonly several years), stored in an APPEND-ONLY or otherwise tamper-evident form (e.g. write-once storage, or a cryptographic hash chain over sequential report batches) so a later modification to historical data would be detectable rather than silently possible. Reporting pipeline: the computation from raw data to the published SLA percentage needs to be fully REPRODUCIBLE, meaning the exact query/computation logic used for a given historical report is itself versioned and retained, so re-running the same computation against the same retained raw data later produces an identical result, not a different one due to a since-changed query.
Worked example
Timezone handling: every raw timestamp needs to be stored with an explicit, unambiguous timezone (UTC is the standard choice) rather than a local time that could be ambiguous around a daylight-saving transition; a report boundary defined as "midnight local time" needs the LOCAL timezone convention documented explicitly and applied consistently, since inconsistent timezone handling is a classic, embarrassing source of an SLA dispute where two parties' independently-computed numbers disagree purely due to a boundary-handling bug, not a real measurement discrepancy. Backup: the raw retained data itself needs backup with the same immutability guarantee as the primary copy, since a backup that CAN be silently altered defeats the tamper-evidence of the primary.
Trade-offs and pitfalls
Legal considerations often drive the retention period and evidentiary format more than pure engineering preference; the specific requirements (how long to retain, what counts as sufficiently tamper-evident, what documentation a regulator would accept) should be confirmed with legal/compliance BEFORE the pipeline is built, not retrofitted after a dispute reveals the existing pipeline doesn't meet the actual evidentiary bar required. It's also worth being explicit that "reproducible" means genuinely re-running the ORIGINAL computation logic against the ORIGINAL raw data, not just re-generating a similar-looking report from today's (possibly since-updated) query logic against the same historical data, which could silently produce a different number even with unchanged raw inputs.
What is an 'error budget policy'? Give two concrete, real-world examples of policies (one automated and one manual) describing thresholds, actors, and actions. Explain why each example helps balance velocity with reliability.
Sample Answer
An error budget policy is the pre-agreed set of rules for what actually happens once budget consumption crosses specific thresholds, turning "we have an error budget" from an abstract concept into concrete, predictable action.
Structured elaboration
A good policy names the threshold, who is notified or empowered to act, and exactly what they do or are authorized to do; it needs both an AUTOMATED tier (fast, mechanical, no human judgment required) and a MANUAL tier (requires a person to weigh context, because the automated response alone isn't sufficient at that severity). Automated example: for a 99.9%-SLO service, if more than 50% of the WEEKLY error budget allocation is consumed, the CI/CD pipeline automatically restricts merges to hotfix-only, no human approval needed to trigger this, though an override exists for genuine emergencies. Manual example: if the error budget is fully exhausted for an entire QUARTER (a sustained, structural reliability problem rather than a one-off incident), engineering leadership and the product owner hold a joint review to reallocate the next quarter's roadmap toward reliability work, since the situation has moved from a tactical release-freeze to a strategic resourcing decision.
Worked example
Applying this to a data-engineering ingestion team illustrates the same mechanism in a different domain: an ingestion pipeline with a 99.9% freshness SLO uses its error budget the same way a request-serving API would, to decide the trade-off between shipping new pipeline features and investing in reliability fixes; when the pipeline's budget burns past the weekly threshold, new feature work pauses automatically (via the same CI gate pattern) in favor of fixing the freshness regression first, and if the budget stays exhausted across a full quarter, the team's roadmap gets the same leadership-level reallocation review as any other service.
Trade-offs and pitfalls
A policy that only defines automated actions without a manual, strategic tier will keep re-triggering the same tactical response (freeze, unfreeze, freeze again) without ever addressing a structural reliability problem; conversely a policy with only manual review and no automated tier is too slow to prevent damage during an active incident. The two tiers are solving different problems (fast tactical response vs slow structural correction) and a mature policy needs both, explicitly distinguished by which threshold triggers which.
Design how an SLI representing user-perceived performance (e.g., time-to-interactive) can be captured, aggregated, and turned into an SLO. Discuss sampling, privacy, instrumentation on mobile/web, and correlation to backend SLIs.
Sample Answer
Capturing user-perceived performance like time-to-interactive means instrumenting on the CLIENT (browser or mobile app), which introduces sampling, privacy, and correlation challenges that a purely backend SLI never has to deal with.
Structured elaboration
Time-to-interactive (or similar) is typically captured via the browser's Navigation Timing / User Timing APIs on web, or platform-specific instrumentation (e.g. app-start-to-first-frame timers) on mobile, then reported back to a collection endpoint, usually sampled (not every single page load, for both cost and privacy reasons) rather than captured exhaustively. Aggregating this into an SLO requires the same percentile-computation discipline as any latency SLI (histogram or sketch-based, since raw per-event storage at scale is expensive), but with an added correlation step: joining the client-observed timing with backend request IDs (via a trace context propagated end-to-end) lets you attribute a slow client experience to a specific backend dependency, rather than treating client and backend telemetry as two disconnected numbers.
Worked example
A mobile app samples time-to-interactive from 10% of sessions (to bound both bandwidth and storage cost), tagging each sample with a trace ID that correlates to the backend request(s) that populated that screen. If p95 time-to-interactive degrades from 1.2s to 2.5s, the trace correlation reveals whether the regression traces back to a specific backend API's latency (in which case the backend SLI should already have caught it) or is purely client-side (a rendering regression, a bundle-size increase, a slow third-party script), which the backend SLIs would never surface on their own.
Trade-offs and pitfalls
Privacy constraints matter more here than for typical backend telemetry: client-side timing data can sometimes be combined with other signals to fingerprint or de-anonymize a user, so sampling rate, retention period, and what additional context gets attached to each sample all need review against your privacy policy and applicable regulation, not just an engineering cost trade-off. Sampling itself introduces bias risk identical to the client-vs-server-instrumentation problem: if the sampling mechanism itself is more likely to succeed on fast, stable connections (a flaky connection might fail to deliver its own telemetry beacon), the reported p95 will systematically understate the true tail experienced by users on worse connections.
Unlock Full Question Bank
Get access to all SLIs, SLOs, SLAs, and Error Budgets interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.