SLIs, SLOs, SLAs, and Error Budgets Questions
Defining and operating reliability targets. Covers choosing service level indicators, setting service level objectives and agreements, computing and spending error budgets, and using them to drive engineering decisions. Includes negotiating reliability targets with stakeholders.
What is an 'error budget policy'? Give two concrete, real-world examples of policies (one automated and one manual) describing thresholds, actors, and actions. Explain why each example helps balance velocity with reliability.
Sample Answer
An error budget policy is the pre-agreed set of rules for what actually happens once budget consumption crosses specific thresholds, turning "we have an error budget" from an abstract concept into concrete, predictable action.
Structured elaboration
A good policy names the threshold, who is notified or empowered to act, and exactly what they do or are authorized to do; it needs both an AUTOMATED tier (fast, mechanical, no human judgment required) and a MANUAL tier (requires a person to weigh context, because the automated response alone isn't sufficient at that severity). Automated example: for a 99.9%-SLO service, if more than 50% of the WEEKLY error budget allocation is consumed, the CI/CD pipeline automatically restricts merges to hotfix-only, no human approval needed to trigger this, though an override exists for genuine emergencies. Manual example: if the error budget is fully exhausted for an entire QUARTER (a sustained, structural reliability problem rather than a one-off incident), engineering leadership and the product owner hold a joint review to reallocate the next quarter's roadmap toward reliability work, since the situation has moved from a tactical release-freeze to a strategic resourcing decision.
Worked example
Applying this to a data-engineering ingestion team illustrates the same mechanism in a different domain: an ingestion pipeline with a 99.9% freshness SLO uses its error budget the same way a request-serving API would, to decide the trade-off between shipping new pipeline features and investing in reliability fixes; when the pipeline's budget burns past the weekly threshold, new feature work pauses automatically (via the same CI gate pattern) in favor of fixing the freshness regression first, and if the budget stays exhausted across a full quarter, the team's roadmap gets the same leadership-level reallocation review as any other service.
Trade-offs and pitfalls
A policy that only defines automated actions without a manual, strategic tier will keep re-triggering the same tactical response (freeze, unfreeze, freeze again) without ever addressing a structural reliability problem; conversely a policy with only manual review and no automated tier is too slow to prevent damage during an active incident. The two tiers are solving different problems (fast tactical response vs slow structural correction) and a mature policy needs both, explicitly distinguished by which threshold triggers which.
Design SLAs and SLOs for an internal platform consumed by 30 teams and operated by a platform team. Define appropriate SLO targets (availability, latency tiers), monitoring and alerting strategy, reporting cadence, support response expectations, remediation steps for missed SLOs, and cultural or contractual enforcement mechanisms to ensure accountability.
Sample Answer
An internal platform serving 30 teams needs SLA-like rigor even without external money changing hands, because the accountability mechanism that makes an SLA work (a real consequence for missing the target) has to exist internally too, or the SLO becomes advisory rather than binding.
Structured elaboration
SLO targets, tiered by criticality (a shared auth service probably needs 99.95%+, a lower-traffic internal reporting tool might be fine at 99%); monitoring and alerting owned by the platform team, with dashboards visible to all 30 consuming teams so status isn't opaque; reporting cadence (e.g. a monthly reliability report circulated to all consuming teams' leads); support response expectations (e.g. P1 acknowledged within 15 minutes, P3 within one business day) that function like an internal support SLA. Since there's no direct revenue on the line, roadmap-level accountability is the real enforcement mechanism: sustained SLO misses should trigger a documented reallocation of the platform team's OWN roadmap toward reliability work, visible to and reviewed by the consuming teams' leadership, rather than being a purely internal, invisible metric with no real consequence.
Worked example
An internal feature-flag platform serving 30 teams sets a tiered target: 99.95% availability with p95 latency < 50ms for its critical, synchronous-lookup tier, versus 99% availability with p95 latency < 500ms for a lower-traffic, non-blocking reporting tier, so the latency commitment scales with criticality the same way availability does, not a single flat number applied regardless of tier. Monthly reports go to all consuming teams; if the SLO is missed two months running, the platform team's next-quarter roadmap must allocate at least 30% of capacity to reliability work, reviewed and confirmed by an internal steering group representing the consuming teams, not decided unilaterally by the platform team itself. For a platform specifically supporting ML workflows, this same mechanism additionally includes a documented ONBOARDING process (new ML teams joining the platform are shown the current SLO status and support tiers up front) and ties platform-improvement PRIORITIZATION directly to which internal SLA commitments are most at risk, so onboarding and roadmap decisions are grounded in the same accountability data rather than treated as a separate, informal process.
Trade-offs and pitfalls
Without a genuine consequence for missing the target, an internal SLO quietly becomes theater: teams cite it in retros but nothing structurally changes when it's repeatedly missed, which is why the roadmap-reallocation mechanism needs external visibility (to the consuming teams, not just the platform team's own management) to have teeth. Incentive design matters here too: if the platform team's own leadership is evaluated purely on feature delivery velocity with no weight given to SLO adherence, the roadmap-reallocation mechanism will face constant internal pressure to be deprioritized, so the incentive structure for the platform team's OWN leadership needs to genuinely reward reliability investment, not just tolerate it.
Design how an SLI representing user-perceived performance (e.g., time-to-interactive) can be captured, aggregated, and turned into an SLO. Discuss sampling, privacy, instrumentation on mobile/web, and correlation to backend SLIs.
Sample Answer
Capturing user-perceived performance like time-to-interactive means instrumenting on the CLIENT (browser or mobile app), which introduces sampling, privacy, and correlation challenges that a purely backend SLI never has to deal with.
Structured elaboration
Time-to-interactive (or similar) is typically captured via the browser's Navigation Timing / User Timing APIs on web, or platform-specific instrumentation (e.g. app-start-to-first-frame timers) on mobile, then reported back to a collection endpoint, usually sampled (not every single page load, for both cost and privacy reasons) rather than captured exhaustively. Aggregating this into an SLO requires the same percentile-computation discipline as any latency SLI (histogram or sketch-based, since raw per-event storage at scale is expensive), but with an added correlation step: joining the client-observed timing with backend request IDs (via a trace context propagated end-to-end) lets you attribute a slow client experience to a specific backend dependency, rather than treating client and backend telemetry as two disconnected numbers.
Worked example
A mobile app samples time-to-interactive from 10% of sessions (to bound both bandwidth and storage cost), tagging each sample with a trace ID that correlates to the backend request(s) that populated that screen. If p95 time-to-interactive degrades from 1.2s to 2.5s, the trace correlation reveals whether the regression traces back to a specific backend API's latency (in which case the backend SLI should already have caught it) or is purely client-side (a rendering regression, a bundle-size increase, a slow third-party script), which the backend SLIs would never surface on their own.
Trade-offs and pitfalls
Privacy constraints matter more here than for typical backend telemetry: client-side timing data can sometimes be combined with other signals to fingerprint or de-anonymize a user, so sampling rate, retention period, and what additional context gets attached to each sample all need review against your privacy policy and applicable regulation, not just an engineering cost trade-off. Sampling itself introduces bias risk identical to the client-vs-server-instrumentation problem: if the sampling mechanism itself is more likely to succeed on fast, stable connections (a flaky connection might fail to deliver its own telemetry beacon), the reported p95 will systematically understate the true tail experienced by users on worse connections.
Define Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for a support ticketing system focused on response and resolution. Propose three alerting thresholds that would escalate from an on-call page to a stakeholder email, and explain how to set each threshold to minimize alert fatigue.
Sample Answer
A support ticketing system's SLIs need to center on the two things that actually matter to a person waiting for help: how fast the first response comes, and how fast the underlying issue actually gets resolved, and the alerting tiers should escalate proportional to how badly those two things are slipping.
Structured elaboration
SLIs: response time (time from ticket creation to first meaningful human response) and resolution time (time from creation to actual issue closure), both typically segmented by ticket priority/severity, since a P1 outage-blocking ticket and a P4 minor-cosmetic-question ticket have entirely different reasonable expectations. SLOs might target: P1 tickets, 95% first-responded within 15 minutes and 90% resolved within 4 hours; P3 tickets, 95% first-responded within 4 business hours and 90% resolved within 3 business days.
Worked example
Three escalation thresholds moving from an on-call page to a stakeholder email: tier 1 (early warning, no page): a single ticket approaching its response-time SLA (e.g. at 80% of its allotted response window with no response yet) generates an internal reminder to the assigned agent, no escalation beyond that. Tier 2 (page on-call): a P1 ticket that has actually BREACHED its response-time SLA with no acknowledgment pages the on-call support lead directly, since a P1 miss is both rare and urgent enough to warrant immediate human attention. Tier 3 (stakeholder email): a SUSTAINED pattern (e.g. more than 10% of all tickets in a rolling 7-day window missing their SLA, not just a single ticket) triggers an email to the support team's leadership, since a single slow ticket is normal noise but a sustained pattern indicates a genuine staffing or process problem worth leadership visibility.
Trade-offs and pitfalls
Paging on every single SLA-approaching ticket (rather than reserving paging for genuinely urgent, rare breaches like a missed P1) would produce exactly the alert-fatigue problem seen elsewhere in high-volume ticketing systems, where dozens of routine near-misses happen daily; the tiered design here reserves the most urgent escalation channel (a page) for the rarest, highest-stakes miss, and uses a slower, less disruptive channel (an aggregate weekly email) for a pattern that matters but isn't individually urgent. It's also worth periodically re-validating that the specific SLA thresholds (15 minutes, 4 hours, etc.) still reflect genuine customer expectations and team capacity, rather than being numbers set once early on and never revisited as ticket volume or team size changes.
Draft an error budget borrowing and repayment policy between teams, including allowed borrow amounts, repayment time windows, visibility requirements, and penalties or incentives. Provide examples of compliant and non-compliant borrowing scenarios.
Sample Answer
Error-budget borrowing lets a team spend AHEAD of its normal allocation for a specific, justified reason (a planned risky migration, a time-sensitive launch), as long as it's repaid, and the policy's whole job is making sure "borrowing" doesn't quietly become "just not tracking budget properly."
Structured elaboration
Allowed borrow amounts: capped at a modest fraction of the team's NEXT period's allocation (e.g. up to 20% of next month's budget can be borrowed against this month), preventing a team from borrowing so much that they're effectively operating with zero real margin for an extended stretch. Repayment time windows: a defined, short horizon (e.g. borrowed budget must be repaid, meaning the team operates with EXTRA caution and reduced risk-taking, within the following month), not an open-ended IOU that can persist indefinitely. Visibility requirements: every borrow event is logged in the same shared, auditable registry used for other governance actions, visible to any team or stakeholder who might reasonably want to know a team is currently operating on borrowed budget. Penalties or incentives: a team that fails to repay within the agreed window loses borrowing privileges for a defined cooldown period (e.g. two subsequent months), a real but proportionate consequence.
Worked example
Compliant scenario: Team A wants to ship a planned, moderately-risky database migration next week that they've assessed will likely consume more budget than their current month's remaining allocation supports; they formally request to borrow 15% of next month's budget, documented with the specific justification and expected repayment plan, approved via the standard lightweight sign-off, and visible to other teams sharing any affected infrastructure. Non-compliant scenario: Team B has been quietly running over budget for three consecutive months without ever formally requesting to borrow, simply treating each month's overage as if it will "sort itself out," with no visibility to anyone else and no explicit repayment plan; this is exactly the undisciplined pattern the formal borrowing policy exists to replace, since an UNTRACKED overage gives none of the planning benefit a formally-documented borrow arrangement provides.
Trade-offs and pitfalls
Without a real, felt consequence for failing to repay, "borrowing" degrades into a polite fiction that lets a team perpetually operate over budget while calling it something more palatable; the cooldown-period consequence needs to be genuinely enforced, not waived informally when a team has a good excuse each time. It's also worth deciding whether TWO teams should be allowed to simultaneously borrow against a SHARED downstream dependency's stability, since two individually-approved, modest borrow requests could combine to create a much larger combined risk than either looks like in isolation; a genuinely lightweight policy still needs at least a basic check for this kind of compounding risk across simultaneous borrow requests.
Unlock Full Question Bank
Get access to all SLIs, SLOs, SLAs, and Error Budgets interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.