On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
How would you design an on-call model that shares ownership fairly between product engineers and the SRE team, without burning either group out?
Sample Answer
Fair, sustainable shared on-call between product engineers and SRE requires splitting the pager by what kind of problem it is, not by headcount: product engineers own their service's application-level pages because they wrote the code, and SRE owns cross-cutting platform incidents plus acts as an escalation and coaching layer, not the default first responder for everything.
Model design
- Two-tier ownership by problem type. Primary on-call for a service's application-level alerts (bad deploy, application error spikes, business-logic bugs) sits with that service's own product engineering team, since they have the fastest path to a fix. Primary on-call for platform-level alerts (cluster capacity, networking, shared infrastructure) sits with SRE, since product engineers generally can't fix those even if paged.
- SRE as secondary/escalation, not default first responder. If a product engineer's on-call shift can't resolve something within a set time, or the issue turns out to be platform-level, it escalates to SRE. This keeps SRE from being paged for every application bug (which is what burns SRE out) while still giving product engineers a safety net.
- Bound the rotation load per person, not just per team. A team of 4 engineers rotating weekly means each person is on-call roughly once a month; a team of 8 halves that frequency. If a team is too small to keep individual on-call frequency reasonable, that's a staffing gap the model should surface, not something the rotation schedule should quietly absorb by making people on-call more often.
- Require a shadow period before solo on-call. New product engineers pair with an experienced on-call responder (from their own team or SRE) for at least one full rotation before taking a shift alone, so the first real page isn't also someone's first exposure to the runbooks and escalation path.
flowchart TD
A[Alert fires] --> B{Application-level or platform-level?}
B -->|Application-level| C[Product Engineer primary]
B -->|Platform-level| D[SRE primary]
C --> E{Resolved within SLA?}
E -->|Yes| F[Closed by product team]
E -->|No| G[Escalate to SRE secondary]
G --> H[SRE helps diagnose]
H --> I[Fix stays owned by product team]
D --> J[SRE resolves directly]
Runbook and readiness split
Product teams own and write their own service's runbooks (they have the context), but SRE owns the runbook template and reviews new runbooks before a service starts taking pages, catching the common gaps (missing verification steps, no rollback path) before they're discovered live. SRE separately owns and maintains the platform-level runbooks (networking, cluster operations) that no single product team should be expected to know.
Worked example
A team of 6 product engineers owns a checkout service. Under this model: one product engineer is primary on-call each week (a six-week gap between an individual's rotations, comfortably sustainable), and SRE's on-call rotation covers checkout along with roughly 8 other services as secondary/platform on-call. A deploy causes checkout to start returning errors on one specific endpoint: this pages the product engineer primary, who has the code context to identify and roll back the bad deploy in the first 15 minutes, no SRE involvement needed. Separately, a shared load balancer starts dropping connections across checkout and 4 other services: this pages SRE primary directly (it's platform-level, tagged as such by the alerting routing, not first-responded by any product team), and SRE resolves it without waking every affected product team's on-call. If the checkout product engineer had instead been unable to diagnose the deploy issue within 20 minutes, the runbook's escalation step would have paged SRE as secondary, who can help diagnose but explicitly hands ownership of the actual fix back to the product team once the platform is ruled out, rather than SRE quietly absorbing the fix and becoming the default responder for checkout's application bugs going forward.
Trade-offs and pitfalls
The main trade-off is that product engineers now carry pager load they didn't have before, which is a real cost to their focus time; the return is that ownership drives better code (engineers who get paged for their own bugs tend to write fewer of them) and it's the only way to keep SRE from being the permanent bottleneck for every application-level issue across dozens of services. A common wrong turn is letting SRE quietly become the actual first responder for application issues "just this once" repeatedly, because it's faster than waiting for a less-experienced product engineer to work through it; each time that happens it erodes the product team's ownership and pushes the model back toward the SRE-does-everything pattern this design is meant to avoid. The fix is a hard rule: SRE helps diagnose during escalation, but the fix and its follow-up stay owned by the product team unless the root cause is confirmed platform-level.
Design how an alert should connect automatically to the right runbook, including letting a responder execute a pre-approved remediation step with one click. What do you log, and what stops the system from taking an unsafe action on its own?
Sample Answer
Direct answer
Wire the alert directly to a runbook execution layer, not just a link: the alert carries a runbook identifier, the paging tool surfaces the matching runbook inline with a one-click execute action for pre-approved steps, and every click passes through validation, RBAC (role-based access control, the system that decides which actions a given user is permitted to run), and policy evaluation before anything runs, with the result and a full audit entry logged back to the incident. What actually stops an unsafe action isn't the button in the UI, it's the policy layer sitting between the click and the execution.
Structured elaboration
Component flow
sequenceDiagram
participant M as Monitoring
participant PD as PagerDuty
participant R as Responder
participant E as Runbook Executor
participant V as Audit Log
M->>PD: Alert fires with runbook_id attached
PD->>R: Page plus incident link
R->>E: Open runbook, click approved step
E->>E: Validate params, RBAC, policy check
E->>V: Write audit entry: who, what, when
E->>PD: Post execution result to incident
What each layer does
- Alert enrichment: alerting rules carry a runbook identifier and a list of allowed actions as labels, so the alert itself knows which remediation it maps to.
- Paging tool integration: the incident view embeds the runbook with each step clearly labeled manual versus automated, along with its risk level.
- Execution broker: on click, it validates the caller's identity through RBAC, confirms the incident is actually open and acknowledged, and evaluates policy, such as whether this action needs extra approval given current severity or a maintenance window.
- Audit log: append-only, recording who clicked, what parameters were used, what the action did, and the result, written before the action is considered complete rather than best-effort afterward.
What logs, and what actually stops an unsafe action
| Logged | Why |
|---|---|
| Who triggered it, real identity | Accountability and RBAC review |
| Incident ID and state at trigger time | Prevents replay against a stale or closed incident |
| Exact parameters passed | Reproducibility and forensics |
| Approval chain, if required | Proves the human-in-loop gate was actually satisfied |
| Result and outcome class | Feeds back into whether this action should keep auto-executing |
Nothing about the UI stops an unsafe action; the broker re-validates everything server-side even if the button already looked available. Any action tagged destructive requires approval regardless of who clicks it, a circuit breaker disables an action automatically after repeated failures, and execution always runs with short-lived, scoped credentials so even a successful unsafe action is bounded in what it can touch.
Worked example
A "restart cache node" action is clicked from the incident view. The broker checks: is the incident open, yes; is the clicker on the paging on-call roster for this service, yes; is this action tagged low-risk, yes, a single-node restart; is it within the rate limit, zero executions in the last five minutes against a cap of one. All checks pass, the executor runs the restart with a scoped, time-limited credential, and writes the audit entry, who, incident ID, action, parameters, result, before returning success to the incident view.
Trade-offs and pitfalls
- Client-side checks are fine for UX, greying out an unavailable button, but the broker must re-validate everything server-side; the client should never be the actual enforcement point.
- Real-time policy evaluation on every click adds latency to remediation; keep the policy check itself a fast local lookup rather than a call to an external system on the critical path, so safety doesn't undercut automation's main benefit.
- Logging "success" based on the action returning without error, rather than verifying the target metric actually recovered, is a common gap; a runbook step can exit cleanly and still not have fixed anything.
What's the difference between a runbook and a playbook, and when would you reach for one instead of the other?
Sample Answer
Direct answer
A runbook is a fixed set of step-by-step instructions for a known failure mode: it tells you exactly what to type. A playbook is a decision framework for handling an incident more broadly: it tells you how to figure out what to do, who needs to be involved, and which runbook to reach for. Reach for a runbook when you already know the cause and the fix is mechanical; reach for a playbook when you're still diagnosing, coordinating multiple people, or the right response depends on judgment.
Structured elaboration
| Runbook | Playbook | |
|---|---|---|
| Scope | One specific, known failure mode or task | A class of incidents, or the overall response process |
| Format | Linear, prescriptive steps | Decision tree or branching guidance |
| Answers | "What do I type" | "What do I decide, and who do I involve" |
| Typical contents | Preconditions, exact commands, verification steps, rollback | Severity thresholds, roles (IC, comms lead), escalation matrix, links to runbooks |
| Usually owned by | The team that owns the specific service | Incident response leadership or SRE |
| Reviewed when | The underlying system changes | The org's escalation structure or tooling changes |
Minimum fields for each:
- Runbook: title, service, owner, trigger/precondition, required permissions and tools, exact step-by-step commands, verification steps, rollback steps, expected impact, last-reviewed date.
- Playbook: title, scope and severity thresholds, incident commander and stakeholder roles, the decision tree itself, links to the relevant runbooks, communication templates, escalation matrix, last-reviewed date.
Worked example
A database replica's lag exceeds a threshold: this is a runbook. It lists the exact commands to promote a replica, the steps to reconfigure the application's read preference, verification queries to confirm the fix, and the rollback commands if the promotion causes a new problem. Now compare that to a major outage affecting payments: this is a playbook. It guides the incident commander through detecting the actual scope, declaring severity, deciding between routing traffic to a fallback payment path versus draining traffic entirely, coordinating the app, infra, and comms teams, and linking out to the specific runbooks (including the replica-promotion one, if that turns out to be the fix) for whichever technical action the decision tree leads to.
Trade-offs and pitfalls
- Pitfall: writing a "runbook" for something that actually needs judgment, for example "when in doubt, restart the service." That hides a decision a playbook should make explicit, and someone follows it verbatim during exactly the incident it doesn't fit.
- Pitfall: letting a playbook go stale is worse than letting one runbook go stale, because the playbook is what everyone reaches for first during ambiguity; a stale escalation matrix (wrong names or numbers) breaks the whole response, not just one specific fix.
- Trade-off: automating a runbook into a one-click execution is great for high-confidence, low-blast-radius fixes (restarting a stateless service) and dangerous for high-blast-radius ones (promoting a database replica). The more damage a wrong click can do, the more the runbook should require an explicit human confirmation step before executing, not less.
- Both belong in a versioned, reviewed repository rather than an unowned wiki page, with periodic review, and runbooks specifically benefit from occasional dry-run or game-day testing to confirm the steps still work against the current system rather than an outdated one.
A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?
Sample Answer
Direct answer
Since you can't fix the vendor directly, you run three things in parallel: mitigate the blast radius with tools you do control (circuit breakers, cached or degraded responses, feature flags), communicate honestly about something outside your control, and push on the vendor relationship itself through support escalation and, if needed, contractual SLA terms. The trade-offs are mostly about how aggressively to degrade functionality versus how much broken or stale behavior your customers will tolerate in the meantime.
Structured elaboration
| Mitigation | What it buys you | What it costs |
|---|---|---|
| Circuit breaker / fail fast | Stops the vendor's failure from cascading into your own services | Feature becomes fully unavailable, more visible outage |
| Serve cached or stale data | Feature stays visibly "up" for the user | Risk of showing wrong or outdated information |
| Queue and retry with backoff | No data loss, eventual consistency once the vendor recovers | User sees delay; adds retry/backoff complexity |
| Feature-flag off (graceful degrade) | Predictable, pre-tested reduced experience | Only works if the flag and the reduced UX already exist before the outage |
Evidence to gather before contacting vendor support. Precise timestamps with timezone noted, representative request/response examples (method, URL, headers, correlation IDs), correlated logs from your own edge/load-balancer and application layers, and a clear scope-and-impact statement (which services, what percentage of traffic, which customers, what SLA is at risk). Vague "your API seems down" tickets sit in a generic queue; a ticket with reproducible evidence and a quantified impact gets triaged faster.
Escalating through vendor support tiers. Open the highest applicable severity case with the evidence attached and explicitly request an engineer and a bridge, not just an acknowledgment. If there's no meaningful response within your own internal SLA for that severity, escalate through the account manager or a phone-based escalation path, citing the specific business impact and contractual SLA terms rather than repeating the original ticket.
Communication cadence, using the same severity-driven pattern as an internal incident: acknowledge to affected customers quickly with what's known and any workaround, then update on a fixed cadence (for example every 30 minutes) until resolved, closing with a summary once the vendor confirms the fix.
Worked example
A payments provider starts returning errors for a subset of transactions. Mitigation: flip a feature flag that routes non-critical calls to a queued-retry path with a "processing" state shown to the user, instead of failing checkout outright; this preserves the customer experience for the subset of traffic where a short delay is tolerable, while transactions that genuinely require a synchronous response fail fast with a clear error rather than hanging. Communication: post an initial status update within roughly 15 minutes acknowledging degraded checkout with the workaround in place, then update every 30 minutes. Vendor escalation: open a high-severity vendor ticket with timestamped request/response examples and the affected transaction volume, request a bridge; if no vendor engineer engages within your internal escalation window, escalate via the account manager's phone line, citing the contractual SLA and quantified customer impact.
Trade-offs and pitfalls
- Pitfall: treating a vendor outage as "not our incident" and skipping the postmortem. Root cause may be external, but your own blast-radius design (whether a circuit breaker or cached fallback existed at all) is exactly what a postmortem should examine, since that's the part you actually control.
- Pitfall: promising customers a fix ETA you don't control. Communicate "investigating, using workaround X, next update in 30 minutes" rather than a timeline that depends on someone else's incident response.
- Trade-off: aggressive circuit-breaking protects your own systems fastest but produces the most visible outage; cached/degraded responses are gentler on the user experience but carry a correctness risk if the vendor's data changes underneath the cache. Which one is right depends on how stale or wrong data is allowed to be for that specific feature, which is a product decision, not just an engineering one.
How would you get a new engineer ready to join the on-call rotation? Walk through what you'd want them to do before their first solo shift.
Sample Answer
Readiness is a checklist, not a countdown: get access and tooling working first, have them study and sign off on the runbooks for their services, shadow several live pages, then run one supervised tabletop and one supervised live (or simulated) incident before they take a shift alone with a mentor reachable but not present.
A four-week ramp
| Week | Focus | Activities | Exit criteria |
|---|---|---|---|
| 1 | Access + orientation | Provision accounts/VPN/MFA/pager, architecture overview, assigned runbook study | All access verified, runbooks read |
| 2 | Guided practice | Shadow 3-4 live alerts with a mentor, pair on small remediation tickets | Runbook sign-offs for owned services |
| 3 | Increasing autonomy | Lead a staging fault-injection drill with mentor observing, handle 1-2 small solo operational tasks with review | Drill led successfully, gaps found in runbooks fixed |
| 4 | Supervised solo shift | First on-call shift with mentor reachable, pre-shift briefing and post-shift debrief | Mentor sign-off, at least one incident handled or correctly escalated |
Sign-off checklist before the first unsupervised shift
- Access confirmed end to end (paging tool, dashboards, deploy/rollback permissions) with a real test, not just "provisioned."
- Runbooks for their assigned services reviewed and any ambiguous steps flagged and fixed.
- At least one supervised tabletop and one supervised live or injected-fault incident completed.
- Mentor sign-off plus the engineer's own confidence self-assessment, not mentor judgment alone.
Extending the ramp for a complex or high-stakes service
Four weeks is often enough for a straightforward service, but for a complex hybrid-cloud system or one with many downstream dependents, add explicit competency checkpoints tied to named systems, for example certifying someone independently on the database failover path as a separate sign-off from general on-call readiness, rather than declaring them ready across the board at once. Longer term, treat the first 90 days as a structured mentorship arc rather than stopping at week four: scheduled 30/60/90-day check-ins, a second mentor pairing on a different service, and a distinct milestone for graduating from "supervised" to "primary" status rather than just a date on the calendar.
Trade-offs and pitfalls
Rushing readiness to fill a rotation gap is the most common failure mode, and it produces confident-sounding but wrong incident responses, which is worse than an obviously under-prepared response because it takes longer to catch. A checklist with no live-incident component only validates that someone can read, not that they can act under time pressure; keep at least one supervised live or simulated incident before signing off. Sign-off criteria should be service-specific rather than one generic "on-call ready" badge, since readiness on a well-instrumented service doesn't automatically transfer to a fragile legacy one with thin runbooks.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.