On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
What would it look like to treat runbooks as code: version-controlled, linted, and tested in CI before a change can merge? Walk through how you'd implement that, and how you'd keep thousands of runbooks discoverable enough that an on-call engineer can find the right one within a couple of minutes.
Sample Answer
Direct answer
Store runbooks as Markdown with structured YAML front-matter in one canonical git repository, gate every change behind CI that lints and schema-validates the metadata and dry-run-tests any automatable step, and require review before merge, the same discipline as application code. Discoverability at scale comes from that same metadata: a required service field feeds a searchable index, and alert annotations link straight to the matching runbook, so an on-call engineer reaches the right doc from the page itself instead of searching for it.
Structured elaboration
CI pipeline
flowchart LR
A[Engineer edits runbook.md] --> B[Open PR]
B --> C[CI: lint YAML and Markdown]
C --> D[CI: schema-validate metadata]
D --> E[CI: dry-run automatable steps in sandbox]
E --> F{Checks pass?}
F -->|No| A
F -->|Yes| G[Owner review plus on-call ack]
G --> H[Merge to main]
H --> I[Publish to searchable index]
I --> J[Alert annotations link runbook URL]
Required metadata
| Field | Purpose |
|---|---|
| id | Stable identifier, referenced from alerts and postmortems |
| service | Ties the runbook to a service registry entry |
| owner | Team and on-call rotation responsible for upkeep |
| severity | Drives which validation tier and review gate applies |
| last_reviewed / last_verified | Freshness signal, updated only after a real validation pass |
| automatable | Whether any step can run through the execution platform |
Making thousands of runbooks findable in minutes
Every runbook's service field is validated against a central service registry at CI time, which catches typos and stale references after a rename. Alert rules carry the runbook's canonical URL in their annotations, so paging a human already surfaces the right doc; searching is the fallback path, not the primary one. A generated, always-fresh index, built from the same metadata at publish time, is searchable by service, tag, and severity, so finding the right runbook is a filtered lookup rather than full-text search over prose.
Rationalizing a messy legacy corpus
When inheriting hundreds of duplicated or contradictory legacy runbooks, don't try to review them all before migrating. Import everything into the new repo behind CI, then triage: retire clear duplicates, keeping the most-recently-verified version and redirecting the rest, and mark anything ambiguous as needing review with an expiry date so it doesn't sit unreviewed indefinitely. Treat the migration as incremental and CI-gated, not a one-time cleanup that has to finish before the new system can go live.
Linking postmortems back to runbooks
Postmortem action items that touch a runbook should reference the runbook's id field directly, not just link to it in prose. That makes it possible to build a reverse index of which incidents cited a given runbook's gaps, so recurring weak spots surface automatically instead of living only in scattered postmortem documents.
Worked example
A runbook's front-matter:
---
id: RB-DA-001
title: "Spark job failing with executor OOM"
service: data-ingest-pipeline
owner: team-data-platform
severity: P1
tags: [spark, etl]
last_reviewed: 2026-05-01
automatable: true
---
A PR that removes the owner field or sets severity: P5 (not a valid value in the schema) fails CI's schema-validation step with a specific, actionable error, rather than merging silently and leaving the on-call escalation path undefined for that runbook.
Trade-offs and pitfalls
- Over-engineering the schema so authoring a runbook feels like filing a support ticket discourages contributions; keep required fields minimal, owner, service, severity, last_reviewed, and let the rest stay freeform Markdown.
- Strict CI gating with owner review and on-call acknowledgment slows updates during an actual incident; document a fast-path for post-incident corrections that still requires review, just not before the fix ships in the moment.
- Building the search index without tying it to alerts leaves on-call still searching manually under pressure. The annotation link on the alert is what actually saves time; the index is the fallback.
Before a new service goes live and starts taking on-call pages, what would you want to see in place? Walk through what a production-readiness review should check.
Sample Answer
A production-readiness review should verify four things before a service starts taking pages: it fails safely (degrades or rolls back instead of cascading), it's observable enough that on-call can diagnose without guessing, on-call actually knows how to respond to it, and someone specific owns it. Structure the review around those four, not a flat checklist, so gaps are obvious by category rather than buried in a long list.
What the review checks, by category
| Category | What to verify | Why it's a gate, not a nice-to-have |
|---|---|---|
| Failure containment | Circuit breakers or timeouts on every downstream call; load-testing evidence at expected peak plus a safety margin; tested rollback path | Without these, a dependency hiccup or a launch-day traffic spike becomes an on-call incident that a healthy service wouldn't have had |
| Observability | Dashboards for the service's key health signals; alerts tied to those signals with sane thresholds (not just "CPU high"); logs/traces sufficient to diagnose the top 3 failure modes without SSH-ing into a box | On-call can't respond to what they can't see; this is the difference between a 10-minute diagnosis and a 2-hour one |
| Runbook readiness | At least one runbook per alert that can actually fire, covering symptom, diagnosis steps, and remediation; runbook has been read (ideally walked through) by the people who'll be paged | An alert with no runbook just wakes someone up with no next step |
| Ownership and escalation | Named on-call rotation for the service, not "whoever's around"; a documented escalation path if the primary can't resolve it; the service is actually in the paging tool's routing, not just assumed to be | Ambiguous ownership is invisible until the first incident, when it costs the most |
Process for running the review
- The owning team self-certifies against the checklist first, providing evidence (load-test results, a link to the rollback runbook, a screenshot of the dashboard) rather than a checked box with no backing.
- A reviewer outside the owning team (SRE or a peer team) spot-checks the evidence, focusing on the failure-containment and observability rows, since those are the ones teams under launch pressure are most likely to overstate.
- Run one live-fire test before go-live: trigger the most likely failure mode in staging (or a controlled prod canary) and confirm the alert fires, the runbook's diagnosis steps actually find the cause, and the rollback works. A checklist that's never been exercised is a hypothesis, not a verified readiness state.
- Sign-off is explicit and time-bound, not a one-time gate that's forgotten: re-review triggers on major architecture changes, not just at initial launch.
flowchart TD
A[Owning team self-certifies checklist] --> B[Provide evidence: load tests, runbook links, dashboards]
B --> C[Outside reviewer spot-checks evidence]
C --> D{Gaps found?}
D -->|Yes| E[Team remediates gap]
E --> C
D -->|No| F[Live-fire test in staging or canary]
F --> G{Alert fires, runbook works, rollback succeeds?}
G -->|No| E
G -->|Yes| H[Sign-off: service takes pages]
Worked example
A new recommendations service is going live. Self-certification claims load testing was done "at expected traffic." The outside reviewer asks for the actual load-test report and finds it tested at the current expected peak (500 req/s) with no margin, while the service also sits behind a feature flag that product plans to ramp to three times that within a month. That's a real gap: the review isn't asking for perfection, but it should require either testing at the higher number now or an explicit, documented plan (with an owner and date) to re-test before the ramp, rather than letting "tested at expected traffic" silently mean "tested at today's traffic." The live-fire test then finds the circuit breaker on the downstream recommendation-model call has no timeout configured, so a slow model response would hang the request instead of failing fast; that's flagged as a blocking issue, not a follow-up ticket, because it directly causes cascading failure under exactly the load condition the service is meant to handle.
Trade-offs and pitfalls
The live-fire test step is the one teams most often skip under launch deadline pressure, and it's also the one that catches the gaps self-certification checklists miss (an alert that's configured but never actually fires, a runbook step that references a dashboard that doesn't exist); treat it as non-negotiable for anything customer-facing, and reserve a lighter self-certification-only path for low-risk internal services. A common wrong turn is treating the checklist as complete once every box is checked, without weighting which gaps are load-bearing; a missing rollback plan for a payments-adjacent service is not the same severity as a missing dashboard for an internal admin tool, and the review should say so explicitly rather than gating everything equally.
How would you actually validate that your runbooks work before you need them in a real incident? Describe a program for testing them under realistic conditions.
Sample Answer
Direct answer
Validating runbooks before a real incident needs three ingredients: a way to safely exercise them, using sandboxes for non-destructive steps and canary execution for destructive ones; measurable acceptance criteria for what "verified" actually means; and a cadence that scales from cheap, frequent tabletop walkthroughs up to expensive, rare full game days. Treat validation as a graded ladder of increasing realism and cost, not a single all-or-nothing chaos exercise.
Structured elaboration
Ladder of validation
| Tier | What happens | Frequency | Blast radius |
|---|---|---|---|
| Tabletop read-through | Team reads the runbook aloud, checks it still matches current architecture | Monthly per critical runbook | None, discussion only |
| Sandboxed dry-run | Non-destructive or dry-run steps run against a synthetic or staging copy | Per runbook change, CI-gated | Isolated sandbox |
| Canary execution | The real, potentially destructive step runs against a single instance or shard in production | Quarterly for the highest-severity runbooks | One instance or shard |
| Full game day | The real trigger condition is simulated and the whole runbook runs end to end with the actual on-call rotation | Quarterly cross-team, and after major architecture changes | Scoped production traffic, with a kill switch |
Acceptance criteria for "verified"
- Every command in the runbook executed successfully against the environment matching its tier.
- Recovery met the documented RTO (recovery time objective: the maximum acceptable downtime) for that exercise.
- The person executing it was not the runbook's original author, which catches "only the author can actually run this" runbooks.
- No manual step was needed beyond what's written, which catches missing steps.
- The runbook's last-verified metadata is updated only after all of the above pass, tied to the specific commit that was tested.
Sandboxing and canary mechanics for destructive steps
- A dry-run flag on any script validates and logs without mutating anything; most cloud SDKs and infrastructure-as-code tools support this natively.
- Ephemeral, synthetic-data environments handle full destructive rehearsals safely, isolated by namespace or project and feature-flagged away from real customer traffic.
- Steps that can only be meaningfully tested in production, like a real failover, get a canary first: a single shard or instance, with an automated rollback path and a pre-agreed abort condition.
Error-budget gate before running in production
Before a production game day, confirm enough error budget remains to absorb the intended, and any accidental, impact. For a monthly SLO of 99.9 percent over a 30-day window, the allowed downtime is:
error budget=(1−SLO)×window minutes
(1−0.999)×43,200=0.001×43,200=43.2 minutes
If the remaining budget is close to that 43.2-minute figure, postpone the exercise rather than spend the safety margin on a drill.
Worked example
A cache-cluster failover runbook is canary-tested against a single shard. The failover command executes without manual intervention, recovery meets the documented RTO target, and the engineer running the drill is not the runbook's original author. All four acceptance criteria above pass, so the runbook's last-verified metadata is updated to the exact commit hash that was tested, and the result feeds into the quarterly decision of whether this runbook is due for a full game day next.
Trade-offs and pitfalls
- Relying only on tabletop reads because full game days are expensive means staleness in the actual commands never gets caught until a real incident does it for you.
- Letting the runbook's author be the only person who can successfully execute it means you've tested the author's tribal knowledge, not the documentation; a different operator running the drill is what actually validates the doc.
- Full production game days build the highest confidence but carry real risk and spend real error budget; the ladder exists so most validation stays cheap, and only the highest-severity runbooks earn a full game day.
Two unrelated incidents hit different services at the same time. How do you decide how to allocate people across them, and when do you escalate to a higher-level incident commander?
Sample Answer
Allocate people by comparing the two incidents' business impact and required expertise, not by splitting the team evenly, and default to running them as separate incidents with separate commanders unless you find a shared root cause; escalate to a higher-level incident commander as soon as the resource conflict itself (not just the technical severity) becomes the bottleneck.
Allocating people across concurrent incidents
- Score each incident independently first: user-facing impact, revenue impact, data-integrity risk, and blast radius (one team's problem versus platform-wide). Two incidents rarely score identically, and the higher-scored one gets first claim on the strongest responders.
- Check for a shared root cause before splitting resources. If both services depend on the same failing component (a shared database, a shared auth service), that's actually one incident with two symptoms, and it should be run as a single incident with one commander, not two competing efforts pulling on the same underlying fix.
- Staff each independent incident with a minimum viable team: one incident commander, one primary responder, one communications owner. Resist over-staffing the incident that's louder or more visible if the other one is actually higher severity but quieter.
- Protect against double-booking the same expert. If one person is the only one who understands a shared piece of infrastructure, they can advise both incidents briefly but should not be the sole owner of fixing both; pull in a secondary responder even if slower.
When to escalate to a higher-level incident commander
Escalate when any of these is true, not just when severity is high:
- Resource contention itself is blocking progress: both incidents need the same scarce specialist or the same change-freeze exception, and someone above both incident commanders needs to arbitrate.
- Combined blast radius crosses an organizational boundary: the two incidents together affect enough of the business (multiple product lines, a shared customer segment) that unified external communication is needed, even if each incident alone wouldn't trigger that.
- One incident commander is starting to context-switch between both incidents. A single IC trying to run two incidents at once is a bigger risk than the incidents themselves; that's a signal to bring in a second commander or an overall coordinator, not to push through.
- Duration crosses a threshold where sustained dual-incident load starts to fatigue the responding team; a higher-level commander can pull in fresh responders or make the call to deprioritize the lower-severity incident explicitly (and communicate that decision) rather than let it silently starve.
Worked example
Two incidents fire eleven minutes apart: the checkout service is returning 500s for roughly a third of requests (revenue-impacting, high severity), and the internal analytics dashboard is showing stale data (no customer impact, low severity). The correct allocation: full incident-commander-plus-primary-plus-comms team goes to checkout immediately; analytics gets a single responder to investigate on a non-paging basis, because pulling more people onto analytics wouldn't shorten its resolution meaningfully and would strip capacity from checkout. If, twenty minutes in, the checkout investigation discovers the 500s trace back to the same message queue that feeds the analytics pipeline, the two incidents are merged under checkout's commander, because they share a root cause and running them separately would mean two people independently investigating the same queue.
Escalation in this example would trigger only if a third, unrelated incident arrived while checkout was still active and unresolved: three concurrent incidents makes single-IC-per-incident coordination itself the bottleneck, which is exactly the resource-contention trigger above.
Trade-offs and pitfalls
A common mistake is allocating headcount proportional to how loud or visible each incident is (how many people are asking about it in Slack) rather than its actual business impact; loud and low-impact will out-compete quiet and high-impact if you let it. Another is treating "escalate to a higher IC" as an admission of failure, so teams delay it past the point where a fresh coordinator would have resolved the resource conflict faster. The trade-off with merging incidents on a suspected shared root cause is real: merge too eagerly and you lose the separate investigation threads that might have found the divergence faster; the mitigation is to merge the coordination and communication, but keep separate technical workstreams until the shared cause is actually confirmed.
How would you decide whether a runbook is actually ready for on-call use, not just written? What would you check before trusting it during a real incident?
Sample Answer
A runbook being written isn't the same as it being trustworthy under pressure: writing tests whether the author understood the system, while readiness tests whether someone else, half-awake at 3am, can follow it and get the right outcome. Check readiness by having someone who didn't write it actually execute it against a real (or realistic) system, not by reading it for completeness.
What to check before trusting a runbook
- Has anyone other than the author run it? A runbook the author has never handed to someone else is unverified by definition; the author's own mental model fills gaps a stranger will trip on (an assumed tool is installed, an assumed permission is already granted, a step that says "check the dashboard" without saying which one).
- Are the steps executable as written, not just described? "Restart the service" is a description; "run
systemctl restart payments-apion each of the 3 hosts listed in the service registry" is executable. If a step requires judgment the runbook doesn't supply (how do you know which hosts?), that's a gap, not an acceptable level of abstraction. - Does it state what success looks like? A remediation step without a stated verification step (what metric or log line confirms this worked) leaves the responder guessing whether to move to the next step or escalate.
- Is it safe to run when the diagnosis is wrong? Incident responders under pressure sometimes run the wrong runbook, or run the right one when the actual cause differs from what it assumes. Check whether each destructive step is reversible, and whether the runbook states a precondition to verify before acting ("only run this if X").
- Is it current? Check for an owner and a last-verified date; a runbook referencing a deprecated tool, an old cluster name, or a rotation that no longer exists is worse than no runbook, because it costs time before the responder realizes it's wrong.
How to actually verify these, not just check for their presence
- Tabletop walkthrough: someone unfamiliar with the runbook reads it aloud, step by step, without help from the author, narrating what they'd actually type or click. Gaps surface immediately as "wait, what do I do here?" moments.
- Staging or canary drill: run the actual remediation against a staging environment or a single canary instance, and time it. This catches steps that look right on paper but fail against the real system (a command with an outdated flag, a permission the on-call role doesn't actually have).
- Cold-open test: hand it to someone with zero context on this specific service (not zero context on the systems generally) and see if they can act on it in under a target time, without pinging the original author. If they can't, the runbook is only usable by the person who wrote it, which defeats the point.
Worked example
A runbook for "database replica lag alert" says: "Check replica lag, if high, failover to standby." A cold-open drill immediately exposes three gaps: no link to where replica lag is displayed, no threshold for what counts as "high" (the alert already fired, so this should already be answered, but the runbook re-asks the question), and "failover to standby" doesn't say which standby if there are multiple, or what to verify afterward to confirm the failover succeeded rather than made things worse. Fixing it: link the specific dashboard panel, state the alert's own threshold so the runbook doesn't require re-deciding it, name the failover command with the specific standby-selection logic, and add a verification step ("confirm write latency on the new primary is under 50ms and replica lag on remaining replicas is decreasing"). Re-running the cold-open drill after the fix, the same tester completes it without asking a clarifying question, which is the actual pass condition.
Trade-offs and pitfalls
Running live drills has a real cost in engineering time and, for staging drills, some risk if the environment isn't well isolated from production; the return is worth it for any runbook covering a high-severity or destructive action, and can be scaled down to tabletop-only for low-risk, easily reversible ones. A common wrong turn is treating runbook review as a documentation-quality pass (is it well written, does it have headers) rather than an execution test; a beautifully formatted runbook that's never been run by anyone but its author is still unverified.
Unlock Full Question Bank
Get access to all 41 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.