On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
Why do runbooks tend to go stale in a large engineering org? What are the common root causes, and what would you actually do about each one?
Sample Answer
Direct answer
Runbooks go stale because nothing automatically ties them to the systems they describe: ownership is unclear, updates aren't triggered by the changes that invalidate them, and nobody is rewarded for maintaining them, so they drift silently until an incident exposes it. The fix for each cause is the same shape: build the update into a workflow that already has to happen, like a deploy, a PR review, or a drill, rather than relying on someone remembering.
Structured elaboration
| Root cause | Why it happens | What to actually do |
|---|---|---|
| No clear owner | Docs feel like everyone's job, so they end up being no one's | Assign a named owner (team and person) per runbook, visible on the doc itself |
| No trigger tied to system changes | Infra or config changes ship without a linked doc update | Require a runbook-touch check in review for infra changes that affect the documented procedure |
| Fragmented across tools | The same procedure exists in a wiki, a chat pin, and a repo, and they diverge | One canonical source, docs-as-code in git; other tools link to it instead of duplicating it |
| Hard to edit | Binary or WYSIWYG pages discourage small fixes | Markdown in git with a low-friction pull-request flow |
| Never verified | Nobody runs the steps until a real incident forces it | Scheduled tabletop or game-day drills that surface breakage before it matters |
| Incentives favor code over docs | Engineers are measured on features shipped, not documentation kept accurate | Include doc currency in the definition of done or the on-call handoff checklist |
Worked example
A payments team migrates from a single database instance to a managed cluster with a different failover tool. The failover runbook still references the old promote command. Nobody touches the runbook because the migration's review process had no requirement to touch documentation tied to it, which is exactly the "no trigger tied to system changes" row above. Months later, an on-call engineer hits a real primary failure, runs the stale command, gets an error, and has to rediscover the correct procedure live instead of following a runbook that already had it. The root cause traces cleanly to the missing trigger, not to the engineer who wrote the original doc.
Trade-offs and pitfalls
- Quarterly "please review this doc" reminders without a named owner tend to become checkbox theater: marked reviewed without anyone actually re-verifying the steps.
- Gating merges on documentation updates adds friction to every infra change; scope the gate to changes that touch a documented procedure specifically, or teams will route around it entirely.
There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?
Sample Answer
Direct answer
Frame the sensitivity-versus-fatigue tension as an explicit cost trade-off rather than a vibes call: assign a rough relative cost to a missed incident versus a false page, pick the threshold that minimizes expected cost given the current false-positive and false-negative rates, and never roll a threshold change straight to paging. Validate it in shadow mode against real traffic first, then canary it on a subset before a full rollout, with an automatic rollback trigger if things get worse.
Structured elaboration
Cost framing. At a candidate threshold τ, define expected cost as:
cost(τ)=CFN⋅P(miss∣τ)+CFP⋅P(false page∣τ)Lowering τ (more sensitive) drives the probability of a miss toward zero but raises the false-page rate, and raising τ does the opposite. The right threshold is wherever this sum is smallest, not wherever either rate alone looks best in isolation.
Safe testing method, three stages with explicit gates:
- Shadow mode: the candidate threshold runs log-only, never pages, and every alert it would have fired gets compared against the real incident record after the fact. No production risk, but also no real-time responder feedback.
- Canary: the candidate threshold pages for real, but only for a subset of services or regions. Watch for missed-detection signals and responder load on that subset before touching anything else.
- Gradual rollout: feature-flagged expansion to the rest of the fleet, with an automatic rollback trigger if the false-positive rate or acknowledgment latency regresses past a predefined bound.
Worked example
Say a one-week shadow test compares the current threshold (A) against a stricter candidate (B) on the same underlying traffic, and every alert is later labeled against the real incident record:
| Threshold | Total alerts | True incidents caught (TP) | False pages (FP) | Missed incidents (FN) |
|---|---|---|---|---|
| A (current) | 200 | 18 | 182 | 2 |
| B (candidate) | 40 | 16 | 24 | 4 |
Assign a rough relative cost: a missed incident costs 500 responder-hour-equivalents (CFN=500), a false page costs 1 (CFP=1). Expected cost per threshold:
costA=CFN⋅FNA+CFP⋅FPA=500×2+1×182=1182 costB=CFN⋅FNB+CFP⋅FPB=500×4+1×24=2024Even though B has far higher precision (16 of 40 alerts were real, versus 18 of 200 for A), A has the lower expected cost: B's two extra missed incidents cost more than the 158 extra false pages A generates. If CFN were much closer to CFP, for example a low-stakes internal tool where a miss is only mildly annoying rather than expensive, B would win instead. The point of doing the arithmetic rather than eyeballing the false-positive rate is that the right threshold depends on getting the relative cost of a miss right for that specific service, not on chasing a universally "less noisy" target.
Trade-offs and pitfalls
- Pitfall: picking CFN and CFP once and never revisiting them. Relative costs shift as customer scale, contractual SLAs, and the service's blast radius change, so the cost model needs the same periodic review as the severity scale it feeds into.
- Pitfall: shadow-testing only against the same traffic period used to design the threshold in the first place, which overfits the result; test against a held-out period the threshold wasn't tuned on.
- Trade-off: assigning explicit costs forces an uncomfortable conversation with stakeholders about what a missed incident is actually worth, but that discomfort produces a threshold the team actually agreed to, rather than one that came from a single engineer's gut feel nobody was consulted on.
How would you keep an organization's runbooks accurate and useful over time instead of letting them rot? Talk through ownership and how you'd catch a runbook that's gone stale before someone relies on it during an incident.
Sample Answer
Runbooks rot because nothing forces them to change when the system they describe changes. Keeping them accurate means giving every runbook a named owner, tying updates to the events that actually invalidate a runbook (a relevant code or infra change, or an incident where it was used), and having a lightweight, recurring check that catches staleness before an incident does, rather than relying on someone remembering to update it.
Ownership
Every runbook has one named primary owner, the team or person who owns the service it covers, recorded in the runbook itself, not a separate spreadsheet that gets forgotten. Ownership isn't honorary: the owner is the one who signs off that the runbook is still accurate at each review point, and the one paged if a stale runbook causes a bad outcome during an incident.
What triggers a review, not just a calendar date
Calendar-only review cadences ("review every quarter") catch some staleness but miss the more common case: a runbook goes stale the day the system changes underneath it, not on a schedule. Two triggers matter more than the calendar:
- Change-linked: any deploy that touches the commands, infra, or thresholds a runbook references should require a runbook update as part of that same change, enforced by a checklist item on the pull request, not a follow-up ticket that competes with the next sprint.
- Incident-linked: every time a runbook is actually used during an incident, the post-incident review includes a specific question, did the runbook match reality, and any gap becomes a tracked follow-up before the incident is closed.
Lifecycle stages
| Stage | Meaning | Who moves it | Trigger |
|---|---|---|---|
| Active | Verified accurate, safe to follow blind | Owner | Passed its last review or was just used successfully in an incident |
| Needs review | A linked change or incident flagged it as possibly stale | Owner, auto-flagged | Change-linked or incident-linked trigger fires |
| Deprecated | Still readable but no longer the source of truth | Owner | Replacement runbook exists, or the failure mode it covers no longer applies |
| Archived | Removed from the on-call surface entirely | Owner | Deprecated for a defined grace period with no further reliance |
Catching a stale runbook before someone relies on it
A stale runbook is genuinely dangerous exactly because it looks trustworthy right up until the moment it's wrong. The cheapest catch is a light automated check, does every command in the runbook reference a tool, dashboard, or endpoint that still exists, run periodically and flagging anything broken for owner review. The more valuable catch is a periodic game-day: pick a runbook, have someone unfamiliar with the system try to follow it against a staging environment, and see where it breaks. Anything that trips someone up in a drill would have tripped up the on-call engineer at 3am.
Worked example: what enforcement actually looks like
A pull request changes the retry/backoff config for a service. The PR template includes a checkbox, "does this change any runbook referenced by this service's on-call docs?" Because the change alters a value a runbook's remediation step depends on, the author checks yes and links the runbook update in the same PR. Two months later, that runbook gets used during an incident; the post-incident review confirms the values matched, so the runbook stays Active with its last-reviewed date updated, no separate ticket needed because the change-linked trigger already did the work.
Trade-offs and pitfalls
Tying every infra change to a mandatory runbook update adds friction to routine PRs, so the check needs to be scoped narrowly (does this specific change affect a specific runbook) rather than a blanket "update all docs" gate that people learn to click through without reading. The most common failure mode is having a lifecycle model on paper but no one actually enforcing the Needs review to Active transition, so runbooks accumulate silently in Needs review and the label stops meaning anything; the fix is making that queue visible (a dashboard, not a buried label) and reviewing it in the same recurring meeting as on-call handoffs.
A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?
Sample Answer
Direct answer
Since you can't fix the vendor directly, you run three things in parallel: mitigate the blast radius with tools you do control (circuit breakers, cached or degraded responses, feature flags), communicate honestly about something outside your control, and push on the vendor relationship itself through support escalation and, if needed, contractual SLA terms. The trade-offs are mostly about how aggressively to degrade functionality versus how much broken or stale behavior your customers will tolerate in the meantime.
Structured elaboration
| Mitigation | What it buys you | What it costs |
|---|---|---|
| Circuit breaker / fail fast | Stops the vendor's failure from cascading into your own services | Feature becomes fully unavailable, more visible outage |
| Serve cached or stale data | Feature stays visibly "up" for the user | Risk of showing wrong or outdated information |
| Queue and retry with backoff | No data loss, eventual consistency once the vendor recovers | User sees delay; adds retry/backoff complexity |
| Feature-flag off (graceful degrade) | Predictable, pre-tested reduced experience | Only works if the flag and the reduced UX already exist before the outage |
Evidence to gather before contacting vendor support. Precise timestamps with timezone noted, representative request/response examples (method, URL, headers, correlation IDs), correlated logs from your own edge/load-balancer and application layers, and a clear scope-and-impact statement (which services, what percentage of traffic, which customers, what SLA is at risk). Vague "your API seems down" tickets sit in a generic queue; a ticket with reproducible evidence and a quantified impact gets triaged faster.
Escalating through vendor support tiers. Open the highest applicable severity case with the evidence attached and explicitly request an engineer and a bridge, not just an acknowledgment. If there's no meaningful response within your own internal SLA for that severity, escalate through the account manager or a phone-based escalation path, citing the specific business impact and contractual SLA terms rather than repeating the original ticket.
Communication cadence, using the same severity-driven pattern as an internal incident: acknowledge to affected customers quickly with what's known and any workaround, then update on a fixed cadence (for example every 30 minutes) until resolved, closing with a summary once the vendor confirms the fix.
Worked example
A payments provider starts returning errors for a subset of transactions. Mitigation: flip a feature flag that routes non-critical calls to a queued-retry path with a "processing" state shown to the user, instead of failing checkout outright; this preserves the customer experience for the subset of traffic where a short delay is tolerable, while transactions that genuinely require a synchronous response fail fast with a clear error rather than hanging. Communication: post an initial status update within roughly 15 minutes acknowledging degraded checkout with the workaround in place, then update every 30 minutes. Vendor escalation: open a high-severity vendor ticket with timestamped request/response examples and the affected transaction volume, request a bridge; if no vendor engineer engages within your internal escalation window, escalate via the account manager's phone line, citing the contractual SLA and quantified customer impact.
Trade-offs and pitfalls
- Pitfall: treating a vendor outage as "not our incident" and skipping the postmortem. Root cause may be external, but your own blast-radius design (whether a circuit breaker or cached fallback existed at all) is exactly what a postmortem should examine, since that's the part you actually control.
- Pitfall: promising customers a fix ETA you don't control. Communicate "investigating, using workaround X, next update in 30 minutes" rather than a timeline that depends on someone else's incident response.
- Trade-off: aggressive circuit-breaking protects your own systems fastest but produces the most visible outage; cached/degraded responses are gentler on the user experience but carry a correctness risk if the vendor's data changes underneath the cache. Which one is right depends on how stale or wrong data is allowed to be for that specific feature, which is a product decision, not just an engineering one.
How would you measure whether an on-call rotation is sustainable or quietly burning people out? What would you actually track?
Sample Answer
No single number proves burnout. Track three families of signals together: raw load (pages per person per week), response burden (after-hours percentage, time-to-resolve), and human signals (fatigue self-reports, PTO usage), and watch for the same people repeatedly crossing thresholds across categories, not just one bad week.
What to track
| Category | Metric | Sustainable guideline | What it flags |
|---|---|---|---|
| Load | Pages per primary on-call per week | Under ~10/week | Rotation or alert volume is too high |
| Load | Share of pages from one service | No single service over ~40% of team pages | One noisy service is dominating the rotation |
| Response burden | After-hours page percentage | Under ~25% | Sleep disruption, needs alert-hours review |
| Response burden | P90 time-to-resolve trend | Flat or improving | Chronic fatigue slowing responders, not just harder incidents |
| Human signal | Post-incident fatigue self-report (1-5) | Sustained score at or below 2 | Early warning before hard metrics move |
| Human signal | PTO usage on the rotation | Not declining quarter over quarter | People avoiding time off is a red flag, not a green one |
Worked example: reading a four-week rotation block
Suppose the primary on-call received a combined 88 pages across the last 4-week rotation block (one week per engineer).
Pages per on-call week=488=22 pages/week≈3.1/dayAgainst the roughly-under-10/week guideline, 22 pages/week is more than double, a sustainability flag on its own. If 39 of those 88 pages fired between 20:00 and 08:00:
After-hours share=8839×100≈44.3%well above the roughly-25% guideline, corroborating that this isn't just a high-volume rotation, it's specifically disrupting sleep.
Trade-offs and pitfalls
These metrics can be gamed by suppressing alerts; pair volume metrics with an independent audit (a sampled review of closed incidents) so under-alerting doesn't masquerade as improvement. Self-reported fatigue data is noisy and subject to survey fatigue itself; treat it as a leading indicator alongside hard metrics, not as the sole trigger for action. A single bad week (one major outage) will spike every metric at once; look for a sustained pattern across at least a full rotation cycle before concluding the rotation itself, rather than the incident, is the problem.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.