On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
Write a runbook a junior on-call engineer with only basic familiarity could follow to resolve a common recurring failure, like a database running out of connections. How much detail do you include, and how do you make sure it's usable by someone half-asleep at 3am?
Sample Answer
For a junior, half-asleep responder, the runbook needs to be a strict linear sequence of copy-pasteable commands with expected output shown at each step, not a description of the problem. A connection pool is just the fixed set of open database connections an application reuses instead of opening a new one per request; "exhausted" means every connection in that set is busy or leaked, so new requests queue or fail waiting for one to free up. Below is the runbook, followed by how the same template extends to a few other common recurring failures.
Runbook: Database connection pool exhausted
Applies when: alert fires for "DB connections at capacity" or the app logs show timeouts waiting for a pooled connection.
1. Confirm the symptom
-- Postgres: how many connections are open right now, and what state are they in
SELECT state, count(*) FROM pg_stat_activity GROUP BY state;
If active + idle in transaction is at or near the configured pool max, this is confirmed.
2. Check for the two common causes
-- Long-running or stuck queries holding a connection
SELECT pid, state, now() - query_start AS duration, query
FROM pg_stat_activity
WHERE state != 'idle'
ORDER BY duration DESC
LIMIT 10;
A handful of queries running for minutes, not seconds, points at a stuck query holding connections hostage. Many short connections in idle in transaction for the same app host points at a connection leak in that service (it opened a connection and never closed it).
3. Immediate mitigation
- If one or two queries are clearly stuck (duration in minutes, not part of a normal batch job): terminate them.
SELECT pg_terminate_backend(pid) FROM pg_stat_activity WHERE pid = <PID>;
- If one app instance is leaking connections: restart that instance only, not the whole fleet, to limit blast radius.
# Example: restart a single app instance behind a load balancer
kubectl rollout restart deployment/<app-name> --namespace=<ns>
4. Verify recovery
Re-run the query from step 1. Connection count should drop back under 80% of pool max within a couple of minutes, and application error logs should stop showing connection-timeout errors.
5. Escalate if
- Step 1 confirms the symptom but step 2 shows no stuck queries and no obvious leaking host, escalate to the on-call DBA.
- Terminating queries and restarting the instance doesn't bring connections back down within 10 minutes, escalate to the on-call DBA and the service owner together.
6. Log it
Note in the incident channel: which queries or instance were involved, what you ran, and the connection count before/after. This becomes the input for the next runbook review.
Applying the same template to other recurring failures
The shape (confirm symptom, check the one or two common causes, apply a scoped mitigation, verify, escalate on a threshold) is the same regardless of what's failing:
- A worker process stuck or crash-looping: confirm via the process manager's status command, check recent logs for the crash reason, restart just that worker (not the whole host), verify it stays up for a few minutes, escalate if it crash-loops again within 15 minutes.
- A Kubernetes service unresponsive: confirm via
kubectl get podsshowingCrashLoopBackOffor 0 ready replicas, checkkubectl describe podfor the failure reason, roll the deployment, verify readiness probes pass, escalate if the rollout itself fails. - High database CPU: confirm via the DB's CPU metric, check
pg_stat_activityfor a small number of expensive queries dominating, kill the worst offender if it's not part of an expected batch job, verify CPU drops, escalate if CPU stays high with no single query responsible (that usually means aggregate load, not one query, and needs a scaling decision above a junior responder's authority).
Keeping it usable day to day, not just during an incident
Part of what keeps a junior responder capable of running this cold at 3am is having already looked at the same dashboard during calmer moments. A daily on-call checklist (glance at connection count, query latency, and error rate once per shift, even with no alert firing) means the responder already knows what "normal" looks like on this system before the night it matters, instead of learning it for the first time under a page.
Trade-offs and pitfalls
The runbook deliberately doesn't cover permanent fixes (increasing pool size, fixing the leak in application code, adding query timeouts), those are follow-up work items, not something a junior responder should do live against production under a page; bundling them in would tempt someone under pressure into a riskier change than the incident calls for. The other pitfall is giving a vague duration threshold like "long-running query"; "minutes, not seconds" is deliberately concrete so a junior responder isn't left guessing whether a 40-second query counts.
How would you decide whether a runbook is actually ready for on-call use, not just written? What would you check before trusting it during a real incident?
Sample Answer
A runbook being written isn't the same as it being trustworthy under pressure: writing tests whether the author understood the system, while readiness tests whether someone else, half-awake at 3am, can follow it and get the right outcome. Check readiness by having someone who didn't write it actually execute it against a real (or realistic) system, not by reading it for completeness.
What to check before trusting a runbook
- Has anyone other than the author run it? A runbook the author has never handed to someone else is unverified by definition; the author's own mental model fills gaps a stranger will trip on (an assumed tool is installed, an assumed permission is already granted, a step that says "check the dashboard" without saying which one).
- Are the steps executable as written, not just described? "Restart the service" is a description; "run
systemctl restart payments-apion each of the 3 hosts listed in the service registry" is executable. If a step requires judgment the runbook doesn't supply (how do you know which hosts?), that's a gap, not an acceptable level of abstraction. - Does it state what success looks like? A remediation step without a stated verification step (what metric or log line confirms this worked) leaves the responder guessing whether to move to the next step or escalate.
- Is it safe to run when the diagnosis is wrong? Incident responders under pressure sometimes run the wrong runbook, or run the right one when the actual cause differs from what it assumes. Check whether each destructive step is reversible, and whether the runbook states a precondition to verify before acting ("only run this if X").
- Is it current? Check for an owner and a last-verified date; a runbook referencing a deprecated tool, an old cluster name, or a rotation that no longer exists is worse than no runbook, because it costs time before the responder realizes it's wrong.
How to actually verify these, not just check for their presence
- Tabletop walkthrough: someone unfamiliar with the runbook reads it aloud, step by step, without help from the author, narrating what they'd actually type or click. Gaps surface immediately as "wait, what do I do here?" moments.
- Staging or canary drill: run the actual remediation against a staging environment or a single canary instance, and time it. This catches steps that look right on paper but fail against the real system (a command with an outdated flag, a permission the on-call role doesn't actually have).
- Cold-open test: hand it to someone with zero context on this specific service (not zero context on the systems generally) and see if they can act on it in under a target time, without pinging the original author. If they can't, the runbook is only usable by the person who wrote it, which defeats the point.
Worked example
A runbook for "database replica lag alert" says: "Check replica lag, if high, failover to standby." A cold-open drill immediately exposes three gaps: no link to where replica lag is displayed, no threshold for what counts as "high" (the alert already fired, so this should already be answered, but the runbook re-asks the question), and "failover to standby" doesn't say which standby if there are multiple, or what to verify afterward to confirm the failover succeeded rather than made things worse. Fixing it: link the specific dashboard panel, state the alert's own threshold so the runbook doesn't require re-deciding it, name the failover command with the specific standby-selection logic, and add a verification step ("confirm write latency on the new primary is under 50ms and replica lag on remaining replicas is decreasing"). Re-running the cold-open drill after the fix, the same tester completes it without asking a clarifying question, which is the actual pass condition.
Trade-offs and pitfalls
Running live drills has a real cost in engineering time and, for staging drills, some risk if the environment isn't well isolated from production; the return is worth it for any runbook covering a high-severity or destructive action, and can be scaled down to tabletop-only for low-risk, easily reversible ones. A common wrong turn is treating runbook review as a documentation-quality pass (is it well written, does it have headers) rather than an execution test; a beautifully formatted runbook that's never been run by anyone but its author is still unverified.
During a postmortem, the incident commander singles out one engineer as the cause of the outage. How do you respond in the moment to preserve a blameless culture, without letting accountability for the fix slide?
Sample Answer
In the moment, redirect the conversation from the person to the timeline: acknowledge what was said without amplifying it, then immediately steer the group back to reconstructing what happened and why the system allowed it, while making clear that accountability for the fix is not going away.
In-the-moment response
- Interrupt with a redirect, not a confrontation. Something like: "Let's hold on names for a second and walk the timeline: what did the system show at each step?" This isn't ignoring what was said; it's refusing to let the postmortem's structure reward the blame framing by continuing down that thread.
- Reframe the specific claim into a system question. If the IC (Incident Commander, the person directing the response) says "this happened because Priya deployed without checking the dashboard," the redirect is: "So the deploy process didn't require a dashboard check before going out. Is that a gap in the checklist, or did the checklist exist and get skipped? Either answer tells us what to fix." This keeps the factual content (a deploy went out without a check) while stripping the blame framing.
- Do not let it pass silently either. Staying quiet when a peer is singled out in front of the team reads as agreement, and it's the fastest way to make the next engineer afraid to be transparent in their own postmortem. A short, calm correction in the room is better than a private word afterward, because the damage (and the culture signal) happened publicly.
- Follow up with the IC privately, separate from the room. The public redirect handles the moment; a private conversation afterward addresses the pattern, especially if this IC does it repeatedly.
Keeping accountability intact
Blameless does not mean no one owns the fix. The distinction to hold onto:
- Blame assigns fault for what already happened, to a person, and looks backward.
- Accountability assigns ownership for what happens next, to a role or system, and looks forward.
So the postmortem should still end with a named owner for each remediation item (a person, because someone has to actually do the work) and a deadline, but the framing is "you're the best person to close this gap because you understand the deploy path," not "this is your fault so you have to fix it." The action items get assigned based on who has the context and capability, independent of who gets blamed.
Worked example
During a payments-outage postmortem, the IC says: "Marcus rolled back the config and that's what caused the second outage." The redirect: "Let's look at what the rollback runbook told him to check before rolling back. Did it call out this specific config's downstream dependency?" The team pulls up the runbook and finds it didn't mention that this particular config was read by two other services; the rollback step existed, but the pre-check for downstream impact didn't. The postmortem action items become: (1) add a downstream-dependency check to the rollback runbook for this config, owned by the platform team, due in two weeks, and (2) audit other high-fanout configs for the same missing check, owned by Marcus, since he now has the clearest picture of what that gap looks like, due in one month. Marcus ends up with an action item, but it's framed as "you're positioned to close this" rather than "you caused this," and the runbook gap, not Marcus's judgment, is recorded as the finding.
Trade-offs and pitfalls
The main pitfall is overcorrecting into vagueness, where "blameless" gets used to avoid naming any specific decision point, and the postmortem ends up too soft to actually change anything; the fix is to be precise about the decision and the missing guardrail while staying impersonal about who made the decision. A related pitfall specific to this scenario: correcting an incident commander in front of the team carries real interpersonal risk if done poorly, so the redirect has to stay factual and calm rather than accusatory itself. This is a leadership-culture issue that shows up at scale too: it typically takes deliberate, sustained work, roughly a couple of quarters of consistent leadership behavior, published blameless postmortems, and visible non-punitive handling of pages, to shift a team's on-call culture away from a punitive default, and it has to be reinforced the same way every time, including in the exact moment someone in authority breaks the pattern.
What's the difference between a runbook and a playbook, and when would you reach for one instead of the other?
Sample Answer
Direct answer
A runbook is a fixed set of step-by-step instructions for a known failure mode: it tells you exactly what to type. A playbook is a decision framework for handling an incident more broadly: it tells you how to figure out what to do, who needs to be involved, and which runbook to reach for. Reach for a runbook when you already know the cause and the fix is mechanical; reach for a playbook when you're still diagnosing, coordinating multiple people, or the right response depends on judgment.
Structured elaboration
| Runbook | Playbook | |
|---|---|---|
| Scope | One specific, known failure mode or task | A class of incidents, or the overall response process |
| Format | Linear, prescriptive steps | Decision tree or branching guidance |
| Answers | "What do I type" | "What do I decide, and who do I involve" |
| Typical contents | Preconditions, exact commands, verification steps, rollback | Severity thresholds, roles (IC, comms lead), escalation matrix, links to runbooks |
| Usually owned by | The team that owns the specific service | Incident response leadership or SRE |
| Reviewed when | The underlying system changes | The org's escalation structure or tooling changes |
Minimum fields for each:
- Runbook: title, service, owner, trigger/precondition, required permissions and tools, exact step-by-step commands, verification steps, rollback steps, expected impact, last-reviewed date.
- Playbook: title, scope and severity thresholds, incident commander and stakeholder roles, the decision tree itself, links to the relevant runbooks, communication templates, escalation matrix, last-reviewed date.
Worked example
A database replica's lag exceeds a threshold: this is a runbook. It lists the exact commands to promote a replica, the steps to reconfigure the application's read preference, verification queries to confirm the fix, and the rollback commands if the promotion causes a new problem. Now compare that to a major outage affecting payments: this is a playbook. It guides the incident commander through detecting the actual scope, declaring severity, deciding between routing traffic to a fallback payment path versus draining traffic entirely, coordinating the app, infra, and comms teams, and linking out to the specific runbooks (including the replica-promotion one, if that turns out to be the fix) for whichever technical action the decision tree leads to.
Trade-offs and pitfalls
- Pitfall: writing a "runbook" for something that actually needs judgment, for example "when in doubt, restart the service." That hides a decision a playbook should make explicit, and someone follows it verbatim during exactly the incident it doesn't fit.
- Pitfall: letting a playbook go stale is worse than letting one runbook go stale, because the playbook is what everyone reaches for first during ambiguity; a stale escalation matrix (wrong names or numbers) breaks the whole response, not just one specific fix.
- Trade-off: automating a runbook into a one-click execution is great for high-confidence, low-blast-radius fixes (restarting a stateless service) and dangerous for high-blast-radius ones (promoting a database replica). The more damage a wrong click can do, the more the runbook should require an explicit human confirmation step before executing, not less.
- Both belong in a versioned, reviewed repository rather than an unowned wiki page, with periodic review, and runbooks specifically benefit from occasional dry-run or game-day testing to confirm the steps still work against the current system rather than an outdated one.
What would it look like to treat runbooks as code: version-controlled, linted, and tested in CI before a change can merge? Walk through how you'd implement that, and how you'd keep thousands of runbooks discoverable enough that an on-call engineer can find the right one within a couple of minutes.
Sample Answer
Direct answer
Store runbooks as Markdown with structured YAML front-matter in one canonical git repository, gate every change behind CI that lints and schema-validates the metadata and dry-run-tests any automatable step, and require review before merge, the same discipline as application code. Discoverability at scale comes from that same metadata: a required service field feeds a searchable index, and alert annotations link straight to the matching runbook, so an on-call engineer reaches the right doc from the page itself instead of searching for it.
Structured elaboration
CI pipeline
flowchart LR
A[Engineer edits runbook.md] --> B[Open PR]
B --> C[CI: lint YAML and Markdown]
C --> D[CI: schema-validate metadata]
D --> E[CI: dry-run automatable steps in sandbox]
E --> F{Checks pass?}
F -->|No| A
F -->|Yes| G[Owner review plus on-call ack]
G --> H[Merge to main]
H --> I[Publish to searchable index]
I --> J[Alert annotations link runbook URL]
Required metadata
| Field | Purpose |
|---|---|
| id | Stable identifier, referenced from alerts and postmortems |
| service | Ties the runbook to a service registry entry |
| owner | Team and on-call rotation responsible for upkeep |
| severity | Drives which validation tier and review gate applies |
| last_reviewed / last_verified | Freshness signal, updated only after a real validation pass |
| automatable | Whether any step can run through the execution platform |
Making thousands of runbooks findable in minutes
Every runbook's service field is validated against a central service registry at CI time, which catches typos and stale references after a rename. Alert rules carry the runbook's canonical URL in their annotations, so paging a human already surfaces the right doc; searching is the fallback path, not the primary one. A generated, always-fresh index, built from the same metadata at publish time, is searchable by service, tag, and severity, so finding the right runbook is a filtered lookup rather than full-text search over prose.
Rationalizing a messy legacy corpus
When inheriting hundreds of duplicated or contradictory legacy runbooks, don't try to review them all before migrating. Import everything into the new repo behind CI, then triage: retire clear duplicates, keeping the most-recently-verified version and redirecting the rest, and mark anything ambiguous as needing review with an expiry date so it doesn't sit unreviewed indefinitely. Treat the migration as incremental and CI-gated, not a one-time cleanup that has to finish before the new system can go live.
Linking postmortems back to runbooks
Postmortem action items that touch a runbook should reference the runbook's id field directly, not just link to it in prose. That makes it possible to build a reverse index of which incidents cited a given runbook's gaps, so recurring weak spots surface automatically instead of living only in scattered postmortem documents.
Worked example
A runbook's front-matter:
---
id: RB-DA-001
title: "Spark job failing with executor OOM"
service: data-ingest-pipeline
owner: team-data-platform
severity: P1
tags: [spark, etl]
last_reviewed: 2026-05-01
automatable: true
---
A PR that removes the owner field or sets severity: P5 (not a valid value in the schema) fails CI's schema-validation step with a specific, actionable error, rather than merging silently and leaving the on-call escalation path undefined for that runbook.
Trade-offs and pitfalls
- Over-engineering the schema so authoring a runbook feels like filing a support ticket discourages contributions; keep required fields minimal, owner, service, severity, last_reviewed, and let the rest stay freeform Markdown.
- Strict CI gating with owner review and on-call acknowledgment slows updates during an actual incident; document a fast-path for post-incident corrections that still requires review, just not before the fix ships in the moment.
- Building the search index without tying it to alerts leaves on-call still searching manually under pressure. The annotation link on the alert is what actually saves time; the index is the fallback.
Unlock Full Question Bank
Get access to all 41 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.