On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?
Sample Answer
Whether to fully automate, keep human-in-the-loop, or leave manual comes down to four questions applied to each specific action: how often does it happen, how bad is it if it goes wrong, can it be safely retried, and can success be verified automatically. High frequency, low blast radius, idempotent, and observable pushes toward full automation; anything destructive or hard to verify stays manual or gated behind a human, no matter how routine it feels.
The criteria
| Criterion | Favors automation | Favors manual / human-in-the-loop |
|---|---|---|
| Frequency | Happens often enough that manual toil adds up | Rare enough that automation investment doesn't pay back |
| Blast radius | Failure is contained (one instance, easily reverted) | Failure can be irreversible or affect data integrity broadly |
| Idempotency | Running it twice is harmless | Running it twice causes a different, possibly worse outcome |
| Verifiability | Success can be checked automatically (health check, row count) | Success requires human judgment to confirm |
Applying it to the three actions
Restarting a cached worker instance: high frequency, low blast radius (stateless, replaceable), fully idempotent, and easily verified with a health check. This is a strong automate candidate: drain connections, spin up a replacement, run a health check, cut traffic over, roll back automatically if the health check fails.
Reattaching a detached volume: lower frequency, meaningfully higher blast radius (attaching to the wrong instance or double-attaching can corrupt data), and only moderately idempotent, reattaching twice isn't necessarily safe. This sits in the middle: automate the pre-checks and the mechanical steps (verify volume ID, verify target instance, snapshot before attaching), but require a human to confirm before the final attach executes.
Running a database schema migration: low frequency, high blast radius (can be destructive and hard to reverse), low idempotency for anything involving DDL (data definition language: schema-altering SQL statements like ALTER TABLE), and success often isn't verifiable by a simple automated check, it needs someone to look at whether the data actually came out right. This stays manual, or more precisely, human-gated: automation handles the mechanical parts (schema diff, pre-migration validation, backup, staged rollout to a canary), but a person approves the production apply.
The underlying argument for phasing automation in gradually
Automating a step doesn't just remove toil, it also removes the moment a human would have caught something unusual about this particular instance of the problem. That's fine for the worker-restart case, where "unusual" mostly doesn't exist, but risky for the migration case, where every migration is a little different. The practical path is phasing: run a new automation in shadow mode first (it proposes the action but a human executes), then human-in-the-loop (it executes after one-click approval), and only promote to full automation once it has a track record across enough real incidents that its false-positive and false-negative rate are actually known, not assumed.
Guarding automated actions with least privilege
Whatever is automated should run with only the permissions that specific action needs, a worker-restart automation shouldn't hold credentials that could also run a schema migration, and every automated action should be logged with who (or what) triggered it and why. For the human-in-the-loop tier, the approval step itself should require a specific person's action (not a shared bot token anyone can trigger), so there's a real approval trail, not a rubber stamp.
Trade-offs and pitfalls
The common wrong turn is automating based on how annoying a task feels rather than how safe it is, restarting workers manually is annoying but safe to automate; migrations are also annoying, but the annoyance is not the variable that should decide it. The other pitfall is leaving a human-in-the-loop step gated behind an approval that nobody actually reads before clicking, if the approval doesn't include enough context (what will run, what's the blast radius, what's the rollback) to make a real judgment, it's automation with an extra click, not a genuine safety gate.
How do you make sure postmortem action items actually get done, and that lessons from one incident reach the teams who didn't experience it directly?
Sample Answer
Direct answer
Action items get done when they're tracked in the same system and cadence as regular engineering work rather than in a postmortem doc nobody revisits, each has exactly one named owner and a due date, and a recurring, lightweight review surfaces overdue items instead of letting them go quiet. Lessons reach teams that didn't experience the incident when postmortems are indexed centrally by system and failure pattern, not just filed per-team, and a short summary is actively pushed to adjacent teams rather than waiting for someone to search the archive.
Structured elaboration
Getting action items done
Categorize each root cause as process, technical, or people and communication before assigning it. A technical fix routes to the engineering backlog with the same prioritization as other work; a process fix might mean updating a runbook or a review gate; making this distinction up front stops "add more monitoring" from becoming the default answer to everything. Each item gets one owner, one due date, and one acceptance criterion, since items without a single owner reliably don't get done. A weekly or sprint-cadence review of all open incident action items across teams escalates anything overdue to the owning manager, not just the assignee, and the remediation deadline itself scales with the severity of the incident that produced it.
Spreading lessons beyond the team that had the incident
A central, searchable postmortem index tagged by system, failure mode, and root-cause category, not filed only under the owning team's folder, makes cross-cutting patterns findable. A short, standardized summary, what broke, why, what changed, gets pushed to a cross-team channel or a recurring digest rather than relying on other teams to search for it. If the same root-cause category, such as "connection-pool exhaustion" or "no alert on saturation," shows up across multiple postmortems, that's a systemic gap worth its own initiative, and it only becomes visible if postmortems are tagged consistently enough to query across them.
Worked example
A quarter's incident action-item tracker opens 12 action items from postmortems and closes 9 of them within their assigned deadline by end of quarter, a 9/12=0.75, or 75 percent, on-time closure rate. The weekly review flags the 3 still-open items each week; two turn out to be cross-team items stuck on a dependency from a third team that was never explicitly assigned ownership of its part. That gap itself becomes a new process action item: cross-team action items need an explicit dependency owner, not just an assignee, added to the next postmortem review template.
Trade-offs and pitfalls
- Tracking action items only in the postmortem document itself means they compete for attention against the team's regular backlog and usually lose; they need to live in the same system as everything else.
- Broadcasting every postmortem summary to the whole organization trains people to ignore the channel; scope the push to teams that own adjacent or dependent systems, and keep a searchable index for everyone else.
- A strict remediation deadline drives closure but can pressure a team into a shallow fix just to hit the date; pairing the deadline with an explicit acceptance criterion is what keeps "closed" meaning "verified" rather than just "closed."
How would you build a cost-benefit case for automating a recurring operational task, rather than continuing to have engineers handle it manually?
Sample Answer
A credible automation case has three parts: the annual cost of doing the task manually (engineer time plus any revenue or SLA impact), the fully loaded cost of building and maintaining the automation, and a residual-risk line for when the automation itself misfires. Turn those into a payback period and an annual ROI, then restate the same numbers for a non-technical audience as "what we spend today versus a one-time investment that pays for itself in N months."
Inputs the model needs
| Variable | Meaning |
|---|---|
| F | Incidents per year requiring this manual task |
| H | Manual engineer-hours per incident |
| C | Fully loaded hourly cost of the engineer |
| R | Revenue or SLA impact per incident |
| D | One-time development cost of the automation |
| N | Years over which you amortize D |
| M | Annual maintenance cost of the automation |
| p | Probability the automation misfires per run |
| C_f | Cost when it misfires (human cleanup time plus any extra impact) |
The formula
ToilannualAutomationannualResidualannualNet savingsPayback (months)=F×(H×C+R)=ND+M=F×p×Cf=Toilannual−Automationannual−Residualannual=Net savings/12DWorked example: a noisy alert that fires 300 times a year
Assumptions (pinned): F = 300/year, H = 0.4 hours, C = $75/hr, R = $100/incident, D = $20,000, N = 2 years, M = $4,000/year, p = 2%, C_f = $175/misfire.
ToilannualAutomationannualResidualannualNet savingsPaybackROI=300×(0.4×75+100)=300×130=$39,000=220,000+4,000=$14,000=300×0.02×175=$1,050=39,000−14,000−1,050=$23,950=23,950/1220,000=1,995.8320,000≈10.0 months=14,00023,950×100≈171%Translating it for a non-technical stakeholder
The pitch a CFO or exec sponsor needs is the same numbers, stripped of the formula: "this task currently costs the team about $39k a year in engineer time and SLA impact. A $20k investment, amortized over two years plus $4k a year to keep it working, pays for itself in about ten months and then keeps saving roughly $24k every year after that, with a 2% chance any given automated run still needs a human to clean up after it." Leading with the payback date and the residual-risk number, not just the savings, is what makes the pitch credible instead of salesy.
Trade-offs and pitfalls
If F is seasonal or unstable, a single annual estimate is misleading; revisit the model quarterly rather than trusting a number computed once. Don't drop the residual-failure term to make the case look better: an automation that works 98% of the time but corrupts state on the other 2% can be a net negative if C_f is large, especially for destructive actions. A common wrong turn is pitching automation using only "engineer-hours saved" and skipping maintenance and residual-risk costs entirely; that consistently overstates ROI and makes the next proposal harder to trust once someone notices the gap.
Some runbook steps involve sensitive actions, like production database admin commands or rotating credentials. How do you control who can run those steps and keep it auditable, without slowing a responder down during a real P1?
Sample Answer
Don't gate sensitive steps behind standing credentials a responder already holds. Gate them behind short-lived, narrowly scoped credentials issued by the runbook orchestrator at the moment of use, with a pre-authorized fast path for the highest severities so speed during a real incident doesn't require quietly bypassing the audit trail.
Comparing access models
| Model | Speed during a P1 | Auditability | Blast radius if leaked |
|---|---|---|---|
| Shared static credential in a vault everyone can read | Fast | Poor, can't tell who actually used it | High, valid indefinitely until manually rotated |
| Manual per-use approval (ticket plus human sign-off) | Slow, adds minutes exactly when they're scarce | Good | Low, but the delay is itself a cost during a P1 |
| Just-in-time ephemeral credential (vault-issued, scoped, short TTL, auto-revoked) | Fast for pre-authorized P1 paths | Excellent, tied to identity, ticket, and TTL window | Low, expires on its own even if forgotten |
Worked example: sizing the credential TTL
If the median observed time to complete a given remediation step across past incidents is 12 minutes, a 15-minute TTL leaves almost no margin:
Margin=15−12=3 mina responder who hits a snag is interrupted 3 minutes short of done, exactly when stopping is most disruptive. A TTL of roughly 20 minutes, the median plus a working buffer rather than an open-ended grant, gives room to finish without leaving a long-lived credential outstanding. A 4-hour TTL "to be safe" instead means a credential compromised from a responder's terminal during that window stays valid for the rest of the shift; that's the trade being made for the extra convenience.
Keeping it auditable without slowing the responder down
- Runbooks reference secret IDs, never raw values, so the document itself is safe to read even if it leaks.
- The orchestrator executes the sensitive step server-side where practical, so the responder never sees the decrypted secret at all, only the outcome.
- Every credential issuance logs identity, ticket or incident ID, scope, and TTL to an immutable log, correlated automatically rather than reconstructed after the fact.
- The highest severities get pre-authorized issuance, no waiting on a human approver, precisely because the TTL and logging, not a manual gate, are what keep it auditable.
Trade-offs and pitfalls
A break-glass path needs more audit rigor than the normal path, not less; pair any emergency bypass with mandatory post-incident review and automatic rotation of whatever it touched. Auto-approval for the highest severities removes a human gate exactly during the highest-risk window (a real incident, adrenaline, and possibly an actor exploiting the chaos), so the TTL and logging have to carry that weight instead. Orchestrator-executed remediation is safer for the responder but adds its own risk surface; the automation itself now needs the same change-review rigor as production code, not less because "it's just a script."
What's the difference between a runbook and a playbook, and when would you reach for one instead of the other?
Sample Answer
Direct answer
A runbook is a fixed set of step-by-step instructions for a known failure mode: it tells you exactly what to type. A playbook is a decision framework for handling an incident more broadly: it tells you how to figure out what to do, who needs to be involved, and which runbook to reach for. Reach for a runbook when you already know the cause and the fix is mechanical; reach for a playbook when you're still diagnosing, coordinating multiple people, or the right response depends on judgment.
Structured elaboration
| Runbook | Playbook | |
|---|---|---|
| Scope | One specific, known failure mode or task | A class of incidents, or the overall response process |
| Format | Linear, prescriptive steps | Decision tree or branching guidance |
| Answers | "What do I type" | "What do I decide, and who do I involve" |
| Typical contents | Preconditions, exact commands, verification steps, rollback | Severity thresholds, roles (IC, comms lead), escalation matrix, links to runbooks |
| Usually owned by | The team that owns the specific service | Incident response leadership or SRE |
| Reviewed when | The underlying system changes | The org's escalation structure or tooling changes |
Minimum fields for each:
- Runbook: title, service, owner, trigger/precondition, required permissions and tools, exact step-by-step commands, verification steps, rollback steps, expected impact, last-reviewed date.
- Playbook: title, scope and severity thresholds, incident commander and stakeholder roles, the decision tree itself, links to the relevant runbooks, communication templates, escalation matrix, last-reviewed date.
Worked example
A database replica's lag exceeds a threshold: this is a runbook. It lists the exact commands to promote a replica, the steps to reconfigure the application's read preference, verification queries to confirm the fix, and the rollback commands if the promotion causes a new problem. Now compare that to a major outage affecting payments: this is a playbook. It guides the incident commander through detecting the actual scope, declaring severity, deciding between routing traffic to a fallback payment path versus draining traffic entirely, coordinating the app, infra, and comms teams, and linking out to the specific runbooks (including the replica-promotion one, if that turns out to be the fix) for whichever technical action the decision tree leads to.
Trade-offs and pitfalls
- Pitfall: writing a "runbook" for something that actually needs judgment, for example "when in doubt, restart the service." That hides a decision a playbook should make explicit, and someone follows it verbatim during exactly the incident it doesn't fit.
- Pitfall: letting a playbook go stale is worse than letting one runbook go stale, because the playbook is what everyone reaches for first during ambiguity; a stale escalation matrix (wrong names or numbers) breaks the whole response, not just one specific fix.
- Trade-off: automating a runbook into a one-click execution is great for high-confidence, low-blast-radius fixes (restarting a stateless service) and dangerous for high-blast-radius ones (promoting a database replica). The more damage a wrong click can do, the more the runbook should require an explicit human confirmation step before executing, not less.
- Both belong in a versioned, reviewed repository rather than an unowned wiki page, with periodic review, and runbooks specifically benefit from occasional dry-run or game-day testing to confirm the steps still work against the current system rather than an outdated one.
Unlock Full Question Bank
Get access to all 41 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.