On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
How would you design a fair approach to compensating engineers for on-call work, balancing pay, time off in lieu, and rotation length?
Sample Answer
Compensation and schedule design are two separate levers, and both need to move: pay a base standby stipend for availability, add per-incident pay or time-off-in-lieu for the work actually done, and size the rotation and rest guarantees so the schedule itself isn't relying on money to make an unsustainable load tolerable.
Compensation model components
| Component | What it covers | Typical structure | Why it's separate |
|---|---|---|---|
| Standby stipend | Being reachable and ready, whether or not paged | Fixed weekly amount | Compensates the constraint on personal time even in a quiet week |
| Per-incident pay or TOIL | Actual time spent responding | Hourly rate, or banked time at 1x to 2x | Rewards work done and discourages treating pages as free to the business |
| Leveling credit | Career recognition for on-call excellence | Counted explicitly in review/promotion criteria | Stops strong on-call performers from being penalized for time not spent on visible project work |
| Rest guarantee | Recovery time | Mandatory hours off after a heavy incident or night shift | Protects sustainability independent of pay |
Worked example: one on-call week
Stipend $250 + 6 hours of actual incident work at $40/hr:
Pay=250+(6×40)=250+240=$490If the same 6 hours bank as TOIL at a 1.5x rate for after-hours work:
TOIL banked=6×1.5=9 hoursSchedule practices that reduce the load pay has to compensate for
Primary/secondary tiers so one page doesn't always land on the same person; a cap on consecutive on-call weeks per engineer; shorter rotations (fewer consecutive days of stress, more handoffs) traded against longer rotations (fewer handoffs, more concentrated fatigue), sized to team headcount rather than picked arbitrarily; and follow-the-sun coverage once the team is large and distributed enough to make timezone handoffs cheaper than overnight pages.
Trade-offs and pitfalls
Per-incident pay can invite gaming in both directions, either padding logged hours or under-reporting to avoid looking like a "high maintenance" service; review incident-hour claims against the paging log rather than trusting self-reports alone. Contractors and salaried employees often need different structures (cash versus TOIL), and a single model rarely fits both cleanly. Regional labor law varies significantly, some jurisdictions treat standby time itself as compensable working time, so confirm with legal or HR before setting a global policy rather than assuming one region's rules generalize. There's no dominant answer on rotation length; it has to be sized to team size and incident frequency, not copied from another team's policy.
Why do runbooks tend to go stale in a large engineering org? What are the common root causes, and what would you actually do about each one?
Sample Answer
Direct answer
Runbooks go stale because nothing automatically ties them to the systems they describe: ownership is unclear, updates aren't triggered by the changes that invalidate them, and nobody is rewarded for maintaining them, so they drift silently until an incident exposes it. The fix for each cause is the same shape: build the update into a workflow that already has to happen, like a deploy, a PR review, or a drill, rather than relying on someone remembering.
Structured elaboration
| Root cause | Why it happens | What to actually do |
|---|---|---|
| No clear owner | Docs feel like everyone's job, so they end up being no one's | Assign a named owner (team and person) per runbook, visible on the doc itself |
| No trigger tied to system changes | Infra or config changes ship without a linked doc update | Require a runbook-touch check in review for infra changes that affect the documented procedure |
| Fragmented across tools | The same procedure exists in a wiki, a chat pin, and a repo, and they diverge | One canonical source, docs-as-code in git; other tools link to it instead of duplicating it |
| Hard to edit | Binary or WYSIWYG pages discourage small fixes | Markdown in git with a low-friction pull-request flow |
| Never verified | Nobody runs the steps until a real incident forces it | Scheduled tabletop or game-day drills that surface breakage before it matters |
| Incentives favor code over docs | Engineers are measured on features shipped, not documentation kept accurate | Include doc currency in the definition of done or the on-call handoff checklist |
Worked example
A payments team migrates from a single database instance to a managed cluster with a different failover tool. The failover runbook still references the old promote command. Nobody touches the runbook because the migration's review process had no requirement to touch documentation tied to it, which is exactly the "no trigger tied to system changes" row above. Months later, an on-call engineer hits a real primary failure, runs the stale command, gets an error, and has to rediscover the correct procedure live instead of following a runbook that already had it. The root cause traces cleanly to the missing trigger, not to the engineer who wrote the original doc.
Trade-offs and pitfalls
- Quarterly "please review this doc" reminders without a named owner tend to become checkbox theater: marked reviewed without anyone actually re-verifying the steps.
- Gating merges on documentation updates adds friction to every infra change; scope the gate to changes that touch a documented procedure specifically, or teams will route around it entirely.
The same incident keeps recurring every month despite repeated fixes. How would you run an RCA that surfaces the systemic process or tooling issue, rather than patching the same symptom again?
Sample Answer
When the same incident keeps recurring despite repeated fixes, the prior RCAs were almost certainly treating a symptom as the root cause; run this RCA by explicitly listing every prior "fix" and asking why each one didn't hold, because the pattern across failed fixes usually points straight at the real systemic gap.
Framework for a systemic RCA
- Build a fix history first, before investigating the current occurrence. For each prior incident of this same recurring issue: what was diagnosed as the cause, what was changed, and did the change actually address that diagnosis or just the immediate symptom? A pattern of "different symptom fixed each time, same underlying trigger every time" is the tell that root cause was never actually found.
- Separate the trigger from the vulnerability. The trigger (a specific deploy, a specific load pattern, a specific external dependency hiccup) may vary each month, but if the same class of trigger keeps causing an outage, the system has a standing vulnerability to that trigger class that no single fix removed. The RCA's job is to name the vulnerability, not just the latest trigger.
- Use the fishbone categories (code, config, infrastructure, process, tooling) to check whether every prior fix landed in the same category. If four consecutive fixes were all code patches but the incident keeps returning, that's evidence the real gap is in process or tooling (no regression test for this class of failure, no canary catching it before full rollout) rather than in any specific line of code.
- Test the systemic hypothesis, don't just assert it. If the hypothesis is "the nightly batch job and the backup window contend for the same database connection pool," verify by reproducing that contention in a controlled environment (isolated run of the batch job during a simulated backup window), not by pattern-matching from the incident timeline alone.
Worked example
A nightly batch job has caused three partial outages in three consecutive months. Prior fixes: month 1, increased the job's timeout (fix addressed "job was timing out"); month 2, added a retry with backoff (fix addressed "job failed transiently"); month 3 is the current incident, and the job is again failing, this time differently, a connection pool exhaustion error. Building the fix history shows a pattern: every fix targeted why the job failed on that specific night, and none asked why the job's failure mode changes every month while the timing (always during the nightly backup window) stays constant. Testing the systemic hypothesis, that the batch job and backup process share a connection pool and the backup's duration has been slowly growing as data volume grows, month-over-month backup duration logs confirm the backup window has grown from roughly 12 minutes to 40 minutes over the quarter, now overlapping the batch job's peak connection usage. That is the systemic cause: the job and backup were never intentionally isolated, and it was invisible for months because the backup was short enough not to overlap.
The durable fix follows from the systemic cause, not the latest symptom: allocate the batch job a dedicated connection pool separate from ad hoc processes, and alert on backup-duration trend (not just backup failure) so a slowly growing resource conflict is visible before it causes an outage again.
Trade-offs and pitfalls
The main pitfall is that a systemic RCA takes longer and produces a less satisfying immediate answer than "here's the line that broke," which creates pressure to ship another symptom-level fix under time pressure; the way to resist that is to make the fix-history review a required first step, not an optional deep-dive, so the systemic question gets asked before anyone commits to a scope. A real trade-off: broadening the RCA to a process or tooling gap usually means the fix is slower to land (new alerting, resource isolation, a regression test suite) than a code patch, so it's worth explicitly stating in the postmortem that a fast interim mitigation (in this example, a manual connection-pool bump) is being paired with the slower structural fix, rather than letting the slow fix block any near-term relief.
A runbook's automated remediation step ran and it caused a partial outage instead of fixing anything. How would you investigate what went wrong, and what would you change to prevent it from happening again?
Sample Answer
When an automated remediation makes things worse, the first move is to stop trusting the automation, not to debug it live: disable the trigger (feature flag or scheduler pause) so it can't fire again while you investigate, then treat the automation's own actions as the incident's primary evidence trail.
Investigation approach
- Pull the automation's own audit log first. What did it decide to do, on what input, at what timestamp? Most remediation frameworks log the triggering condition and the action taken; if this one doesn't, that's itself a finding.
- Reconstruct the precondition it evaluated against. Was the health check it used stale (cached metrics, delayed scrape) or narrower than reality (checked one replica's health, not cluster quorum)?
- Check for concurrency. Did two instances of the same remediation run at once, or did it run while a human was mid-deploy? Interleaved writes to the same resource are a common cause of "fix that broke things."
- Diff the assumed environment against the actual one. Runbooks and remediation scripts encode assumptions (resource names, API versions, cluster topology) that drift silently; check whether the automation was written against a topology that's since changed.
Framework for the fix
- Add a pre-check gate: the remediation must verify the system is in the state it assumes (quorum present, no in-flight deploy, dependency healthy) before acting, and abort loudly if not.
- Make the action idempotent and reversible: re-running it, or running it against a system already in the target state, should be a no-op, and every destructive step needs a paired rollback.
- Bound the blast radius: act on one node/instance first (canary), verify success, then proceed, rather than acting cluster-wide in one shot.
- Add a concurrency guard: a lock or lease so two triggers of the same remediation can't run simultaneously.
- Gate high-impact actions behind a second signal: require the automation to see the problem confirmed by two independent signals (e.g., an alert plus a direct health check) before taking a destructive action, not just one noisy metric.
Worked example
Suppose the remediation is: "if a node reports high memory for 3 consecutive scrapes, cordon and drain it." The postmortem finds the metrics scraper had a 90-second collection lag during a load spike, so by the time the automation cordoned the third node it was actually reading data that was already 4.5 minutes stale (three 90-second-lagged scrapes), and it drained three nodes in the same 2-minute window because the memory spike was cluster-wide, not node-specific. Losing three nodes at once dropped the cluster below quorum for its replicated service, which is the partial outage.
The fix that follows directly from that trace: (a) the pre-check should compare current live memory, not the lagged scrape, before acting; (b) the automation should check how many nodes it has already drained in the current window and refuse to exceed a cap (e.g., no more than one node per 10 minutes) until a human confirms; (c) it should check that the remaining fleet still satisfies quorum before draining another node.
Trade-offs and pitfalls
Adding pre-checks and rate caps makes the remediation slower to react, which is the right trade for anything that can cause an outage of its own; reserve fully unthrottled auto-remediation for actions that are cheap to reverse (like restarting a single stateless pod) and keep caps and human gates on anything that removes capacity or touches shared state. A common wrong turn is to respond to this incident by simply disabling the automation permanently and reverting to manual remediation: that trades a rare automation bug for a much larger population of slower, inconsistent manual responses. The senior move is to narrow what the automation is trusted to do unsupervised, not to abandon automation.
What's the difference between a runbook and a playbook, and when would you reach for one instead of the other?
Sample Answer
Direct answer
A runbook is a fixed set of step-by-step instructions for a known failure mode: it tells you exactly what to type. A playbook is a decision framework for handling an incident more broadly: it tells you how to figure out what to do, who needs to be involved, and which runbook to reach for. Reach for a runbook when you already know the cause and the fix is mechanical; reach for a playbook when you're still diagnosing, coordinating multiple people, or the right response depends on judgment.
Structured elaboration
| Runbook | Playbook | |
|---|---|---|
| Scope | One specific, known failure mode or task | A class of incidents, or the overall response process |
| Format | Linear, prescriptive steps | Decision tree or branching guidance |
| Answers | "What do I type" | "What do I decide, and who do I involve" |
| Typical contents | Preconditions, exact commands, verification steps, rollback | Severity thresholds, roles (IC, comms lead), escalation matrix, links to runbooks |
| Usually owned by | The team that owns the specific service | Incident response leadership or SRE |
| Reviewed when | The underlying system changes | The org's escalation structure or tooling changes |
Minimum fields for each:
- Runbook: title, service, owner, trigger/precondition, required permissions and tools, exact step-by-step commands, verification steps, rollback steps, expected impact, last-reviewed date.
- Playbook: title, scope and severity thresholds, incident commander and stakeholder roles, the decision tree itself, links to the relevant runbooks, communication templates, escalation matrix, last-reviewed date.
Worked example
A database replica's lag exceeds a threshold: this is a runbook. It lists the exact commands to promote a replica, the steps to reconfigure the application's read preference, verification queries to confirm the fix, and the rollback commands if the promotion causes a new problem. Now compare that to a major outage affecting payments: this is a playbook. It guides the incident commander through detecting the actual scope, declaring severity, deciding between routing traffic to a fallback payment path versus draining traffic entirely, coordinating the app, infra, and comms teams, and linking out to the specific runbooks (including the replica-promotion one, if that turns out to be the fix) for whichever technical action the decision tree leads to.
Trade-offs and pitfalls
- Pitfall: writing a "runbook" for something that actually needs judgment, for example "when in doubt, restart the service." That hides a decision a playbook should make explicit, and someone follows it verbatim during exactly the incident it doesn't fit.
- Pitfall: letting a playbook go stale is worse than letting one runbook go stale, because the playbook is what everyone reaches for first during ambiguity; a stale escalation matrix (wrong names or numbers) breaks the whole response, not just one specific fix.
- Trade-off: automating a runbook into a one-click execution is great for high-confidence, low-blast-radius fixes (restarting a stateless service) and dangerous for high-blast-radius ones (promoting a database replica). The more damage a wrong click can do, the more the runbook should require an explicit human confirmation step before executing, not less.
- Both belong in a versioned, reviewed repository rather than an unowned wiki page, with periodic review, and runbooks specifically benefit from occasional dry-run or game-day testing to confirm the steps still work against the current system rather than an outdated one.
Unlock Full Question Bank
Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.