On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
The same incident keeps recurring every month despite repeated fixes. How would you run an RCA that surfaces the systemic process or tooling issue, rather than patching the same symptom again?
Sample Answer
When the same incident keeps recurring despite repeated fixes, the prior RCAs were almost certainly treating a symptom as the root cause; run this RCA by explicitly listing every prior "fix" and asking why each one didn't hold, because the pattern across failed fixes usually points straight at the real systemic gap.
Framework for a systemic RCA
- Build a fix history first, before investigating the current occurrence. For each prior incident of this same recurring issue: what was diagnosed as the cause, what was changed, and did the change actually address that diagnosis or just the immediate symptom? A pattern of "different symptom fixed each time, same underlying trigger every time" is the tell that root cause was never actually found.
- Separate the trigger from the vulnerability. The trigger (a specific deploy, a specific load pattern, a specific external dependency hiccup) may vary each month, but if the same class of trigger keeps causing an outage, the system has a standing vulnerability to that trigger class that no single fix removed. The RCA's job is to name the vulnerability, not just the latest trigger.
- Use the fishbone categories (code, config, infrastructure, process, tooling) to check whether every prior fix landed in the same category. If four consecutive fixes were all code patches but the incident keeps returning, that's evidence the real gap is in process or tooling (no regression test for this class of failure, no canary catching it before full rollout) rather than in any specific line of code.
- Test the systemic hypothesis, don't just assert it. If the hypothesis is "the nightly batch job and the backup window contend for the same database connection pool," verify by reproducing that contention in a controlled environment (isolated run of the batch job during a simulated backup window), not by pattern-matching from the incident timeline alone.
Worked example
A nightly batch job has caused three partial outages in three consecutive months. Prior fixes: month 1, increased the job's timeout (fix addressed "job was timing out"); month 2, added a retry with backoff (fix addressed "job failed transiently"); month 3 is the current incident, and the job is again failing, this time differently, a connection pool exhaustion error. Building the fix history shows a pattern: every fix targeted why the job failed on that specific night, and none asked why the job's failure mode changes every month while the timing (always during the nightly backup window) stays constant. Testing the systemic hypothesis, that the batch job and backup process share a connection pool and the backup's duration has been slowly growing as data volume grows, month-over-month backup duration logs confirm the backup window has grown from roughly 12 minutes to 40 minutes over the quarter, now overlapping the batch job's peak connection usage. That is the systemic cause: the job and backup were never intentionally isolated, and it was invisible for months because the backup was short enough not to overlap.
The durable fix follows from the systemic cause, not the latest symptom: allocate the batch job a dedicated connection pool separate from ad hoc processes, and alert on backup-duration trend (not just backup failure) so a slowly growing resource conflict is visible before it causes an outage again.
Trade-offs and pitfalls
The main pitfall is that a systemic RCA takes longer and produces a less satisfying immediate answer than "here's the line that broke," which creates pressure to ship another symptom-level fix under time pressure; the way to resist that is to make the fix-history review a required first step, not an optional deep-dive, so the systemic question gets asked before anyone commits to a scope. A real trade-off: broadening the RCA to a process or tooling gap usually means the fix is slower to land (new alerting, resource isolation, a regression test suite) than a code patch, so it's worth explicitly stating in the postmortem that a fast interim mitigation (in this example, a manual connection-pool bump) is being paired with the slower structural fix, rather than letting the slow fix block any near-term relief.
You have limited engineering capacity and a high on-call load from frequent alerts. How would you prioritize technical debt, alert tuning, and feature work over the next quarter to bring the pager volume down?
Sample Answer
Spend the first two weeks measuring where pages actually come from, not guessing, then rank the recurring drivers by pages eliminated per engineer-day of effort and fund the top of that list first. Feature work gets whatever capacity is left after the pager-volume target for the quarter is funded, not the other way around.
Step 1: baseline before you prioritize anything
Pull a two-week alert log and count pages by source, time of day, and whether each one required a real action or was noise. Prioritizing off memory or the loudest recent incident produces a list that optimizes for what people remember, not what's actually costing the most on-call time.
Step 2: score every recurring driver
| Item | Pages eliminated/month | Effort (engineer-days) | Score (pages/day) |
|---|---|---|---|
| Silence a flapping disk-alert threshold | 40 | 1 | 40.0 |
| Add retry/backoff to a flaky downstream call | 25 | 3 | 8.3 |
| Refactor alert routing to dedupe fan-out | 20 | 4 | 5.0 |
| Auto-remediate a stuck queue consumer | 15 | 5 | 3.0 |
| Full service redesign to remove the root cause | 10 | 15 | 0.7 |
Worked example: allocating a 20-day quarterly toil budget
Take the items in score order until the budget runs out: item 1 (1 day) + item 2 (3 days) + item 3 (4 days) + item 4 (5 days) = 13 of 20 days, leaving 7 days of buffer rather than starting the 15-day redesign this quarter.
Pages eliminated=40+25+20+15=100 pages/monthAgainst a baseline of 220 pages/month, that's a reduction of
220100×100≈45.5%funded by 13 of 20 available toil-reduction days, with the remaining 7 days as margin for whatever the alert log surfaces next.
Trade-offs and pitfalls
Silencing an alert to hit the score is only a win if it was genuinely non-actionable; verify that before suppressing it, since a silenced alert that was catching a real problem just moves the cost from "pages" to "undetected incidents." The full redesign scores lowest on pages-per-day but may be the only fix that prevents an outage-class failure; the score is an input to prioritization, not the whole decision, and a high-blast-radius item can justify funding even at a low score. Communicate the capacity trade explicitly to product as a quarter-long commitment rather than letting it get silently reprioritized sprint by sprint. Re-score the remaining list after each phase; fixing the top four items usually promotes a previously mid-ranked item to the top as the dominant offender changes.
How do you hand off an on-call shift so nothing falls through the cracks? What does a good handoff actually need to include?
Sample Answer
Direct answer
A good handoff transfers three things: current state (what's broken or at risk right now), context (what's already been tried and what's scheduled), and ownership (who is accountable for what next), and it has to be verifiable rather than a status dump, meaning the incoming engineer confirms they can actually act (they have access, they can reproduce the symptom, they understand the next step) before the outgoing engineer is done.
Structured elaboration
Handoff checklist, in order:
- Status snapshot: current owner and contact, "all green" or a one-line count of active incidents/unstable services.
- Active incidents: for each, severity, start time, impact, current owner, and a link to the ticket, not a re-explanation from scratch.
- Recent alert history: what's fired in the last several hours, and specifically any alert that's been flapping, since that's exactly what an incoming responder will misdiagnose as new if nobody flags it.
- Ongoing mitigations and runbook links: what's been tried, what's blocked, and the specific next action, with a link to the runbook rather than a paraphrase of it.
- Scheduled changes: upcoming deploys, migrations, or maintenance windows during the next shift, with rollback plans linked.
- Degraded-but-not-incident services: anything running hot or close to a threshold that isn't paging yet.
- Access and tooling: pager rotation, incident channel, dashboards, runbook repo, and an explicit note if the incoming engineer is missing access to any of them.
- Explicit confirmation, not implied: "Can you access the dashboards I linked? Can you reproduce the symptom in incident #123? Do you agree you own the DB migration follow-up?" Each gets an actual yes, not a thumbs-up emoji on a wall of text.
Making it lightweight (chatops mechanics). For day-to-day handoffs without an active incident, a short structured chat message beats a long document nobody reads: header (shift window, owner, escalation contact), one-line status, action items with owners, upcoming risks, and links. Anything in that message that turns out to matter beyond the shift boundary (a workaround that becomes permanent, a gotcha that will recur) gets tagged for follow-up and folded into the actual runbook within a day or two, so the team's durable documentation doesn't quietly live and die in chat history.
Automating the tedious part. The status snapshot, active-incident list, and recent-alert-history sections don't need to be typed by hand: a handoff template can be pre-populated from the monitoring and paging systems (current alert state, open incident IDs, last-deploy timestamp) so the outgoing engineer is editing and confirming pre-filled facts rather than writing a report from a blank page. That reduces both the time cost of handoff and the chance that something gets left out because the outgoing engineer forgot it existed.
Worked example
Friday, 6pm, end of a shift. One active Sev2 incident (checkout latency degraded for a subset of EU traffic, mitigation in progress: a feature flag was flipped to route around a slow dependency, error rate has dropped but root cause isn't fixed), and a database migration scheduled for 2am that night. The outgoing engineer posts:
Handoff | Fri 18:00-Sat 02:00 UTC
Owner: @outgoing -> @incoming | Escalation: @oncall-lead
Status: Degraded (checkout latency, EU) -> incident #482, mitigated not resolved
Action items:
- Watch checkout error rate; if it climbs above 2% again, re-check the feature flag is still on
- DB migration at 02:00 UTC (runbook: <link>, rollback: <link>) - I'll be asleep, this is yours
Risks: migration touches the same table implicated in incident #482; if latency spikes right after, check the migration first
Links: <dashboard> <incident #482> <migration runbook>
The incoming engineer confirms: dashboard access works, they can see incident #482's current state, and they explicitly acknowledge owning the migration watch. That confirmation, not the message itself, is what makes the handoff complete.
Trade-offs and pitfalls
- Pitfall: a "read the ticket" handoff with no verification step lets the incoming responder discover gaps at 3am instead of at 6pm when the person with context is still reachable.
- Pitfall: too much ceremony (a mandatory 45-minute call every single handoff) burns out the outgoing engineer and makes people avoid going on-call at all; reserve synchronous overlap for when there's an active Sev1/Sev2, not as the default for a quiet shift.
- Trade-off: synchronous handoff transfers tacit knowledge best but costs both people's time; async structured notes are cheaper but only as good as their template and discipline. The right default is async-by-default, with synchronous overlap triggered automatically whenever an incident is still open at shift boundary.
Two unrelated incidents hit different services at the same time. How do you decide how to allocate people across them, and when do you escalate to a higher-level incident commander?
Sample Answer
Allocate people by comparing the two incidents' business impact and required expertise, not by splitting the team evenly, and default to running them as separate incidents with separate commanders unless you find a shared root cause; escalate to a higher-level incident commander as soon as the resource conflict itself (not just the technical severity) becomes the bottleneck.
Allocating people across concurrent incidents
- Score each incident independently first: user-facing impact, revenue impact, data-integrity risk, and blast radius (one team's problem versus platform-wide). Two incidents rarely score identically, and the higher-scored one gets first claim on the strongest responders.
- Check for a shared root cause before splitting resources. If both services depend on the same failing component (a shared database, a shared auth service), that's actually one incident with two symptoms, and it should be run as a single incident with one commander, not two competing efforts pulling on the same underlying fix.
- Staff each independent incident with a minimum viable team: one incident commander, one primary responder, one communications owner. Resist over-staffing the incident that's louder or more visible if the other one is actually higher severity but quieter.
- Protect against double-booking the same expert. If one person is the only one who understands a shared piece of infrastructure, they can advise both incidents briefly but should not be the sole owner of fixing both; pull in a secondary responder even if slower.
When to escalate to a higher-level incident commander
Escalate when any of these is true, not just when severity is high:
- Resource contention itself is blocking progress: both incidents need the same scarce specialist or the same change-freeze exception, and someone above both incident commanders needs to arbitrate.
- Combined blast radius crosses an organizational boundary: the two incidents together affect enough of the business (multiple product lines, a shared customer segment) that unified external communication is needed, even if each incident alone wouldn't trigger that.
- One incident commander is starting to context-switch between both incidents. A single IC trying to run two incidents at once is a bigger risk than the incidents themselves; that's a signal to bring in a second commander or an overall coordinator, not to push through.
- Duration crosses a threshold where sustained dual-incident load starts to fatigue the responding team; a higher-level commander can pull in fresh responders or make the call to deprioritize the lower-severity incident explicitly (and communicate that decision) rather than let it silently starve.
Worked example
Two incidents fire eleven minutes apart: the checkout service is returning 500s for roughly a third of requests (revenue-impacting, high severity), and the internal analytics dashboard is showing stale data (no customer impact, low severity). The correct allocation: full incident-commander-plus-primary-plus-comms team goes to checkout immediately; analytics gets a single responder to investigate on a non-paging basis, because pulling more people onto analytics wouldn't shorten its resolution meaningfully and would strip capacity from checkout. If, twenty minutes in, the checkout investigation discovers the 500s trace back to the same message queue that feeds the analytics pipeline, the two incidents are merged under checkout's commander, because they share a root cause and running them separately would mean two people independently investigating the same queue.
Escalation in this example would trigger only if a third, unrelated incident arrived while checkout was still active and unresolved: three concurrent incidents makes single-IC-per-incident coordination itself the bottleneck, which is exactly the resource-contention trigger above.
Trade-offs and pitfalls
A common mistake is allocating headcount proportional to how loud or visible each incident is (how many people are asking about it in Slack) rather than its actual business impact; loud and low-impact will out-compete quiet and high-impact if you let it. Another is treating "escalate to a higher IC" as an admission of failure, so teams delay it past the point where a fresh coordinator would have resolved the resource conflict faster. The trade-off with merging incidents on a suspected shared root cause is real: merge too eagerly and you lose the separate investigation threads that might have found the divergence faster; the mitigation is to merge the coordination and communication, but keep separate technical workstreams until the shared cause is actually confirmed.
You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?
Sample Answer
Whether to fully automate, keep human-in-the-loop, or leave manual comes down to four questions applied to each specific action: how often does it happen, how bad is it if it goes wrong, can it be safely retried, and can success be verified automatically. High frequency, low blast radius, idempotent, and observable pushes toward full automation; anything destructive or hard to verify stays manual or gated behind a human, no matter how routine it feels.
The criteria
| Criterion | Favors automation | Favors manual / human-in-the-loop |
|---|---|---|
| Frequency | Happens often enough that manual toil adds up | Rare enough that automation investment doesn't pay back |
| Blast radius | Failure is contained (one instance, easily reverted) | Failure can be irreversible or affect data integrity broadly |
| Idempotency | Running it twice is harmless | Running it twice causes a different, possibly worse outcome |
| Verifiability | Success can be checked automatically (health check, row count) | Success requires human judgment to confirm |
Applying it to the three actions
Restarting a cached worker instance: high frequency, low blast radius (stateless, replaceable), fully idempotent, and easily verified with a health check. This is a strong automate candidate: drain connections, spin up a replacement, run a health check, cut traffic over, roll back automatically if the health check fails.
Reattaching a detached volume: lower frequency, meaningfully higher blast radius (attaching to the wrong instance or double-attaching can corrupt data), and only moderately idempotent, reattaching twice isn't necessarily safe. This sits in the middle: automate the pre-checks and the mechanical steps (verify volume ID, verify target instance, snapshot before attaching), but require a human to confirm before the final attach executes.
Running a database schema migration: low frequency, high blast radius (can be destructive and hard to reverse), low idempotency for anything involving DDL (data definition language: schema-altering SQL statements like ALTER TABLE), and success often isn't verifiable by a simple automated check, it needs someone to look at whether the data actually came out right. This stays manual, or more precisely, human-gated: automation handles the mechanical parts (schema diff, pre-migration validation, backup, staged rollout to a canary), but a person approves the production apply.
The underlying argument for phasing automation in gradually
Automating a step doesn't just remove toil, it also removes the moment a human would have caught something unusual about this particular instance of the problem. That's fine for the worker-restart case, where "unusual" mostly doesn't exist, but risky for the migration case, where every migration is a little different. The practical path is phasing: run a new automation in shadow mode first (it proposes the action but a human executes), then human-in-the-loop (it executes after one-click approval), and only promote to full automation once it has a track record across enough real incidents that its false-positive and false-negative rate are actually known, not assumed.
Guarding automated actions with least privilege
Whatever is automated should run with only the permissions that specific action needs, a worker-restart automation shouldn't hold credentials that could also run a schema migration, and every automated action should be logged with who (or what) triggered it and why. For the human-in-the-loop tier, the approval step itself should require a specific person's action (not a shared bot token anyone can trigger), so there's a real approval trail, not a rubber stamp.
Trade-offs and pitfalls
The common wrong turn is automating based on how annoying a task feels rather than how safe it is, restarting workers manually is annoying but safe to automate; migrations are also annoying, but the annoyance is not the variable that should decide it. The other pitfall is leaving a human-in-the-loop step gated behind an approval that nobody actually reads before clicking, if the approval doesn't include enough context (what will run, what's the blast radius, what's the rollback) to make a real judgment, it's automation with an extra click, not a genuine safety gate.
Unlock Full Question Bank
Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.