On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?
Sample Answer
Direct answer
Frame the sensitivity-versus-fatigue tension as an explicit cost trade-off rather than a vibes call: assign a rough relative cost to a missed incident versus a false page, pick the threshold that minimizes expected cost given the current false-positive and false-negative rates, and never roll a threshold change straight to paging. Validate it in shadow mode against real traffic first, then canary it on a subset before a full rollout, with an automatic rollback trigger if things get worse.
Structured elaboration
Cost framing. At a candidate threshold τ, define expected cost as:
cost(τ)=CFN⋅P(miss∣τ)+CFP⋅P(false page∣τ)Lowering τ (more sensitive) drives the probability of a miss toward zero but raises the false-page rate, and raising τ does the opposite. The right threshold is wherever this sum is smallest, not wherever either rate alone looks best in isolation.
Safe testing method, three stages with explicit gates:
- Shadow mode: the candidate threshold runs log-only, never pages, and every alert it would have fired gets compared against the real incident record after the fact. No production risk, but also no real-time responder feedback.
- Canary: the candidate threshold pages for real, but only for a subset of services or regions. Watch for missed-detection signals and responder load on that subset before touching anything else.
- Gradual rollout: feature-flagged expansion to the rest of the fleet, with an automatic rollback trigger if the false-positive rate or acknowledgment latency regresses past a predefined bound.
Worked example
Say a one-week shadow test compares the current threshold (A) against a stricter candidate (B) on the same underlying traffic, and every alert is later labeled against the real incident record:
| Threshold | Total alerts | True incidents caught (TP) | False pages (FP) | Missed incidents (FN) |
|---|---|---|---|---|
| A (current) | 200 | 18 | 182 | 2 |
| B (candidate) | 40 | 16 | 24 | 4 |
Assign a rough relative cost: a missed incident costs 500 responder-hour-equivalents (CFN=500), a false page costs 1 (CFP=1). Expected cost per threshold:
costA=CFN⋅FNA+CFP⋅FPA=500×2+1×182=1182 costB=CFN⋅FNB+CFP⋅FPB=500×4+1×24=2024Even though B has far higher precision (16 of 40 alerts were real, versus 18 of 200 for A), A has the lower expected cost: B's two extra missed incidents cost more than the 158 extra false pages A generates. If CFN were much closer to CFP, for example a low-stakes internal tool where a miss is only mildly annoying rather than expensive, B would win instead. The point of doing the arithmetic rather than eyeballing the false-positive rate is that the right threshold depends on getting the relative cost of a miss right for that specific service, not on chasing a universally "less noisy" target.
Trade-offs and pitfalls
- Pitfall: picking CFN and CFP once and never revisiting them. Relative costs shift as customer scale, contractual SLAs, and the service's blast radius change, so the cost model needs the same periodic review as the severity scale it feeds into.
- Pitfall: shadow-testing only against the same traffic period used to design the threshold in the first place, which overfits the result; test against a held-out period the threshold wasn't tuned on.
- Trade-off: assigning explicit costs forces an uncomfortable conversation with stakeholders about what a missed incident is actually worth, but that discomfort produces a threshold the team actually agreed to, rather than one that came from a single engineer's gut feel nobody was consulted on.
During a postmortem, the incident commander singles out one engineer as the cause of the outage. How do you respond in the moment to preserve a blameless culture, without letting accountability for the fix slide?
Sample Answer
In the moment, redirect the conversation from the person to the timeline: acknowledge what was said without amplifying it, then immediately steer the group back to reconstructing what happened and why the system allowed it, while making clear that accountability for the fix is not going away.
In-the-moment response
- Interrupt with a redirect, not a confrontation. Something like: "Let's hold on names for a second and walk the timeline: what did the system show at each step?" This isn't ignoring what was said; it's refusing to let the postmortem's structure reward the blame framing by continuing down that thread.
- Reframe the specific claim into a system question. If the IC (Incident Commander, the person directing the response) says "this happened because Priya deployed without checking the dashboard," the redirect is: "So the deploy process didn't require a dashboard check before going out. Is that a gap in the checklist, or did the checklist exist and get skipped? Either answer tells us what to fix." This keeps the factual content (a deploy went out without a check) while stripping the blame framing.
- Do not let it pass silently either. Staying quiet when a peer is singled out in front of the team reads as agreement, and it's the fastest way to make the next engineer afraid to be transparent in their own postmortem. A short, calm correction in the room is better than a private word afterward, because the damage (and the culture signal) happened publicly.
- Follow up with the IC privately, separate from the room. The public redirect handles the moment; a private conversation afterward addresses the pattern, especially if this IC does it repeatedly.
Keeping accountability intact
Blameless does not mean no one owns the fix. The distinction to hold onto:
- Blame assigns fault for what already happened, to a person, and looks backward.
- Accountability assigns ownership for what happens next, to a role or system, and looks forward.
So the postmortem should still end with a named owner for each remediation item (a person, because someone has to actually do the work) and a deadline, but the framing is "you're the best person to close this gap because you understand the deploy path," not "this is your fault so you have to fix it." The action items get assigned based on who has the context and capability, independent of who gets blamed.
Worked example
During a payments-outage postmortem, the IC says: "Marcus rolled back the config and that's what caused the second outage." The redirect: "Let's look at what the rollback runbook told him to check before rolling back. Did it call out this specific config's downstream dependency?" The team pulls up the runbook and finds it didn't mention that this particular config was read by two other services; the rollback step existed, but the pre-check for downstream impact didn't. The postmortem action items become: (1) add a downstream-dependency check to the rollback runbook for this config, owned by the platform team, due in two weeks, and (2) audit other high-fanout configs for the same missing check, owned by Marcus, since he now has the clearest picture of what that gap looks like, due in one month. Marcus ends up with an action item, but it's framed as "you're positioned to close this" rather than "you caused this," and the runbook gap, not Marcus's judgment, is recorded as the finding.
Trade-offs and pitfalls
The main pitfall is overcorrecting into vagueness, where "blameless" gets used to avoid naming any specific decision point, and the postmortem ends up too soft to actually change anything; the fix is to be precise about the decision and the missing guardrail while staying impersonal about who made the decision. A related pitfall specific to this scenario: correcting an incident commander in front of the team carries real interpersonal risk if done poorly, so the redirect has to stay factual and calm rather than accusatory itself. This is a leadership-culture issue that shows up at scale too: it typically takes deliberate, sustained work, roughly a couple of quarters of consistent leadership behavior, published blameless postmortems, and visible non-punitive handling of pages, to shift a team's on-call culture away from a punitive default, and it has to be reinforced the same way every time, including in the exact moment someone in authority breaks the pattern.
Walk me through how you'd run a postmortem after a Sev1 incident: what data you'd gather, how you separate contributing factors from the root cause, and how you turn it into action items that actually get done.
Sample Answer
Direct answer
Run the postmortem in three phases: reconstruct a fact-based timeline from logs, metrics, and deploy history before the meeting even starts; use a structured technique during the meeting, like the "5 whys" or a fishbone breakdown, to separate genuine contributing factors from the actual root cause; and convert findings into specific, owned, time-boxed action items that get tracked to closure rather than filed and forgotten.
Structured elaboration
Data to gather before the meeting
- A precise timeline: when the alert fired, was acknowledged, mitigated, and resolved, with timestamps and who acted at each step.
- Dashboards showing the SLO breach before, during, and after.
- Logs and traces from the affected window, correlated across every service involved, not just the one that alerted first, since the alerting service is often a symptom rather than the cause.
- Deploy and config change history around the incident window, including commit references, feature flags, and infra changes.
- The runbook steps actually executed and their outcomes.
Separating contributing factors from root cause
| Technique | What it's for |
|---|---|
| Timeline correlation | Narrows the causal window before anyone starts theorizing |
| 5 whys | Drills from the symptom to an actionable, systemic cause; stop at the point where fixing it would have prevented the outage |
| Fishbone (Ishikawa) | Categorizes contributors, people, process, code, infra, monitoring, so several independent factors aren't collapsed into one "root cause" |
| Forensic log correlation | For a non-obvious failure spanning multiple components, joins logs and traces by request ID or timestamp across services instead of trusting any single service's self-report |
A factor is a contributing factor if removing it alone would not have prevented the outage. It's the root cause only if removing it would have. Multiple genuine root causes are possible and should be named as such rather than forced into one.
Turning findings into action items that get done
Each item is specific, owned by exactly one person, carries a due date and an explicit acceptance criterion, such as "add an alert that fires when X, verified by injecting a synthetic breach and confirming it pages." Items are prioritized by risk reduced versus effort and tagged by urgency the same way incidents are, so remediation work competes visibly against other roadmap work. They're tracked in the same system as regular engineering work, with a scheduled follow-up check that closure actually happened and actually worked.
Worked example
- 14:02 UTC: a config deploy ships a change to the connection-pool size for checkout-service.
- 14:07 UTC: a checkout-service latency alert fires.
- 14:09 UTC: on-call acknowledges and begins triage on checkout-service logs, the service that alerted.
- 14:18 UTC: correlating logs across checkout-service and its downstream payments-service shows connection-pool exhaustion actually originated in payments-service; the earlier deploy had reduced its pool size, and checkout-service was just the first caller to see timeouts.
- 14:22 UTC: on-call rolls back the payments-service config; latency recovers.
- Root cause: the connection-pool config change on payments-service, since removing it would have prevented the outage. Contributing factor: no alert existed on payments-service's own pool saturation, only on downstream symptoms, which meant triage started on the wrong service before cross-service log correlation found the real source.
- Action items: add a connection-pool saturation alert directly on payments-service, owned by the payments team, verified by injecting a synthetic pool-exhaustion test and confirming it pages before the downstream symptom would; and require connection-pool config changes to go through the same review gate as code changes, owned by the platform team.
Trade-offs and pitfalls
- Stopping at the first plausible "why," usually the service that alerted, instead of correlating across the actual dependency chain is how a downstream root cause gets misattributed to whichever service merely surfaced the symptom first.
- Writing action items as vague intentions, like "improve monitoring," instead of specific, verifiable changes with an owner and an acceptance test, is the single biggest reason action items don't get done.
- A longer, more thorough timeline reconstruction produces a more accurate root cause but delays the meeting; worth it for a top-severity incident, disproportionate for a lower-severity one.
An automated remediation keeps firing and the service flips between healthy and unhealthy as a result, a feedback loop. How would you design the automation to avoid this kind of flapping?
Sample Answer
Direct answer
Stop the flap by putting a circuit breaker on top of the remediation itself: require a minimum number of consecutive failures before acting at all, cap how many automated attempts happen before backing off exponentially, and open the breaker entirely, stopping automated action, after repeated failures so the automation isn't the thing keeping the service oscillating. The automation should get less aggressive the more it fails, not equally aggressive every time.
Structured elaboration
State machine
stateDiagram-v2
[*] --> Closed
Closed --> Remediating: threshold breached
Remediating --> Closed: health recovers
Remediating --> Open: still failing after attempt
Open --> HalfOpen: cooldown elapses
HalfOpen --> Closed: canary healthy
HalfOpen --> Open: canary fails
Open --> HumanApproval: max attempts reached
Why each piece prevents flapping
- A failure-count threshold before acting at all, requiring multiple consecutive failed health checks within a window rather than one, filters out single transient blips that don't need remediation.
- An exponential cooldown between attempts means each retry backs off further, so a persistent problem gets fewer, more spaced-out remediation attempts instead of a tight restart loop.
- The circuit breaker's open state, where the system stops acting entirely and escalates to a human after enough failed attempts, is what actually breaks the flap, not just slows it down.
- The half-open probe tests recovery with a single canary instance or request after the cooldown, before trusting the system enough to resume full automated remediation.
- Staggered, single-instance restarts that respect the dependency graph prevent one service's remediation from cascading into a restart storm on services that depend on it.
Worked example
With a base cooldown of 60 seconds and a cap of 1800 seconds (30 minutes), the cooldown after the nth failed attempt is:
cooldownn=min(base⋅2n−1,max)
cooldown1=min(60⋅20,1800)=60 s
cooldown2=min(60⋅21,1800)=120 s
cooldown3=min(60⋅22,1800)=240 s
cooldown4=min(60⋅23,1800)=480 s
cooldown5=min(60⋅24,1800)=960 s
cooldown6=min(60⋅25,1800)=min(1920,1800)=1800 s
After the sixth failed attempt the breaker opens: no further automated action until a human intervenes. This pattern applies whether the underlying trigger is a transient availability-zone network partition, a database connection-pool spike, or intermittent upstream timeouts. The automation doesn't need to know which cause it's looking at, because the backoff and breaker behave the same regardless.
Trade-offs and pitfalls
- Tuning the failure-count threshold too low means genuinely fast-recurring problems get no remediation attempt before the breaker opens; too high delays help for the transient blips the system was meant to catch.
- A longer maximum cooldown reduces flap risk but leaves a genuinely fixable problem unremediated for longer; the cap should be informed by how long a typical transient cause, like an AZ partition or a connection spike, usually takes to self-resolve.
- Forgetting the dependency graph is the most common miss: fixing one service's flap by restarting it can trigger a second flap downstream if that service health-checks against the still-recovering upstream.
You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?
Sample Answer
Whether to fully automate, keep human-in-the-loop, or leave manual comes down to four questions applied to each specific action: how often does it happen, how bad is it if it goes wrong, can it be safely retried, and can success be verified automatically. High frequency, low blast radius, idempotent, and observable pushes toward full automation; anything destructive or hard to verify stays manual or gated behind a human, no matter how routine it feels.
The criteria
| Criterion | Favors automation | Favors manual / human-in-the-loop |
|---|---|---|
| Frequency | Happens often enough that manual toil adds up | Rare enough that automation investment doesn't pay back |
| Blast radius | Failure is contained (one instance, easily reverted) | Failure can be irreversible or affect data integrity broadly |
| Idempotency | Running it twice is harmless | Running it twice causes a different, possibly worse outcome |
| Verifiability | Success can be checked automatically (health check, row count) | Success requires human judgment to confirm |
Applying it to the three actions
Restarting a cached worker instance: high frequency, low blast radius (stateless, replaceable), fully idempotent, and easily verified with a health check. This is a strong automate candidate: drain connections, spin up a replacement, run a health check, cut traffic over, roll back automatically if the health check fails.
Reattaching a detached volume: lower frequency, meaningfully higher blast radius (attaching to the wrong instance or double-attaching can corrupt data), and only moderately idempotent, reattaching twice isn't necessarily safe. This sits in the middle: automate the pre-checks and the mechanical steps (verify volume ID, verify target instance, snapshot before attaching), but require a human to confirm before the final attach executes.
Running a database schema migration: low frequency, high blast radius (can be destructive and hard to reverse), low idempotency for anything involving DDL (data definition language: schema-altering SQL statements like ALTER TABLE), and success often isn't verifiable by a simple automated check, it needs someone to look at whether the data actually came out right. This stays manual, or more precisely, human-gated: automation handles the mechanical parts (schema diff, pre-migration validation, backup, staged rollout to a canary), but a person approves the production apply.
The underlying argument for phasing automation in gradually
Automating a step doesn't just remove toil, it also removes the moment a human would have caught something unusual about this particular instance of the problem. That's fine for the worker-restart case, where "unusual" mostly doesn't exist, but risky for the migration case, where every migration is a little different. The practical path is phasing: run a new automation in shadow mode first (it proposes the action but a human executes), then human-in-the-loop (it executes after one-click approval), and only promote to full automation once it has a track record across enough real incidents that its false-positive and false-negative rate are actually known, not assumed.
Guarding automated actions with least privilege
Whatever is automated should run with only the permissions that specific action needs, a worker-restart automation shouldn't hold credentials that could also run a schema migration, and every automated action should be logged with who (or what) triggered it and why. For the human-in-the-loop tier, the approval step itself should require a specific person's action (not a shared bot token anyone can trigger), so there's a real approval trail, not a rubber stamp.
Trade-offs and pitfalls
The common wrong turn is automating based on how annoying a task feels rather than how safe it is, restarting workers manually is annoying but safe to automate; migrations are also annoying, but the annoyance is not the variable that should decide it. The other pitfall is leaving a human-in-the-loop step gated behind an approval that nobody actually reads before clicking, if the approval doesn't include enough context (what will run, what's the blast radius, what's the rollback) to make a real judgment, it's automation with an extra click, not a genuine safety gate.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.