On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?
Sample Answer
Direct answer
Frame the sensitivity-versus-fatigue tension as an explicit cost trade-off rather than a vibes call: assign a rough relative cost to a missed incident versus a false page, pick the threshold that minimizes expected cost given the current false-positive and false-negative rates, and never roll a threshold change straight to paging. Validate it in shadow mode against real traffic first, then canary it on a subset before a full rollout, with an automatic rollback trigger if things get worse.
Structured elaboration
Cost framing. At a candidate threshold τ, define expected cost as:
cost(τ)=CFN⋅P(miss∣τ)+CFP⋅P(false page∣τ)Lowering τ (more sensitive) drives the probability of a miss toward zero but raises the false-page rate, and raising τ does the opposite. The right threshold is wherever this sum is smallest, not wherever either rate alone looks best in isolation.
Safe testing method, three stages with explicit gates:
- Shadow mode: the candidate threshold runs log-only, never pages, and every alert it would have fired gets compared against the real incident record after the fact. No production risk, but also no real-time responder feedback.
- Canary: the candidate threshold pages for real, but only for a subset of services or regions. Watch for missed-detection signals and responder load on that subset before touching anything else.
- Gradual rollout: feature-flagged expansion to the rest of the fleet, with an automatic rollback trigger if the false-positive rate or acknowledgment latency regresses past a predefined bound.
Worked example
Say a one-week shadow test compares the current threshold (A) against a stricter candidate (B) on the same underlying traffic, and every alert is later labeled against the real incident record:
| Threshold | Total alerts | True incidents caught (TP) | False pages (FP) | Missed incidents (FN) |
|---|---|---|---|---|
| A (current) | 200 | 18 | 182 | 2 |
| B (candidate) | 40 | 16 | 24 | 4 |
Assign a rough relative cost: a missed incident costs 500 responder-hour-equivalents (CFN=500), a false page costs 1 (CFP=1). Expected cost per threshold:
costA=CFN⋅FNA+CFP⋅FPA=500×2+1×182=1182 costB=CFN⋅FNB+CFP⋅FPB=500×4+1×24=2024Even though B has far higher precision (16 of 40 alerts were real, versus 18 of 200 for A), A has the lower expected cost: B's two extra missed incidents cost more than the 158 extra false pages A generates. If CFN were much closer to CFP, for example a low-stakes internal tool where a miss is only mildly annoying rather than expensive, B would win instead. The point of doing the arithmetic rather than eyeballing the false-positive rate is that the right threshold depends on getting the relative cost of a miss right for that specific service, not on chasing a universally "less noisy" target.
Trade-offs and pitfalls
- Pitfall: picking CFN and CFP once and never revisiting them. Relative costs shift as customer scale, contractual SLAs, and the service's blast radius change, so the cost model needs the same periodic review as the severity scale it feeds into.
- Pitfall: shadow-testing only against the same traffic period used to design the threshold in the first place, which overfits the result; test against a held-out period the threshold wasn't tuned on.
- Trade-off: assigning explicit costs forces an uncomfortable conversation with stakeholders about what a missed incident is actually worth, but that discomfort produces a threshold the team actually agreed to, rather than one that came from a single engineer's gut feel nobody was consulted on.
Design a severity rubric, say P0 through P3, for a SaaS product. What determines the level, what SLA applies at each, and who has to be paged?
Sample Answer
Direct answer
Base the rubric on customer-facing impact and business consequence, not internal technical severity: P0 is a full outage or a safety or data-loss risk with immediate paging and an aggressive SLA, scaling down through partial-impact and degraded-but-functional states to P3, a backlog item with no paging at all. Each level pairs a concrete definition with a response SLA, a resolution SLA, and exactly who gets paged, so classification is a lookup an on-call engineer can make under pressure, not a judgment call.
Structured elaboration
| Severity | Impact | Response SLA | Resolution SLA | Who's paged |
|---|---|---|---|---|
| P0 | Full outage, data loss or corruption, or safety or legal risk, affecting all or most customers | Page immediately, acknowledge within 5 minutes | Continuous work until mitigated, target under 4 hours | On-call engineer, service owner, incident commander, security or legal if data exposure |
| P1 | Major feature broken for many customers, or one high-value customer severely impacted; SLA-covered functionality degraded | Page, acknowledge within 15 minutes | Target under 24 hours, mitigation expected sooner | On-call engineer, team tech lead, customer success for affected accounts |
| P2 | Partial degradation, a subset of users, intermittent errors, core flows still work | Notify without paging, acknowledge within 1 hour during business hours | Target under 3 business days | Feature owner; on-call optional |
| P3 | Cosmetic issue, edge case, no measurable customer impact | Acknowledge within 1 business day | Scheduled into normal backlog | No paging; filed and triaged by product and engineering |
What determines the level
- Scope: how many customers or what fraction of traffic is affected.
- Reversibility of harm: data being lost or corrupted pushes toward P0 regardless of how few customers are affected, versus a fully recoverable error.
- Whether a workaround exists for the customer.
- Contractual exposure: whether this breaches an SLA the company is financially on the hook for.
Reclassifying as more information arrives
Declare an initial severity fast, from the first available signal, and treat it as provisional. Many incidents start looking like P1 and get upgraded to P0 once data loss is confirmed, or start as P0, a total outage, and downgrade to P1 once a workaround is found. Reclassification in either direction should be cheap and require no approval, because holding onto an inaccurate severity either under-pages a real emergency or burns unnecessary on-call attention.
Worked example
A payment-processing API returns errors for all merchants for several minutes before an automatic circuit breaker reroutes traffic to a backup provider, after which errors drop back to baseline. Applying the table: the initial signal, all merchants affected and revenue-blocking, classifies this as P0 and pages the on-call engineer, the service owner, and the incident commander immediately. Once the reroute confirms the impact is contained and no data was lost, the incident is reclassified to P1, SLA-covered functionality degraded with a workaround in place via the backup provider, for the remainder of the response. That reclassification changes the resolution SLA from continuous work under 4 hours to a target under 24 hours, but it doesn't stand down the already-paged responders mid-incident.
Trade-offs and pitfalls
- Defining severity by an internal technical signal, like an error-rate threshold, instead of customer impact, treats a high error rate on a low-traffic internal endpoint the same as the same error rate on checkout, which it isn't.
- Too many severity levels creates ambiguity at classification time under pressure; four tiers is usually enough resolution to route paging and SLAs correctly without forcing a judgment call between two levels that don't functionally differ.
- Making reclassification require a meeting or approval means on-call will just leave the severity wrong, which quietly corrupts incident metrics later.
Mid-incident during a Sev1, you discover the runbook you're following has outdated commands that don't work on the current cluster configuration. What do you do to keep the response moving, and how do you make sure the runbook gets fixed afterward?
Sample Answer
Direct answer
Keep the incident moving without trusting the stale command: switch to safe, read-only discovery to re-derive the actual current state instead of assuming the runbook's exact syntax still matches reality, get a second responder to sanity-check any ad-hoc workaround before running it, and narrate every command and its outcome in the incident channel as you go. Afterward, treat the correction as a first-class, owned follow-up: file it with the incident as evidence and require the same review the runbook normally gets, rather than merging a fix drafted under adrenaline with no second pair of eyes.
Structured elaboration
Keeping the response moving
- Don't keep retrying the stale command hoping it starts working; that burns clock on a Sev1.
- Fall back to read-only discovery to find what actually changed: list the current resource names or config instead of assuming the runbook's exact prior values, and check whether a known change (a migration, a rename, a tool upgrade) explains the mismatch.
- If a workaround command is genuinely needed, treat it like an experiment: scope it to the smallest possible blast radius (a single pod, host, or canary) if at all possible, and have a second responder review it before running anything destructive.
- Narrate in the incident channel as you go: exact command, who ran it, what happened. This live log is what makes the eventual runbook fix accurate instead of a reconstruction from memory the next morning.
Getting the runbook actually fixed afterward
- File the correction as its own owned action item, not "someone should update this."
- Attach evidence: the incident timeline entries showing what actually worked, captured live rather than recalled later.
- Route the fix through the same review the runbook would normally require. A fix drafted under incident pressure is exactly the kind of change that benefits from a second reviewer, not an exception to needing one.
- If the drift has a systemic cause, such as no process tying documentation updates to the change that invalidated it, raise that separately as its own postmortem action item, not just a one-off doc patch.
Worked example
kubectl rollout restart deployment/checkout -n prod fails with "deployment not found." The responder falls back to read-only discovery: kubectl get deploy -A | grep checkout shows the deployment now lives in namespace checkout-prod, after a namespace-per-service migration weeks earlier that never touched the runbook. The corrected command runs, the service recovers, and both the failed and working commands are logged with timestamps in the incident doc as they happen. The resulting follow-up action item reads: "update runbook RB-042's namespace reference and add a namespace-lookup step instead of a hardcoded name; owner: platform team; verified by a peer review plus a sandboxed dry-run before merge."
Trade-offs and pitfalls
- Fixing the runbook file directly, mid-incident, with no review is a common shortcut; the correct fix ships as a follow-up change through the normal review gate, informed by what was learned live.
- Treating the ad-hoc working command as tribal knowledge instead of writing it down immediately is how the same staleness reappears at the next incident; capture it in the channel the moment it works, not after the retro.
- Verifying a corrected command in a sandbox before trusting it live is safer, but a Sev1 usually doesn't have that time; this is really an argument for building runbooks with idempotent, safe-to-retry commands and dynamic lookups in the first place, so on-call isn't forced into that trade-off during the incident.
How would you actually validate that your runbooks work before you need them in a real incident? Describe a program for testing them under realistic conditions.
Sample Answer
Direct answer
Validating runbooks before a real incident needs three ingredients: a way to safely exercise them, using sandboxes for non-destructive steps and canary execution for destructive ones; measurable acceptance criteria for what "verified" actually means; and a cadence that scales from cheap, frequent tabletop walkthroughs up to expensive, rare full game days. Treat validation as a graded ladder of increasing realism and cost, not a single all-or-nothing chaos exercise.
Structured elaboration
Ladder of validation
| Tier | What happens | Frequency | Blast radius |
|---|---|---|---|
| Tabletop read-through | Team reads the runbook aloud, checks it still matches current architecture | Monthly per critical runbook | None, discussion only |
| Sandboxed dry-run | Non-destructive or dry-run steps run against a synthetic or staging copy | Per runbook change, CI-gated | Isolated sandbox |
| Canary execution | The real, potentially destructive step runs against a single instance or shard in production | Quarterly for the highest-severity runbooks | One instance or shard |
| Full game day | The real trigger condition is simulated and the whole runbook runs end to end with the actual on-call rotation | Quarterly cross-team, and after major architecture changes | Scoped production traffic, with a kill switch |
Acceptance criteria for "verified"
- Every command in the runbook executed successfully against the environment matching its tier.
- Recovery met the documented RTO (recovery time objective: the maximum acceptable downtime) for that exercise.
- The person executing it was not the runbook's original author, which catches "only the author can actually run this" runbooks.
- No manual step was needed beyond what's written, which catches missing steps.
- The runbook's last-verified metadata is updated only after all of the above pass, tied to the specific commit that was tested.
Sandboxing and canary mechanics for destructive steps
- A dry-run flag on any script validates and logs without mutating anything; most cloud SDKs and infrastructure-as-code tools support this natively.
- Ephemeral, synthetic-data environments handle full destructive rehearsals safely, isolated by namespace or project and feature-flagged away from real customer traffic.
- Steps that can only be meaningfully tested in production, like a real failover, get a canary first: a single shard or instance, with an automated rollback path and a pre-agreed abort condition.
Error-budget gate before running in production
Before a production game day, confirm enough error budget remains to absorb the intended, and any accidental, impact. For a monthly SLO of 99.9 percent over a 30-day window, the allowed downtime is:
error budget=(1−SLO)×window minutes
(1−0.999)×43,200=0.001×43,200=43.2 minutes
If the remaining budget is close to that 43.2-minute figure, postpone the exercise rather than spend the safety margin on a drill.
Worked example
A cache-cluster failover runbook is canary-tested against a single shard. The failover command executes without manual intervention, recovery meets the documented RTO target, and the engineer running the drill is not the runbook's original author. All four acceptance criteria above pass, so the runbook's last-verified metadata is updated to the exact commit hash that was tested, and the result feeds into the quarterly decision of whether this runbook is due for a full game day next.
Trade-offs and pitfalls
- Relying only on tabletop reads because full game days are expensive means staleness in the actual commands never gets caught until a real incident does it for you.
- Letting the runbook's author be the only person who can successfully execute it means you've tested the author's tribal knowledge, not the documentation; a different operator running the drill is what actually validates the doc.
- Full production game days build the highest confidence but carry real risk and spend real error budget; the ladder exists so most validation stays cheap, and only the highest-severity runbooks earn a full game day.
Walk me through how you'd run a postmortem after a Sev1 incident: what data you'd gather, how you separate contributing factors from the root cause, and how you turn it into action items that actually get done.
Sample Answer
Direct answer
Run the postmortem in three phases: reconstruct a fact-based timeline from logs, metrics, and deploy history before the meeting even starts; use a structured technique during the meeting, like the "5 whys" or a fishbone breakdown, to separate genuine contributing factors from the actual root cause; and convert findings into specific, owned, time-boxed action items that get tracked to closure rather than filed and forgotten.
Structured elaboration
Data to gather before the meeting
- A precise timeline: when the alert fired, was acknowledged, mitigated, and resolved, with timestamps and who acted at each step.
- Dashboards showing the SLO breach before, during, and after.
- Logs and traces from the affected window, correlated across every service involved, not just the one that alerted first, since the alerting service is often a symptom rather than the cause.
- Deploy and config change history around the incident window, including commit references, feature flags, and infra changes.
- The runbook steps actually executed and their outcomes.
Separating contributing factors from root cause
| Technique | What it's for |
|---|---|
| Timeline correlation | Narrows the causal window before anyone starts theorizing |
| 5 whys | Drills from the symptom to an actionable, systemic cause; stop at the point where fixing it would have prevented the outage |
| Fishbone (Ishikawa) | Categorizes contributors, people, process, code, infra, monitoring, so several independent factors aren't collapsed into one "root cause" |
| Forensic log correlation | For a non-obvious failure spanning multiple components, joins logs and traces by request ID or timestamp across services instead of trusting any single service's self-report |
A factor is a contributing factor if removing it alone would not have prevented the outage. It's the root cause only if removing it would have. Multiple genuine root causes are possible and should be named as such rather than forced into one.
Turning findings into action items that get done
Each item is specific, owned by exactly one person, carries a due date and an explicit acceptance criterion, such as "add an alert that fires when X, verified by injecting a synthetic breach and confirming it pages." Items are prioritized by risk reduced versus effort and tagged by urgency the same way incidents are, so remediation work competes visibly against other roadmap work. They're tracked in the same system as regular engineering work, with a scheduled follow-up check that closure actually happened and actually worked.
Worked example
- 14:02 UTC: a config deploy ships a change to the connection-pool size for checkout-service.
- 14:07 UTC: a checkout-service latency alert fires.
- 14:09 UTC: on-call acknowledges and begins triage on checkout-service logs, the service that alerted.
- 14:18 UTC: correlating logs across checkout-service and its downstream payments-service shows connection-pool exhaustion actually originated in payments-service; the earlier deploy had reduced its pool size, and checkout-service was just the first caller to see timeouts.
- 14:22 UTC: on-call rolls back the payments-service config; latency recovers.
- Root cause: the connection-pool config change on payments-service, since removing it would have prevented the outage. Contributing factor: no alert existed on payments-service's own pool saturation, only on downstream symptoms, which meant triage started on the wrong service before cross-service log correlation found the real source.
- Action items: add a connection-pool saturation alert directly on payments-service, owned by the payments team, verified by injecting a synthetic pool-exhaustion test and confirming it pages before the downstream symptom would; and require connection-pool config changes to go through the same review gate as code changes, owned by the platform team.
Trade-offs and pitfalls
- Stopping at the first plausible "why," usually the service that alerted, instead of correlating across the actual dependency chain is how a downstream root cause gets misattributed to whichever service merely surfaced the symptom first.
- Writing action items as vague intentions, like "improve monitoring," instead of specific, verifiable changes with an owner and an acceptance test, is the single biggest reason action items don't get done.
- A longer, more thorough timeline reconstruction produces a more accurate root cause but delays the meeting; worth it for a top-severity incident, disproportionate for a lower-severity one.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.