On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
How do you hand off an on-call shift so nothing falls through the cracks? What does a good handoff actually need to include?
Sample Answer
Direct answer
A good handoff transfers three things: current state (what's broken or at risk right now), context (what's already been tried and what's scheduled), and ownership (who is accountable for what next), and it has to be verifiable rather than a status dump, meaning the incoming engineer confirms they can actually act (they have access, they can reproduce the symptom, they understand the next step) before the outgoing engineer is done.
Structured elaboration
Handoff checklist, in order:
- Status snapshot: current owner and contact, "all green" or a one-line count of active incidents/unstable services.
- Active incidents: for each, severity, start time, impact, current owner, and a link to the ticket, not a re-explanation from scratch.
- Recent alert history: what's fired in the last several hours, and specifically any alert that's been flapping, since that's exactly what an incoming responder will misdiagnose as new if nobody flags it.
- Ongoing mitigations and runbook links: what's been tried, what's blocked, and the specific next action, with a link to the runbook rather than a paraphrase of it.
- Scheduled changes: upcoming deploys, migrations, or maintenance windows during the next shift, with rollback plans linked.
- Degraded-but-not-incident services: anything running hot or close to a threshold that isn't paging yet.
- Access and tooling: pager rotation, incident channel, dashboards, runbook repo, and an explicit note if the incoming engineer is missing access to any of them.
- Explicit confirmation, not implied: "Can you access the dashboards I linked? Can you reproduce the symptom in incident #123? Do you agree you own the DB migration follow-up?" Each gets an actual yes, not a thumbs-up emoji on a wall of text.
Making it lightweight (chatops mechanics). For day-to-day handoffs without an active incident, a short structured chat message beats a long document nobody reads: header (shift window, owner, escalation contact), one-line status, action items with owners, upcoming risks, and links. Anything in that message that turns out to matter beyond the shift boundary (a workaround that becomes permanent, a gotcha that will recur) gets tagged for follow-up and folded into the actual runbook within a day or two, so the team's durable documentation doesn't quietly live and die in chat history.
Automating the tedious part. The status snapshot, active-incident list, and recent-alert-history sections don't need to be typed by hand: a handoff template can be pre-populated from the monitoring and paging systems (current alert state, open incident IDs, last-deploy timestamp) so the outgoing engineer is editing and confirming pre-filled facts rather than writing a report from a blank page. That reduces both the time cost of handoff and the chance that something gets left out because the outgoing engineer forgot it existed.
Worked example
Friday, 6pm, end of a shift. One active Sev2 incident (checkout latency degraded for a subset of EU traffic, mitigation in progress: a feature flag was flipped to route around a slow dependency, error rate has dropped but root cause isn't fixed), and a database migration scheduled for 2am that night. The outgoing engineer posts:
Handoff | Fri 18:00-Sat 02:00 UTC
Owner: @outgoing -> @incoming | Escalation: @oncall-lead
Status: Degraded (checkout latency, EU) -> incident #482, mitigated not resolved
Action items:
- Watch checkout error rate; if it climbs above 2% again, re-check the feature flag is still on
- DB migration at 02:00 UTC (runbook: <link>, rollback: <link>) - I'll be asleep, this is yours
Risks: migration touches the same table implicated in incident #482; if latency spikes right after, check the migration first
Links: <dashboard> <incident #482> <migration runbook>
The incoming engineer confirms: dashboard access works, they can see incident #482's current state, and they explicitly acknowledge owning the migration watch. That confirmation, not the message itself, is what makes the handoff complete.
Trade-offs and pitfalls
- Pitfall: a "read the ticket" handoff with no verification step lets the incoming responder discover gaps at 3am instead of at 6pm when the person with context is still reachable.
- Pitfall: too much ceremony (a mandatory 45-minute call every single handoff) burns out the outgoing engineer and makes people avoid going on-call at all; reserve synchronous overlap for when there's an active Sev1/Sev2, not as the default for a quiet shift.
- Trade-off: synchronous handoff transfers tacit knowledge best but costs both people's time; async structured notes are cheaper but only as good as their template and discipline. The right default is async-by-default, with synchronous overlap triggered automatically whenever an incident is still open at shift boundary.
During a P1 outage, the first responder doesn't restore service within a few minutes and doesn't acknowledge the page. Walk through what happens next: escalation timeouts, who gets paged, which channels you use, and who ultimately declares a major incident.
Sample Answer
Direct answer
Escalation has to be a timed staircase, not a hope that someone notices: each tier gets an explicit timeout, and crossing it triggers the next tier automatically rather than waiting for a human to decide to escalate. The critical branch is distinguishing "primary hasn't acknowledged at all" (a much stronger, faster signal that they're unreachable) from "primary acknowledged but hasn't fixed it yet" (a slower, more judgment-based signal), and a specific person needs to be unambiguously the incident commander once the incident crosses into major-incident territory.
Structured elaboration
flowchart TD
A[Page fires to primary] --> B{Acked within 5 min?}
B -->|No| C[Auto-escalate to secondary]
B -->|Yes| D{Mitigated or owned by 15 min?}
C --> D
D -->|No| E[Notify SRE lead and EM at 20 min]
E -->|Sev1 criteria still met at 30 min| F[IC declares Major Incident]
F --> G[Comms lead posts updates every 15 min]
| Time | Trigger | Action | Channel |
|---|---|---|---|
| T+0 | Page fires | Primary paged | PagerDuty + phone |
| T+5m | No ack from primary | Auto-escalate to secondary (unreachable-primary path) | PagerDuty + phone |
| T+15m | Acked but not mitigated or no clear owner | Secondary/backup engages as acting responder | Incident Slack channel |
| T+20m | Still unresolved | SRE lead and engineering manager notified | PagerDuty + phone |
| T+30m | Sev1 criteria still met, no imminent fix | Incident Commander formally declares a Major Incident, opens a bridge | Incident channel + status page |
Redundancy for an unreachable primary. The T+5m no-ack escalation exists specifically because "didn't acknowledge" is not the same failure mode as "acknowledged but stuck." An unreachable primary should trigger the fastest possible escalation, on multiple channels at once (page, SMS, phone call) rather than a single retry, since every minute the whole rotation is effectively uncovered.
Who declares Major Incident, and when. The role of Incident Commander should be predefined (whoever is the senior-most engaged responder at the moment of declaration, or a designated on-call IC rotation), not assigned ad hoc during the incident. Declaration criteria should be objective: customer-facing impact confirmed and no imminent fix, not a vibes call.
Audit trail. Every escalation event (page sent, ack received, tier crossed, IC declared) should be logged automatically by the paging tool with timestamps, not reconstructed from memory afterward. This is what makes the eventual postmortem's timeline accurate instead of approximate.
Worked example
A P1 outage affects payment processing. T+0: primary is paged, acknowledges within 2 minutes, and starts investigating (restart attempts, checking recent deploys). T+15m: the issue isn't mitigated and the primary flags they need help; the secondary engages and effectively becomes acting IC. T+20m: still unresolved, SRE lead and EM are notified via page and phone. T+30m: payment processing is still materially impaired with no fix in sight, so the SRE lead formally declares a Major Incident: opens an incident bridge, assigns a comms lead to post external status-page updates and an internal lead to keep driving the technical fix, and sets a 15-minute update cadence. Every one of these steps (ack time, escalation trigger, IC declaration) is logged automatically with a timestamp by the paging tool, which becomes the backbone of the postmortem timeline.
Trade-offs and pitfalls
- Pitfall: "someone will notice and escalate" with no explicit timer is the default at many small companies and reliably fails exactly when it matters most, at 3am with a skeleton crew.
- Pitfall: leaving IC as "whoever showed up" instead of a defined selection rule creates coordination ambiguity in the first minutes of a major incident, which is the worst possible time to negotiate who's in charge.
- Trade-off: the T+5m no-ack timeout should be noticeably tighter than the T+20m and T+30m escalation-for-non-progress timeouts, because "didn't acknowledge at all" is a much stronger signal of a coverage gap than "acknowledged and is still working on it." Making every tier the same length either escalates too aggressively on ordinary, in-progress incidents or too slowly on a genuinely unreachable primary.
Design the guardrails for a system that lets on-call engineers trigger automated runbook actions directly from an alert. How do you prevent a misfire, or a compromised trigger, from causing a bigger outage than the one it was meant to fix?
Sample Answer
Direct answer
Guardrails come from three layers working together: classify every automatable action by risk, reversibility and blast radius, so low-risk actions can run unattended while destructive ones require signed, multi-party approval; rate-limit and circuit-break execution so a misfiring trigger can't repeat itself into a bigger outage; and make every execution auditable and traceable to a specific signed, version-pinned runbook so a compromised trigger's blast radius stays bounded even if it does fire.
Structured elaboration
Threat model
- Misfire: a legitimate alert misclassifies severity or triggers the wrong action, such as restarting the wrong service.
- Compromised automation pipeline: an attacker forges or replays an alert to trigger a destructive action.
- Insider abuse: someone with legitimate access triggers an action outside its intended use.
Each needs a different control. Misfire needs validation and rate limits. Compromise needs signing and short-lived credentials. Insider abuse needs approval gates and an audit log that can't be edited after the fact.
Guardrail decision flow
flowchart TD
A[Alert triggers automated action] --> B{Reversible and low blast radius?}
B -->|Yes, low risk| C{Under rate-limit threshold?}
C -->|Yes| D[Execute in sandboxed, least-privilege runner]
C -->|No| E[Circuit-break: block further auto-actions]
B -->|No, high risk or destructive| F[Require signed approval from two operators]
F --> G{Approved within TTL?}
G -->|Yes| D
G -->|No| H[Escalate to human on-call, no auto-execution]
D --> I[Write signed, tamper-evident audit entry]
Core controls
- Risk classification: every automatable action is tagged low, medium, or high at authoring time, reviewed like code, based on reversibility and blast radius.
- Rate limiting and circuit breaking: cap executions per action per time window, and auto-disable an action after repeated failures instead of letting it keep firing.
- Signing and provenance: only signed, version-pinned runbook releases execute in production; the executor verifies the signature first, so a compromised trigger can request an action but can't smuggle in unreviewed logic.
- Least privilege and short-lived credentials: the executor pulls scoped, time-limited credentials per execution rather than holding standing broad access.
- Immutable audit trail: every execution, who or what triggered it, the parameters, and the outcome, is written to an append-only log kept separate from the systems it can act on, so it survives a compromise of the executor itself.
Worked example
| Action | Reversibility | Blast radius | Risk tier | Guardrail |
|---|---|---|---|---|
| Restart a single stateless pod | Fully reversible | One instance | Low | Auto-execute, rate-limited to one per 5 minutes per pod |
| Drop and rebuild a search index | Not reversible without a full rebuild | Whole service | High | Requires two signed operator approvals within a 10-minute TTL |
Consider an attacker who can forge an alert payload. Against the pod-restart action, they're bounded by the rate limit and the action's inherently small blast radius. Against the index-drop action, they're blocked at the approval gate no matter how convincing the forged alert looks, because approval requires a human signature that an alert payload can't fake.
Trade-offs and pitfalls
- Classifying everything as needing approval defeats the point of automation, which is unattended response to the common case. Keep destructive steps as a separate, explicitly gated action rather than bundling them with routine remediation, so most actions land in the low-risk tier by design.
- More approval gates reduce blast radius but increase mean time to remediate; mitigate by keeping the low-risk tier wide and reserving gates for the genuinely destructive minority.
- Signing the runbook but not validating its parameters leaves a gap: a signed, but freely parameterizable, action like "restart <service>" can still be misused if the parameter itself isn't checked against an allowlist.
What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?
Sample Answer
MTTD is how long a problem existed before anything noticed it. MTTA is how long a human took to acknowledge the alert once it fired. MTTR is how long it took to fully resolve once someone was working it. The most common mistake is reporting a single blended average and treating it as typical, when one long outage in the set is doing all the work.
Definitions
| Metric | Starts at | Ends at | What it measures |
|---|---|---|---|
| MTTD | Failure begins | Alert fires / someone notices | How good detection is |
| MTTA | Alert fires | Human acknowledges | How well paging and routing work |
| MTTR | Acknowledgment | Service fully restored | How fast the response process fixes it, once someone owns it |
(Some teams instead measure MTTR from detection to resolve rather than ack to resolve; either is defensible, but the convention has to be fixed and stated, because mixing them across teams silently changes what the number means.)
Worked example: one incident timeline
| Event | Time |
|---|---|
| Failure begins | 14:00:00 |
| Alert fires (detection) | 14:06:00 |
| Engineer acknowledges | 14:11:00 |
| Service restored | 14:47:00 |
Worked example: averaging across three incidents, and where it goes wrong
| Incident | MTTD | MTTA | MTTR |
|---|---|---|---|
| 1 | 6 | 5 | 36 |
| 2 | 2 | 3 | 20 |
| 3 | 15 | 8 | 54 |
The mean of 36.7 minutes is being pulled up almost entirely by incident 3's 54-minute outlier: the mean sits above two of the three data points (20 and 36), with only the outlier itself larger. The median of {20, 36, 54} is 36, the middle value itself rather than a value inflated by the outlier, so it is a better single-number stand-in for the typical incident than the mean. Reporting mean MTTR alone, without the incident count or a percentile, makes a single bad incident look like the typical case.
Trade-offs and pitfalls
Comparing MTTR across teams that use different start-point conventions is comparing two different metrics wearing the same name; agree on the convention org-wide before benchmarking teams against each other. A dropping mean MTTR can hide a rising incident count: if you're resolving more small incidents faster while one rare severe incident still takes hours, the mean improves and the tail risk hasn't moved at all. Improving MTTD without improving MTTA or MTTR just means you find out about the same slow response faster; treat the three as stages of one pipeline, not independent wins to report separately.
Design an on-call escalation system for an organization with multiple teams that need to coordinate coverage across time zones. How do you route pages, prevent alert-noise from cascading into unnecessary escalations, and decide who gets pulled in for a revenue-impacting versus a data-sensitive incident?
Sample Answer
An escalation system for a multi-team, multi-timezone org needs three separate mechanisms working together: routing (getting an alert to the right on-duty person without a human deciding that in the moment), noise suppression (so one root cause doesn't fan out into ten pages), and a severity model that decides who gets pulled in and how fast, because a revenue-impacting outage and a data-sensitive incident need different people in the room, not just different urgency.
Core building blocks
- Alert gateway: every alert is deduplicated by fingerprint, enriched with service/team/severity tags, and checked against maintenance windows before it's allowed to page anyone.
- Routing table: maps service ownership and team schedule (including timezone-local business-hours windows) to whoever is currently on duty, kept in the paging tool as the single source of truth rather than a wiki page someone forgets to update.
- Severity model: decides who gets paged first and how many people, based on impact type, not just raw error rate.
- Escalation ladder: a fixed sequence of who gets paged next if no one acknowledges, with a hard time budget at each step.
Severity and routing matrix
| Impact type | First page | Ack SLA | If unacked | Extra routing |
|---|---|---|---|---|
| Revenue-impacting (checkout, payments down) | Primary on-call for the affected service | 5 min | Escalate to secondary, then service lead | Auto-opens a major-incident bridge if still unacked at 15 min |
| Data-sensitive (PII exposure, access-control gap) | Primary on-call and security/compliance on-call, paged together | 5 min | Escalate both chains in parallel | Legal/compliance notified regardless of ack status, on a fixed clock, not gated on resolution |
| BI/dashboard degradation (stale or broken dashboards, no customer-facing impact) | Data platform on-call only | 30 min | Escalate to data platform lead | No bridge; tracked as a ticket unless it crosses a staleness threshold (e.g. data older than its documented freshness SLA) |
Escalation flow
flowchart TD
A[Alert fires] --> B[Gateway: dedupe, enrich, tag severity]
B --> C{Severity type}
C -->|Revenue-impacting| D[Page Primary, ack SLA 5m]
C -->|Data-sensitive| E["Page Data on-call AND Compliance lead in parallel; notify Legal/Compliance on fixed clock, independent of ack"]
D --> F{Acked by 5m?}
F -->|No| G[Escalate to Secondary, ack SLA +10m]
G --> H{Acked by 15m?}
H -->|No| I[Escalate to Service Lead, open incident bridge]
F -->|Yes| J[Primary mitigates]
H -->|Yes| J
E --> K{Acked by 5m?}
K -->|No| L[Escalate both chains in parallel, ack SLA +10m]
K -->|Yes| M[Data on-call + Compliance mitigate]
L --> M
Preventing alert-noise from cascading into unnecessary escalations
Most alert storms come from one root cause tripping many downstream checks at once (a database going down pages every service that depends on it). The gateway groups alerts by a correlation key (same root dependency, same time window) before routing, so the escalation ladder above runs once for the incident, not once per symptom. Escalation timers also only start on the first page for a correlated group; late-arriving duplicates reset nothing.
Keeping the matrix trustworthy
A routing matrix that's wrong is worse than no matrix, because it creates false confidence. Two things keep it honest: primary/backup contacts are pulled live from the scheduling tool rather than hand-maintained, and if the on-duty person marks themselves absent (leave, travel) in that same tool, pages route to the next person automatically instead of timing out silently first. A silent timeout during a real on-call absence is exactly the failure mode that erodes trust in the whole system.
Trade-offs and pitfalls
Stricter deduplication reduces noise but risks folding two genuinely unrelated incidents into one correlation group if the correlation key is too broad; the fix is scoping correlation to a real dependency graph, not just a time window. A common wrong turn is building one severity ladder for everything, which either pages security teams for routine downtime or under-escalates a compliance-relevant incident because it didn't look revenue-critical on the dashboard. The severity model has to be impact-type aware, not just impact-size aware.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.