On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
How do you hand off an on-call shift so nothing falls through the cracks? What does a good handoff actually need to include?
Sample Answer
Direct answer
A good handoff transfers three things: current state (what's broken or at risk right now), context (what's already been tried and what's scheduled), and ownership (who is accountable for what next), and it has to be verifiable rather than a status dump, meaning the incoming engineer confirms they can actually act (they have access, they can reproduce the symptom, they understand the next step) before the outgoing engineer is done.
Structured elaboration
Handoff checklist, in order:
- Status snapshot: current owner and contact, "all green" or a one-line count of active incidents/unstable services.
- Active incidents: for each, severity, start time, impact, current owner, and a link to the ticket, not a re-explanation from scratch.
- Recent alert history: what's fired in the last several hours, and specifically any alert that's been flapping, since that's exactly what an incoming responder will misdiagnose as new if nobody flags it.
- Ongoing mitigations and runbook links: what's been tried, what's blocked, and the specific next action, with a link to the runbook rather than a paraphrase of it.
- Scheduled changes: upcoming deploys, migrations, or maintenance windows during the next shift, with rollback plans linked.
- Degraded-but-not-incident services: anything running hot or close to a threshold that isn't paging yet.
- Access and tooling: pager rotation, incident channel, dashboards, runbook repo, and an explicit note if the incoming engineer is missing access to any of them.
- Explicit confirmation, not implied: "Can you access the dashboards I linked? Can you reproduce the symptom in incident #123? Do you agree you own the DB migration follow-up?" Each gets an actual yes, not a thumbs-up emoji on a wall of text.
Making it lightweight (chatops mechanics). For day-to-day handoffs without an active incident, a short structured chat message beats a long document nobody reads: header (shift window, owner, escalation contact), one-line status, action items with owners, upcoming risks, and links. Anything in that message that turns out to matter beyond the shift boundary (a workaround that becomes permanent, a gotcha that will recur) gets tagged for follow-up and folded into the actual runbook within a day or two, so the team's durable documentation doesn't quietly live and die in chat history.
Automating the tedious part. The status snapshot, active-incident list, and recent-alert-history sections don't need to be typed by hand: a handoff template can be pre-populated from the monitoring and paging systems (current alert state, open incident IDs, last-deploy timestamp) so the outgoing engineer is editing and confirming pre-filled facts rather than writing a report from a blank page. That reduces both the time cost of handoff and the chance that something gets left out because the outgoing engineer forgot it existed.
Worked example
Friday, 6pm, end of a shift. One active Sev2 incident (checkout latency degraded for a subset of EU traffic, mitigation in progress: a feature flag was flipped to route around a slow dependency, error rate has dropped but root cause isn't fixed), and a database migration scheduled for 2am that night. The outgoing engineer posts:
Handoff | Fri 18:00-Sat 02:00 UTC
Owner: @outgoing -> @incoming | Escalation: @oncall-lead
Status: Degraded (checkout latency, EU) -> incident #482, mitigated not resolved
Action items:
- Watch checkout error rate; if it climbs above 2% again, re-check the feature flag is still on
- DB migration at 02:00 UTC (runbook: <link>, rollback: <link>) - I'll be asleep, this is yours
Risks: migration touches the same table implicated in incident #482; if latency spikes right after, check the migration first
Links: <dashboard> <incident #482> <migration runbook>
The incoming engineer confirms: dashboard access works, they can see incident #482's current state, and they explicitly acknowledge owning the migration watch. That confirmation, not the message itself, is what makes the handoff complete.
Trade-offs and pitfalls
- Pitfall: a "read the ticket" handoff with no verification step lets the incoming responder discover gaps at 3am instead of at 6pm when the person with context is still reachable.
- Pitfall: too much ceremony (a mandatory 45-minute call every single handoff) burns out the outgoing engineer and makes people avoid going on-call at all; reserve synchronous overlap for when there's an active Sev1/Sev2, not as the default for a quiet shift.
- Trade-off: synchronous handoff transfers tacit knowledge best but costs both people's time; async structured notes are cheaper but only as good as their template and discipline. The right default is async-by-default, with synchronous overlap triggered automatically whenever an incident is still open at shift boundary.
Design how an alert should connect automatically to the right runbook, including letting a responder execute a pre-approved remediation step with one click. What do you log, and what stops the system from taking an unsafe action on its own?
Sample Answer
Direct answer
Wire the alert directly to a runbook execution layer, not just a link: the alert carries a runbook identifier, the paging tool surfaces the matching runbook inline with a one-click execute action for pre-approved steps, and every click passes through validation, RBAC (role-based access control, the system that decides which actions a given user is permitted to run), and policy evaluation before anything runs, with the result and a full audit entry logged back to the incident. What actually stops an unsafe action isn't the button in the UI, it's the policy layer sitting between the click and the execution.
Structured elaboration
Component flow
sequenceDiagram
participant M as Monitoring
participant PD as PagerDuty
participant R as Responder
participant E as Runbook Executor
participant V as Audit Log
M->>PD: Alert fires with runbook_id attached
PD->>R: Page plus incident link
R->>E: Open runbook, click approved step
E->>E: Validate params, RBAC, policy check
E->>V: Write audit entry: who, what, when
E->>PD: Post execution result to incident
What each layer does
- Alert enrichment: alerting rules carry a runbook identifier and a list of allowed actions as labels, so the alert itself knows which remediation it maps to.
- Paging tool integration: the incident view embeds the runbook with each step clearly labeled manual versus automated, along with its risk level.
- Execution broker: on click, it validates the caller's identity through RBAC, confirms the incident is actually open and acknowledged, and evaluates policy, such as whether this action needs extra approval given current severity or a maintenance window.
- Audit log: append-only, recording who clicked, what parameters were used, what the action did, and the result, written before the action is considered complete rather than best-effort afterward.
What logs, and what actually stops an unsafe action
| Logged | Why |
|---|---|
| Who triggered it, real identity | Accountability and RBAC review |
| Incident ID and state at trigger time | Prevents replay against a stale or closed incident |
| Exact parameters passed | Reproducibility and forensics |
| Approval chain, if required | Proves the human-in-loop gate was actually satisfied |
| Result and outcome class | Feeds back into whether this action should keep auto-executing |
Nothing about the UI stops an unsafe action; the broker re-validates everything server-side even if the button already looked available. Any action tagged destructive requires approval regardless of who clicks it, a circuit breaker disables an action automatically after repeated failures, and execution always runs with short-lived, scoped credentials so even a successful unsafe action is bounded in what it can touch.
Worked example
A "restart cache node" action is clicked from the incident view. The broker checks: is the incident open, yes; is the clicker on the paging on-call roster for this service, yes; is this action tagged low-risk, yes, a single-node restart; is it within the rate limit, zero executions in the last five minutes against a cap of one. All checks pass, the executor runs the restart with a scoped, time-limited credential, and writes the audit entry, who, incident ID, action, parameters, result, before returning success to the incident view.
Trade-offs and pitfalls
- Client-side checks are fine for UX, greying out an unavailable button, but the broker must re-validate everything server-side; the client should never be the actual enforcement point.
- Real-time policy evaluation on every click adds latency to remediation; keep the policy check itself a fast local lookup rather than a call to an external system on the critical path, so safety doesn't undercut automation's main benefit.
- Logging "success" based on the action returning without error, rather than verifying the target metric actually recovered, is a common gap; a runbook step can exit cleanly and still not have fixed anything.
What does 'blameless' actually mean in a blameless postmortem, and why does it matter? What are the essential components of a good postmortem document?
Sample Answer
Direct answer
"Blameless" means the postmortem analyzes the incident as a failure of systems and processes that made a reasonable person's normal action produce a bad outcome, not as a failure of that person's competence or effort. It matters because the moment a postmortem starts assigning individual fault, people stop giving you the honest details (what they clicked, what they assumed, what they skipped) that you actually need to fix the underlying gap, and near-misses stop getting reported at all.
Structured elaboration
What blameless does not mean. It is not "no accountability." Individuals still own action items and are still expected to do their jobs well. Blameless means the analysis stops at "why did this look like the right thing to do at the time, given the information and tools available" instead of stopping at "who made the mistake."
Essential components of a postmortem document:
- Summary and impact: severity, duration, which customers/systems were affected, business impact.
- Timeline: timestamped sequence from first signal to full resolution, including detection and every mitigation attempt.
- Contributing factors: plural, not a single root cause. A "five whys" style chain that includes technical gaps (missing test, no lock-duration check) and process gaps (review checklist didn't require it, no staging environment with prod-sized data).
- What went well / what didn't: honest assessment of detection speed, mitigation effectiveness, and communication, separate from the technical cause.
- Action items: each with an owner, a due date, and a verification step, tracked to closure rather than left as a paragraph nobody revisits.
- Lessons learned: shared broadly enough that other teams with the same pattern can act on it before they hit the same incident.
Worked example
A deploy runs a database migration that locks a hot table for several minutes, causing a customer-facing outage. A blameful writeup says: "Engineer X pushed a migration without checking lock duration." A blameless writeup for the same incident says: "The migration tool doesn't warn about lock duration before merge, and the review checklist doesn't require a dry run against a prod-sized dataset. Both are now action items: add a lock-duration check to the migration tool (owner: platform team, verify by testing against a 10M-row table), and add a prod-sized dry-run step to the migration checklist (owner: DBA lead, verify by auditing the next five migrations)." The second version identifies exactly the same failure but produces two concrete, ownable fixes instead of a warning to be more careful next time, which is not something that reliably prevents a repeat.
Trade-offs and pitfalls
- Pitfall: blameless in the document but not in the room. Teams sometimes write a passive-voice, name-free doc while the actual retro meeting is full of pointed questions at one person. The doc's tone has to match the meeting's tone, or the culture stays blameful regardless of what's written down.
- Pitfall: treating blameless as zero consequences ever. Repeated, egregious negligence (ignoring a known policy, skipping a required review on purpose) is a people-management conversation, but it happens outside the technical postmortem, not inside it.
- Senior signal: distinguishing proximate cause from contributing factors. A junior answer stops at "the migration locked the table." A senior answer keeps asking why the tooling, the review process, and the testing environment all failed to catch it, because single-root-cause thinking tends to produce a single, shallow fix that doesn't survive the next incident with a different trigger but the same underlying gap.
How do you define severity levels for production incidents (say Sev1 through Sev4), and how does severity map to expected response time and who gets notified?
Sample Answer
Direct answer
Severity is a fixed classification of an incident's technical and business impact right now (how bad is it), and each severity tier maps to a specific acknowledgment SLA, escalation path, and notification list so the response scales automatically with how bad things are. Severity is often confused with priority: severity measures blast radius and impact, while priority additionally weighs urgency and business context, and the two usually move together but can diverge.
Structured elaboration
| Severity | Definition | Ack SLA | Who's paged | Update cadence |
|---|---|---|---|---|
| Sev1 | Full outage, data loss, or security breach affecting all or most customers | 5 minutes | Primary + secondary on-call, engineering manager, exec on-call | Every 15-30 min until resolved |
| Sev2 | Major feature broken or severe degradation for a large subset of users | 15 minutes | Primary on-call, secondary auto-paged if unacked | Every 30-60 min |
| Sev3 | Partial degradation with a workaround, or impact limited to a small subset | Next business hour | Routed to on-call as a ticket, no page | Daily until closed |
| Sev4 | Cosmetic or non-user-facing issue | Best effort | Backlog, no page | None required |
Severity vs. priority. Severity is a property of the system: what fraction of functionality is broken and for whom. Priority is a property of the response: how urgently the organization needs to act on it right now, which factors in severity plus things like contract SLAs, timing, and who is affected. A Sev2 bug (partial degradation, workaround exists) affecting one enterprise customer with a contractual one-hour response commitment can get treated with P1 urgency even though its technical severity classification stays Sev2. Conversely, a technically Sev1-caliber bug discovered in a staging-only environment has low priority because there is no live customer impact yet. Conflating the two leads to two failure modes: under-resourcing a contractually urgent-but-technically-narrow issue, or paging the whole org for something with real severity but zero current business urgency.
Worked example
Two incidents happen the same week. Incident A: the primary API returns errors for 70% of requests across all customers. That's Sev1 by impact (majority of users, core path) and P1 by urgency (acknowledge in 5 minutes, exec on-call notified). Incident B: a non-critical reporting endpoint used by one enterprise customer returns stale data. By impact alone that's Sev3 (small subset, workaround exists: refresh manually). But that customer's contract has a 30-minute response SLA for any reported defect, so it gets routed with P1 priority: acknowledged within the contract window and staffed immediately, even though the severity label on the incident stays Sev3. The postmortem for B should note this divergence explicitly, since it's exactly the kind of nuance a severity-only view misses.
Trade-offs and pitfalls
- Pitfall: over-classifying everything as Sev1 "to be safe." This burns out on-call and trains people to treat pages as noise, defeating the purpose of having tiers at all.
- Pitfall: assigning severity once at triage and never revisiting it. Initial severity is frequently wrong (scope looks narrow until the second wave of impact shows up); the postmortem should include a severity-accuracy check as a standard field.
- Pitfall: letting priority silently override severity without documenting why, which erodes trust in the severity scale over time because people start reading "severity" as "whatever got the fastest response," rather than a consistent, calibratable measure of impact.
What is the role of an Incident Commander during a live incident, and how does it differ from the other roles typically involved in incident response?
Sample Answer
The Incident Commander (IC) owns the incident, not the fix. Their job is coordination and decision-making, deciding priorities, deciding when to escalate, keeping the response moving, while the technical work of diagnosing and fixing the problem belongs to subject matter experts (SMEs) the IC coordinates but doesn't have to be one of. That separation is the whole point: it lets someone stay focused on the shape of the response instead of getting pulled into a single technical rabbit hole.
Roles and how they differ
| Role | Owns | Does NOT do |
|---|---|---|
| Incident Commander (IC) | Overall response: sets priorities, makes the call/rollback/escalate decisions, declares severity, decides when the incident is resolved | Doesn't personally debug the system or write the fix |
| Subject Matter Expert (SME) | Diagnosis and remediation for their area (database, networking, the specific service) | Doesn't own communications or overall sequencing across teams |
| Communications Lead | Status updates to stakeholders, customers, and the incident channel on a fixed cadence | Doesn't make technical decisions about the fix |
| Scribe | Timeline of what happened, when, and by whom, feeding the postmortem | Doesn't participate in the technical response itself |
Why the IC role has to be distinct from the SME role
If the IC is also the person elbow-deep in a stack trace, two things suffer at once: the technical dive doesn't get their full attention, and no one is watching the overall picture (are we escalating too slowly, is communications falling behind, has severity changed). Separating the roles means the IC can pull in a second or third SME without needing to personally understand every system, and can make a call like "stop investigating, roll back now" even when an SME would rather keep digging for the root cause, because the IC's job is minimizing impact, not necessarily finding the deepest explanation in the moment.
Escalation triggers and handoffs
An IC should escalate (bring in a more senior IC, or a specific SME) when the current responder hits the edge of their authority or context: severity increasing beyond what the current team can safely own, the fix requiring a decision (like a risky rollback) above the current IC's authorization level, or the incident running long enough that fatigue is a real risk. Handoffs between ICs during a long-running incident follow the same discipline as an on-call shift handoff: the outgoing IC states current status, open decisions, and what's already been tried, in the incident channel, with the incoming IC explicitly confirming they've taken over before the outgoing IC steps back. An IC handoff that happens silently, with no explicit confirmation, is a common source of dropped context in long incidents.
Trade-offs and pitfalls
On a small team, it's tempting to skip a dedicated IC and let the most senior engineer both fix and coordinate; that works for short, simple incidents but breaks down exactly when it matters most, a complex, multi-team incident, because that's when coordination and deep technical focus can no longer be done well by the same person at once. The other common pitfall is an IC who defers every decision back to SMEs instead of actually deciding, which turns the incident into a discussion instead of a response; the IC's authority to make the call, even an imperfect one, quickly, is the actual value of the role.
Unlock Full Question Bank
Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.