On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
How would you design a fair approach to compensating engineers for on-call work, balancing pay, time off in lieu, and rotation length?
Sample Answer
Compensation and schedule design are two separate levers, and both need to move: pay a base standby stipend for availability, add per-incident pay or time-off-in-lieu for the work actually done, and size the rotation and rest guarantees so the schedule itself isn't relying on money to make an unsustainable load tolerable.
Compensation model components
| Component | What it covers | Typical structure | Why it's separate |
|---|---|---|---|
| Standby stipend | Being reachable and ready, whether or not paged | Fixed weekly amount | Compensates the constraint on personal time even in a quiet week |
| Per-incident pay or TOIL | Actual time spent responding | Hourly rate, or banked time at 1x to 2x | Rewards work done and discourages treating pages as free to the business |
| Leveling credit | Career recognition for on-call excellence | Counted explicitly in review/promotion criteria | Stops strong on-call performers from being penalized for time not spent on visible project work |
| Rest guarantee | Recovery time | Mandatory hours off after a heavy incident or night shift | Protects sustainability independent of pay |
Worked example: one on-call week
Stipend $250 + 6 hours of actual incident work at $40/hr:
Pay=250+(6×40)=250+240=$490If the same 6 hours bank as TOIL at a 1.5x rate for after-hours work:
TOIL banked=6×1.5=9 hoursSchedule practices that reduce the load pay has to compensate for
Primary/secondary tiers so one page doesn't always land on the same person; a cap on consecutive on-call weeks per engineer; shorter rotations (fewer consecutive days of stress, more handoffs) traded against longer rotations (fewer handoffs, more concentrated fatigue), sized to team headcount rather than picked arbitrarily; and follow-the-sun coverage once the team is large and distributed enough to make timezone handoffs cheaper than overnight pages.
Trade-offs and pitfalls
Per-incident pay can invite gaming in both directions, either padding logged hours or under-reporting to avoid looking like a "high maintenance" service; review incident-hour claims against the paging log rather than trusting self-reports alone. Contractors and salaried employees often need different structures (cash versus TOIL), and a single model rarely fits both cleanly. Regional labor law varies significantly, some jurisdictions treat standby time itself as compensable working time, so confirm with legal or HR before setting a global policy rather than assuming one region's rules generalize. There's no dominant answer on rotation length; it has to be sized to team size and incident frequency, not copied from another team's policy.
A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?
Sample Answer
Direct answer
Since you can't fix the vendor directly, you run three things in parallel: mitigate the blast radius with tools you do control (circuit breakers, cached or degraded responses, feature flags), communicate honestly about something outside your control, and push on the vendor relationship itself through support escalation and, if needed, contractual SLA terms. The trade-offs are mostly about how aggressively to degrade functionality versus how much broken or stale behavior your customers will tolerate in the meantime.
Structured elaboration
| Mitigation | What it buys you | What it costs |
|---|---|---|
| Circuit breaker / fail fast | Stops the vendor's failure from cascading into your own services | Feature becomes fully unavailable, more visible outage |
| Serve cached or stale data | Feature stays visibly "up" for the user | Risk of showing wrong or outdated information |
| Queue and retry with backoff | No data loss, eventual consistency once the vendor recovers | User sees delay; adds retry/backoff complexity |
| Feature-flag off (graceful degrade) | Predictable, pre-tested reduced experience | Only works if the flag and the reduced UX already exist before the outage |
Evidence to gather before contacting vendor support. Precise timestamps with timezone noted, representative request/response examples (method, URL, headers, correlation IDs), correlated logs from your own edge/load-balancer and application layers, and a clear scope-and-impact statement (which services, what percentage of traffic, which customers, what SLA is at risk). Vague "your API seems down" tickets sit in a generic queue; a ticket with reproducible evidence and a quantified impact gets triaged faster.
Escalating through vendor support tiers. Open the highest applicable severity case with the evidence attached and explicitly request an engineer and a bridge, not just an acknowledgment. If there's no meaningful response within your own internal SLA for that severity, escalate through the account manager or a phone-based escalation path, citing the specific business impact and contractual SLA terms rather than repeating the original ticket.
Communication cadence, using the same severity-driven pattern as an internal incident: acknowledge to affected customers quickly with what's known and any workaround, then update on a fixed cadence (for example every 30 minutes) until resolved, closing with a summary once the vendor confirms the fix.
Worked example
A payments provider starts returning errors for a subset of transactions. Mitigation: flip a feature flag that routes non-critical calls to a queued-retry path with a "processing" state shown to the user, instead of failing checkout outright; this preserves the customer experience for the subset of traffic where a short delay is tolerable, while transactions that genuinely require a synchronous response fail fast with a clear error rather than hanging. Communication: post an initial status update within roughly 15 minutes acknowledging degraded checkout with the workaround in place, then update every 30 minutes. Vendor escalation: open a high-severity vendor ticket with timestamped request/response examples and the affected transaction volume, request a bridge; if no vendor engineer engages within your internal escalation window, escalate via the account manager's phone line, citing the contractual SLA and quantified customer impact.
Trade-offs and pitfalls
- Pitfall: treating a vendor outage as "not our incident" and skipping the postmortem. Root cause may be external, but your own blast-radius design (whether a circuit breaker or cached fallback existed at all) is exactly what a postmortem should examine, since that's the part you actually control.
- Pitfall: promising customers a fix ETA you don't control. Communicate "investigating, using workaround X, next update in 30 minutes" rather than a timeline that depends on someone else's incident response.
- Trade-off: aggressive circuit-breaking protects your own systems fastest but produces the most visible outage; cached/degraded responses are gentler on the user experience but carry a correctness risk if the vendor's data changes underneath the cache. Which one is right depends on how stale or wrong data is allowed to be for that specific feature, which is a product decision, not just an engineering one.
How would you keep an organization's runbooks accurate and useful over time instead of letting them rot? Talk through ownership and how you'd catch a runbook that's gone stale before someone relies on it during an incident.
Sample Answer
Runbooks rot because nothing forces them to change when the system they describe changes. Keeping them accurate means giving every runbook a named owner, tying updates to the events that actually invalidate a runbook (a relevant code or infra change, or an incident where it was used), and having a lightweight, recurring check that catches staleness before an incident does, rather than relying on someone remembering to update it.
Ownership
Every runbook has one named primary owner, the team or person who owns the service it covers, recorded in the runbook itself, not a separate spreadsheet that gets forgotten. Ownership isn't honorary: the owner is the one who signs off that the runbook is still accurate at each review point, and the one paged if a stale runbook causes a bad outcome during an incident.
What triggers a review, not just a calendar date
Calendar-only review cadences ("review every quarter") catch some staleness but miss the more common case: a runbook goes stale the day the system changes underneath it, not on a schedule. Two triggers matter more than the calendar:
- Change-linked: any deploy that touches the commands, infra, or thresholds a runbook references should require a runbook update as part of that same change, enforced by a checklist item on the pull request, not a follow-up ticket that competes with the next sprint.
- Incident-linked: every time a runbook is actually used during an incident, the post-incident review includes a specific question, did the runbook match reality, and any gap becomes a tracked follow-up before the incident is closed.
Lifecycle stages
| Stage | Meaning | Who moves it | Trigger |
|---|---|---|---|
| Active | Verified accurate, safe to follow blind | Owner | Passed its last review or was just used successfully in an incident |
| Needs review | A linked change or incident flagged it as possibly stale | Owner, auto-flagged | Change-linked or incident-linked trigger fires |
| Deprecated | Still readable but no longer the source of truth | Owner | Replacement runbook exists, or the failure mode it covers no longer applies |
| Archived | Removed from the on-call surface entirely | Owner | Deprecated for a defined grace period with no further reliance |
Catching a stale runbook before someone relies on it
A stale runbook is genuinely dangerous exactly because it looks trustworthy right up until the moment it's wrong. The cheapest catch is a light automated check, does every command in the runbook reference a tool, dashboard, or endpoint that still exists, run periodically and flagging anything broken for owner review. The more valuable catch is a periodic game-day: pick a runbook, have someone unfamiliar with the system try to follow it against a staging environment, and see where it breaks. Anything that trips someone up in a drill would have tripped up the on-call engineer at 3am.
Worked example: what enforcement actually looks like
A pull request changes the retry/backoff config for a service. The PR template includes a checkbox, "does this change any runbook referenced by this service's on-call docs?" Because the change alters a value a runbook's remediation step depends on, the author checks yes and links the runbook update in the same PR. Two months later, that runbook gets used during an incident; the post-incident review confirms the values matched, so the runbook stays Active with its last-reviewed date updated, no separate ticket needed because the change-linked trigger already did the work.
Trade-offs and pitfalls
Tying every infra change to a mandatory runbook update adds friction to routine PRs, so the check needs to be scoped narrowly (does this specific change affect a specific runbook) rather than a blanket "update all docs" gate that people learn to click through without reading. The most common failure mode is having a lifecycle model on paper but no one actually enforcing the Needs review to Active transition, so runbooks accumulate silently in Needs review and the label stops meaning anything; the fix is making that queue visible (a dashboard, not a buried label) and reviewing it in the same recurring meeting as on-call handoffs.
During an incident, what would make you immediately escalate to the security team or an external vendor rather than continuing to handle it yourself? Give concrete examples of the signals that would trigger that call.
Sample Answer
Direct answer
Escalate immediately, before finishing your own triage, when the signal suggests the failure mode is adversarial or requires expertise and authority you don't have: active data movement out of the environment, evidence of privilege escalation or persistence, or anything (like a ransom note) that means you're dealing with an attacker's decisions, not a bug. The test is simple: could my next troubleshooting action destroy evidence or tip off an active adversary, and is there plausibly someone on the other side making decisions right now? If either is plausibly yes, escalate now and treat a false alarm as an acceptable cost.
Structured elaboration
Signals that trigger immediate escalation:
The two categories below, exfiltration and privilege escalation, are what an on-call responder most commonly actually sees; ransomware and unexplained vendor compromise are rarer but warrant the same immediate escalation when they happen.
- Data exfiltration (data leaving the environment to somewhere it shouldn't): unusual large outbound transfers to unfamiliar destinations, atypical bulk access to data stores outside normal patterns, DLP (data loss prevention, tooling that flags sensitive data leaving the network) or IDS (intrusion detection system, tooling that flags suspicious network or host activity) hits correlated with credential use outside its normal baseline.
- Privilege escalation (an account gaining access it shouldn't have) or persistence (an attacker planting a way to keep that access after you think you've cleaned up): unexpected new admin accounts, IAM (identity and access management, the system controlling who can access what) role or policy changes nobody on the team made, SSH keys added across multiple hosts, processes or webshells (a script an attacker leaves behind for remote access) that survive a reboot.
- Extortion or ransomware: encrypted files, a ransom note, or backup systems suddenly and specifically inaccessible.
- A vendor or dependency failure that doesn't match any known bug or outage pattern, especially if the vendor confirms unauthorized access on their side.
The self-handle-vs-escalate test, three fast questions to ask before touching anything further:
- Is this reversible with a config change I already understand, or does it plausibly require forensics, legal, or vendor-security expertise I don't have?
- Could my next action (reboot, delete, revoke, patch) destroy evidence that security or legal will need?
- Is there a plausible person on the other side actively making decisions right now?
Any "yes" means escalate now, not after you've tried a few things yourself.
Worked example
An alert fires showing an outbound transfer of several gigabytes to an IP address with no history in that account's traffic baseline, immediately after a service account's credentials were used from an unfamiliar region. This hits two of the three escalate-now questions: it plausibly needs forensics expertise (a compromised credential's blast radius isn't something to guess at), and clumsy remediation could tip off an active attacker or destroy evidence. What not to do: reboot the host (destroys volatile memory that forensics needs) or immediately revoke every credential in sight without coordinating (can alert an attacker who's still active and cause them to accelerate or cover tracks). What to do instead: isolate the host at the network layer rather than powering it off, preserve logs and a memory snapshot, and page security immediately with the specific evidence (timestamps, transfer volume, destination IP, the credential and region involved) rather than a vague "something looks off." Security, not the on-call responder, then decides the sequencing of credential revocation, further isolation, and any legal or disclosure steps.
Trade-offs and pitfalls
- Over-escalating everything ambiguous as a security incident burns the security team's trust and slows down genuinely urgent calls later, the same alert-fatigue dynamic that applies to paging in general applies to security escalation specifically.
- Under-escalating to avoid looking alarmist risks destroying evidence or letting an active attacker continue while you troubleshoot as if it were an ordinary bug.
- The resolution isn't "use better judgment in the moment," because incident stress predictably degrades judgment on exactly these ambiguous calls; it's having a short, memorized decision list (the three questions above) that doesn't require clear thinking under pressure to apply correctly.
- Pitfall: treating "call security" as an admission that you failed. In a genuinely blameless culture, escalating on a plausible-but-unconfirmed signal is the correct, expected behavior, not something a responder should be second-guessed for after the fact if it turns out to be a false alarm.
What signals, the 'golden signals', would you monitor for a web service, and how do you decide which ones should actually page a human versus just show up on a dashboard?
Sample Answer
The four golden signals for a web service are latency, traffic, errors, and saturation. Which of them should page a human comes down to one rule: page on symptoms that are actively hurting users right now, and let everything else, including most resource metrics, sit on a dashboard until it either crosses into user-visible territory or a slower trend alert catches it.
The four signals
- Latency: how long requests take, measured at p95/p99 (the 95th and 99th percentile: the response time that 95%, or 99%, of requests come in under) rather than the average, since the average hides the slow tail that users actually feel.
- Traffic: request volume and shape; a sudden drop is often as meaningful as a spike.
- Errors: the rate of failed requests, both hard failures (5xx) and soft ones (200s carrying wrong data).
- Saturation: how full a resource is, CPU, memory, connection pools, queue depth, which predicts trouble before it becomes user-visible.
What should page, versus what belongs on a dashboard
The dividing line is whether the signal is symptom-based (something a user is experiencing right now) or cause-based (something that predicts a symptom later). Symptom-based signals, elevated error rate, elevated p95 latency, should page immediately because every minute of delay is a minute of real user pain. Cause-based signals, CPU at 85%, a queue growing, belong on a dashboard and get an alert only if they're trending toward actually breaching a symptom threshold; otherwise the on-call person gets paged for problems that haven't happened yet and often self-resolve.
SLO burn-rate alerting, derived
A cleaner way to decide the paging threshold than picking a number by feel is to work from the error budget itself. Take a service with a 99.9% monthly availability SLO evaluated over a 30-day, 720-hour window.
error budget=1−0.999=0.001=0.1% budget in minutes=0.001×30×24×60=43.2 minutes per monthThat's the total "allowed" downtime for the month. A burn rate of 1x means the service is consuming budget exactly on pace to use all 43.2 minutes by month end. To find the burn rate that should page immediately versus the one that should just open a ticket, fix a target: an alert should fire fast enough that a real outage doesn't quietly eat the whole month's budget before anyone notices.
For a 1-hour detection window, solve for the burn multiplier that consumes 2% of the monthly budget within that hour:
m×7201=0.02⟹m=0.02×720=14.4For a 6-hour detection window, solve for the multiplier that consumes 5% of the budget:
m×7206=0.05⟹m=0.05×6720=6So: a 14.4x burn sustained for 1 hour (meaning the error rate is running at 14.4 times what the SLO allows) pages a human immediately, because at that rate the entire month's budget would be gone in about 50 hours. A 6x burn sustained for 6 hours only needs a ticket, since it's real but slow enough to catch on the next business day without waking anyone.
Putting it together for a web service
- Page immediately: 5xx rate or p95 latency crossing a threshold that maps to a fast SLO burn (roughly 14x or higher).
- Page on a slower cadence or ticket: a sustained but slow burn (single digits), or a saturation metric trending toward its symptom threshold.
- Dashboard only, no alert: traffic shape changes, CPU/memory levels that haven't threatened a symptom, anything already explained by a known deploy or maintenance window.
Trade-offs and pitfalls
Paging on every saturation metric produces alert fatigue fast, because resource usage fluctuates constantly without breaking anything; the fix is always tying a page to a symptom or a budget-burn calculation, not a raw resource number. The opposite mistake, paging only on hard 5xx errors, misses slow degradation (elevated latency, soft failures returning 200 with bad data) that erodes user trust just as much; that's exactly what the multi-window burn-rate approach is for, since it catches both fast, dramatic breaches and slow, sustained ones on two different clocks.
Unlock Full Question Bank
Get access to all 49 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.