On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?
Sample Answer
MTTD is how long a problem existed before anything noticed it. MTTA is how long a human took to acknowledge the alert once it fired. MTTR is how long it took to fully resolve once someone was working it. The most common mistake is reporting a single blended average and treating it as typical, when one long outage in the set is doing all the work.
Definitions
| Metric | Starts at | Ends at | What it measures |
|---|---|---|---|
| MTTD | Failure begins | Alert fires / someone notices | How good detection is |
| MTTA | Alert fires | Human acknowledges | How well paging and routing work |
| MTTR | Acknowledgment | Service fully restored | How fast the response process fixes it, once someone owns it |
(Some teams instead measure MTTR from detection to resolve rather than ack to resolve; either is defensible, but the convention has to be fixed and stated, because mixing them across teams silently changes what the number means.)
Worked example: one incident timeline
| Event | Time |
|---|---|
| Failure begins | 14:00:00 |
| Alert fires (detection) | 14:06:00 |
| Engineer acknowledges | 14:11:00 |
| Service restored | 14:47:00 |
Worked example: averaging across three incidents, and where it goes wrong
| Incident | MTTD | MTTA | MTTR |
|---|---|---|---|
| 1 | 6 | 5 | 36 |
| 2 | 2 | 3 | 20 |
| 3 | 15 | 8 | 54 |
The mean of 36.7 minutes is being pulled up almost entirely by incident 3's 54-minute outlier: the mean sits above two of the three data points (20 and 36), with only the outlier itself larger. The median of {20, 36, 54} is 36, the middle value itself rather than a value inflated by the outlier, so it is a better single-number stand-in for the typical incident than the mean. Reporting mean MTTR alone, without the incident count or a percentile, makes a single bad incident look like the typical case.
Trade-offs and pitfalls
Comparing MTTR across teams that use different start-point conventions is comparing two different metrics wearing the same name; agree on the convention org-wide before benchmarking teams against each other. A dropping mean MTTR can hide a rising incident count: if you're resolving more small incidents faster while one rare severe incident still takes hours, the mean improves and the tail risk hasn't moved at all. Improving MTTD without improving MTTA or MTTR just means you find out about the same slow response faster; treat the three as stages of one pipeline, not independent wins to report separately.
How would you communicate about an ongoing outage differently to your own engineering team versus to non-technical stakeholders or customers? What changes, and what stays the same?
Sample Answer
Direct answer
What changes between internal and external incident communication is technical depth and certainty: engineering updates include specifics, hypotheses, exact systems and logs, mitigation steps, because that audience can act on them, while external updates stick to confirmed customer-visible impact and next-update timing, because speculation shared with customers erodes trust if it turns out wrong. What stays the same is cadence discipline and honesty: both audiences get updates on a predictable schedule, and neither gets a false ETA.
Structured elaboration
| Internal (engineering) | External (customers, stakeholders) | |
|---|---|---|
| Content | Hypotheses, specific systems and logs, exact mitigation steps, owners | Confirmed customer-visible impact, current status, general next steps |
| Language | Technical shorthand is fine | Plain language, no internal system names or jargon |
| Certainty | Unconfirmed working theories can be shared, labeled as such | Only confirmed facts; no speculation presented as cause |
| Cadence | Every 15 to 30 minutes while active | Every 30 to 60 minutes, or on a material status change |
| Who approves | Incident commander, informally | Incident commander plus communications or product sign-off |
| Tone | Direct, action-oriented | Calm, factual, acknowledges impact without over-promising |
Tailoring further by audience
Enterprise customers with contractual SLAs often get a more detailed, sometimes one-to-one update from their account team in addition to the public status page, while free-tier customers rely on the public page alone. The underlying facts should be identical; only the delivery channel and level of individual attention differ. Escalation also has its own trigger: open a live bridge call and involve leadership once the outage crosses a defined severity or duration threshold, such as a sustained top-severity incident running longer than the org's own SLA commitment, not just because internal chat is busy.
What never changes regardless of audience
No blame is assigned, internally or externally, while the incident is still active; naming a cause before it's confirmed just gets walked back later. No ETA gets promised that isn't actually known; "investigating" is an honest status, a fabricated timeline is not.
Worked example
Internal update at the 15-minute mark: "checkout-service 5xx rate is elevated since 14:02 UTC, correlates with a config deploy at 13:58; on-call is rolling back now, expect confirmation in a few minutes; questions in the incident channel." External status page at the 30-minute mark, built only from the confirmed facts in that internal update: "We're investigating an issue causing checkout errors for a subset of users. Our team has identified a likely cause and is applying a fix. Next update in 30 minutes." The external version omits the specific deploy detail, since it isn't customer-actionable and would need walking back if the rollback doesn't fully resolve things, but keeps the same cadence commitment as the internal update.
Trade-offs and pitfalls
- Copy-pasting the internal technical update into the external channel either leaks unnecessary detail or reads as more alarming than the plain-language version would.
- More frequent external updates build trust but risk announcing something not yet confirmed if the situation is still moving fast; anchor the cadence to confirmed milestones, not the clock alone.
- Skipping communications or product sign-off on external messages risks a technically accurate but poorly worded update going out while everyone is heads-down on the actual fix.
A runbook's automated remediation step ran and it caused a partial outage instead of fixing anything. How would you investigate what went wrong, and what would you change to prevent it from happening again?
Sample Answer
When an automated remediation makes things worse, the first move is to stop trusting the automation, not to debug it live: disable the trigger (feature flag or scheduler pause) so it can't fire again while you investigate, then treat the automation's own actions as the incident's primary evidence trail.
Investigation approach
- Pull the automation's own audit log first. What did it decide to do, on what input, at what timestamp? Most remediation frameworks log the triggering condition and the action taken; if this one doesn't, that's itself a finding.
- Reconstruct the precondition it evaluated against. Was the health check it used stale (cached metrics, delayed scrape) or narrower than reality (checked one replica's health, not cluster quorum)?
- Check for concurrency. Did two instances of the same remediation run at once, or did it run while a human was mid-deploy? Interleaved writes to the same resource are a common cause of "fix that broke things."
- Diff the assumed environment against the actual one. Runbooks and remediation scripts encode assumptions (resource names, API versions, cluster topology) that drift silently; check whether the automation was written against a topology that's since changed.
Framework for the fix
- Add a pre-check gate: the remediation must verify the system is in the state it assumes (quorum present, no in-flight deploy, dependency healthy) before acting, and abort loudly if not.
- Make the action idempotent and reversible: re-running it, or running it against a system already in the target state, should be a no-op, and every destructive step needs a paired rollback.
- Bound the blast radius: act on one node/instance first (canary), verify success, then proceed, rather than acting cluster-wide in one shot.
- Add a concurrency guard: a lock or lease so two triggers of the same remediation can't run simultaneously.
- Gate high-impact actions behind a second signal: require the automation to see the problem confirmed by two independent signals (e.g., an alert plus a direct health check) before taking a destructive action, not just one noisy metric.
Worked example
Suppose the remediation is: "if a node reports high memory for 3 consecutive scrapes, cordon and drain it." The postmortem finds the metrics scraper had a 90-second collection lag during a load spike, so by the time the automation cordoned the third node it was actually reading data that was already 4.5 minutes stale (three 90-second-lagged scrapes), and it drained three nodes in the same 2-minute window because the memory spike was cluster-wide, not node-specific. Losing three nodes at once dropped the cluster below quorum for its replicated service, which is the partial outage.
The fix that follows directly from that trace: (a) the pre-check should compare current live memory, not the lagged scrape, before acting; (b) the automation should check how many nodes it has already drained in the current window and refuse to exceed a cap (e.g., no more than one node per 10 minutes) until a human confirms; (c) it should check that the remaining fleet still satisfies quorum before draining another node.
Trade-offs and pitfalls
Adding pre-checks and rate caps makes the remediation slower to react, which is the right trade for anything that can cause an outage of its own; reserve fully unthrottled auto-remediation for actions that are cheap to reverse (like restarting a single stateless pod) and keep caps and human gates on anything that removes capacity or touches shared state. A common wrong turn is to respond to this incident by simply disabling the automation permanently and reverting to manual remediation: that trades a rare automation bug for a much larger population of slower, inconsistent manual responses. The senior move is to narrow what the automation is trusted to do unsupervised, not to abandon automation.
A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?
Sample Answer
Direct answer
Since you can't fix the vendor directly, you run three things in parallel: mitigate the blast radius with tools you do control (circuit breakers, cached or degraded responses, feature flags), communicate honestly about something outside your control, and push on the vendor relationship itself through support escalation and, if needed, contractual SLA terms. The trade-offs are mostly about how aggressively to degrade functionality versus how much broken or stale behavior your customers will tolerate in the meantime.
Structured elaboration
| Mitigation | What it buys you | What it costs |
|---|---|---|
| Circuit breaker / fail fast | Stops the vendor's failure from cascading into your own services | Feature becomes fully unavailable, more visible outage |
| Serve cached or stale data | Feature stays visibly "up" for the user | Risk of showing wrong or outdated information |
| Queue and retry with backoff | No data loss, eventual consistency once the vendor recovers | User sees delay; adds retry/backoff complexity |
| Feature-flag off (graceful degrade) | Predictable, pre-tested reduced experience | Only works if the flag and the reduced UX already exist before the outage |
Evidence to gather before contacting vendor support. Precise timestamps with timezone noted, representative request/response examples (method, URL, headers, correlation IDs), correlated logs from your own edge/load-balancer and application layers, and a clear scope-and-impact statement (which services, what percentage of traffic, which customers, what SLA is at risk). Vague "your API seems down" tickets sit in a generic queue; a ticket with reproducible evidence and a quantified impact gets triaged faster.
Escalating through vendor support tiers. Open the highest applicable severity case with the evidence attached and explicitly request an engineer and a bridge, not just an acknowledgment. If there's no meaningful response within your own internal SLA for that severity, escalate through the account manager or a phone-based escalation path, citing the specific business impact and contractual SLA terms rather than repeating the original ticket.
Communication cadence, using the same severity-driven pattern as an internal incident: acknowledge to affected customers quickly with what's known and any workaround, then update on a fixed cadence (for example every 30 minutes) until resolved, closing with a summary once the vendor confirms the fix.
Worked example
A payments provider starts returning errors for a subset of transactions. Mitigation: flip a feature flag that routes non-critical calls to a queued-retry path with a "processing" state shown to the user, instead of failing checkout outright; this preserves the customer experience for the subset of traffic where a short delay is tolerable, while transactions that genuinely require a synchronous response fail fast with a clear error rather than hanging. Communication: post an initial status update within roughly 15 minutes acknowledging degraded checkout with the workaround in place, then update every 30 minutes. Vendor escalation: open a high-severity vendor ticket with timestamped request/response examples and the affected transaction volume, request a bridge; if no vendor engineer engages within your internal escalation window, escalate via the account manager's phone line, citing the contractual SLA and quantified customer impact.
Trade-offs and pitfalls
- Pitfall: treating a vendor outage as "not our incident" and skipping the postmortem. Root cause may be external, but your own blast-radius design (whether a circuit breaker or cached fallback existed at all) is exactly what a postmortem should examine, since that's the part you actually control.
- Pitfall: promising customers a fix ETA you don't control. Communicate "investigating, using workaround X, next update in 30 minutes" rather than a timeline that depends on someone else's incident response.
- Trade-off: aggressive circuit-breaking protects your own systems fastest but produces the most visible outage; cached/degraded responses are gentler on the user experience but carry a correctness risk if the vendor's data changes underneath the cache. Which one is right depends on how stale or wrong data is allowed to be for that specific feature, which is a product decision, not just an engineering one.
What would it look like to treat runbooks as code: version-controlled, linted, and tested in CI before a change can merge? Walk through how you'd implement that, and how you'd keep thousands of runbooks discoverable enough that an on-call engineer can find the right one within a couple of minutes.
Sample Answer
Direct answer
Store runbooks as Markdown with structured YAML front-matter in one canonical git repository, gate every change behind CI that lints and schema-validates the metadata and dry-run-tests any automatable step, and require review before merge, the same discipline as application code. Discoverability at scale comes from that same metadata: a required service field feeds a searchable index, and alert annotations link straight to the matching runbook, so an on-call engineer reaches the right doc from the page itself instead of searching for it.
Structured elaboration
CI pipeline
flowchart LR
A[Engineer edits runbook.md] --> B[Open PR]
B --> C[CI: lint YAML and Markdown]
C --> D[CI: schema-validate metadata]
D --> E[CI: dry-run automatable steps in sandbox]
E --> F{Checks pass?}
F -->|No| A
F -->|Yes| G[Owner review plus on-call ack]
G --> H[Merge to main]
H --> I[Publish to searchable index]
I --> J[Alert annotations link runbook URL]
Required metadata
| Field | Purpose |
|---|---|
| id | Stable identifier, referenced from alerts and postmortems |
| service | Ties the runbook to a service registry entry |
| owner | Team and on-call rotation responsible for upkeep |
| severity | Drives which validation tier and review gate applies |
| last_reviewed / last_verified | Freshness signal, updated only after a real validation pass |
| automatable | Whether any step can run through the execution platform |
Making thousands of runbooks findable in minutes
Every runbook's service field is validated against a central service registry at CI time, which catches typos and stale references after a rename. Alert rules carry the runbook's canonical URL in their annotations, so paging a human already surfaces the right doc; searching is the fallback path, not the primary one. A generated, always-fresh index, built from the same metadata at publish time, is searchable by service, tag, and severity, so finding the right runbook is a filtered lookup rather than full-text search over prose.
Rationalizing a messy legacy corpus
When inheriting hundreds of duplicated or contradictory legacy runbooks, don't try to review them all before migrating. Import everything into the new repo behind CI, then triage: retire clear duplicates, keeping the most-recently-verified version and redirecting the rest, and mark anything ambiguous as needing review with an expiry date so it doesn't sit unreviewed indefinitely. Treat the migration as incremental and CI-gated, not a one-time cleanup that has to finish before the new system can go live.
Linking postmortems back to runbooks
Postmortem action items that touch a runbook should reference the runbook's id field directly, not just link to it in prose. That makes it possible to build a reverse index of which incidents cited a given runbook's gaps, so recurring weak spots surface automatically instead of living only in scattered postmortem documents.
Worked example
A runbook's front-matter:
---
id: RB-DA-001
title: "Spark job failing with executor OOM"
service: data-ingest-pipeline
owner: team-data-platform
severity: P1
tags: [spark, etl]
last_reviewed: 2026-05-01
automatable: true
---
A PR that removes the owner field or sets severity: P5 (not a valid value in the schema) fails CI's schema-validation step with a specific, actionable error, rather than merging silently and leaving the on-call escalation path undefined for that runbook.
Trade-offs and pitfalls
- Over-engineering the schema so authoring a runbook feels like filing a support ticket discourages contributions; keep required fields minimal, owner, service, severity, last_reviewed, and let the rest stay freeform Markdown.
- Strict CI gating with owner review and on-call acknowledgment slows updates during an actual incident; document a fast-path for post-incident corrections that still requires review, just not before the fix ships in the moment.
- Building the search index without tying it to alerts leaves on-call still searching manually under pressure. The annotation link on the alert is what actually saves time; the index is the fallback.
Unlock Full Question Bank
Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.