On-Call Practices and Runbook Design Questions
Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.
The same incident keeps recurring every month despite repeated fixes. How would you run an RCA that surfaces the systemic process or tooling issue, rather than patching the same symptom again?
Sample Answer
When the same incident keeps recurring despite repeated fixes, the prior RCAs were almost certainly treating a symptom as the root cause; run this RCA by explicitly listing every prior "fix" and asking why each one didn't hold, because the pattern across failed fixes usually points straight at the real systemic gap.
Framework for a systemic RCA
- Build a fix history first, before investigating the current occurrence. For each prior incident of this same recurring issue: what was diagnosed as the cause, what was changed, and did the change actually address that diagnosis or just the immediate symptom? A pattern of "different symptom fixed each time, same underlying trigger every time" is the tell that root cause was never actually found.
- Separate the trigger from the vulnerability. The trigger (a specific deploy, a specific load pattern, a specific external dependency hiccup) may vary each month, but if the same class of trigger keeps causing an outage, the system has a standing vulnerability to that trigger class that no single fix removed. The RCA's job is to name the vulnerability, not just the latest trigger.
- Use the fishbone categories (code, config, infrastructure, process, tooling) to check whether every prior fix landed in the same category. If four consecutive fixes were all code patches but the incident keeps returning, that's evidence the real gap is in process or tooling (no regression test for this class of failure, no canary catching it before full rollout) rather than in any specific line of code.
- Test the systemic hypothesis, don't just assert it. If the hypothesis is "the nightly batch job and the backup window contend for the same database connection pool," verify by reproducing that contention in a controlled environment (isolated run of the batch job during a simulated backup window), not by pattern-matching from the incident timeline alone.
Worked example
A nightly batch job has caused three partial outages in three consecutive months. Prior fixes: month 1, increased the job's timeout (fix addressed "job was timing out"); month 2, added a retry with backoff (fix addressed "job failed transiently"); month 3 is the current incident, and the job is again failing, this time differently, a connection pool exhaustion error. Building the fix history shows a pattern: every fix targeted why the job failed on that specific night, and none asked why the job's failure mode changes every month while the timing (always during the nightly backup window) stays constant. Testing the systemic hypothesis, that the batch job and backup process share a connection pool and the backup's duration has been slowly growing as data volume grows, month-over-month backup duration logs confirm the backup window has grown from roughly 12 minutes to 40 minutes over the quarter, now overlapping the batch job's peak connection usage. That is the systemic cause: the job and backup were never intentionally isolated, and it was invisible for months because the backup was short enough not to overlap.
The durable fix follows from the systemic cause, not the latest symptom: allocate the batch job a dedicated connection pool separate from ad hoc processes, and alert on backup-duration trend (not just backup failure) so a slowly growing resource conflict is visible before it causes an outage again.
Trade-offs and pitfalls
The main pitfall is that a systemic RCA takes longer and produces a less satisfying immediate answer than "here's the line that broke," which creates pressure to ship another symptom-level fix under time pressure; the way to resist that is to make the fix-history review a required first step, not an optional deep-dive, so the systemic question gets asked before anyone commits to a scope. A real trade-off: broadening the RCA to a process or tooling gap usually means the fix is slower to land (new alerting, resource isolation, a regression test suite) than a code patch, so it's worth explicitly stating in the postmortem that a fast interim mitigation (in this example, a manual connection-pool bump) is being paired with the slower structural fix, rather than letting the slow fix block any near-term relief.
During an incident, what would make you immediately escalate to the security team or an external vendor rather than continuing to handle it yourself? Give concrete examples of the signals that would trigger that call.
Sample Answer
Direct answer
Escalate immediately, before finishing your own triage, when the signal suggests the failure mode is adversarial or requires expertise and authority you don't have: active data movement out of the environment, evidence of privilege escalation or persistence, or anything (like a ransom note) that means you're dealing with an attacker's decisions, not a bug. The test is simple: could my next troubleshooting action destroy evidence or tip off an active adversary, and is there plausibly someone on the other side making decisions right now? If either is plausibly yes, escalate now and treat a false alarm as an acceptable cost.
Structured elaboration
Signals that trigger immediate escalation:
The two categories below, exfiltration and privilege escalation, are what an on-call responder most commonly actually sees; ransomware and unexplained vendor compromise are rarer but warrant the same immediate escalation when they happen.
- Data exfiltration (data leaving the environment to somewhere it shouldn't): unusual large outbound transfers to unfamiliar destinations, atypical bulk access to data stores outside normal patterns, DLP (data loss prevention, tooling that flags sensitive data leaving the network) or IDS (intrusion detection system, tooling that flags suspicious network or host activity) hits correlated with credential use outside its normal baseline.
- Privilege escalation (an account gaining access it shouldn't have) or persistence (an attacker planting a way to keep that access after you think you've cleaned up): unexpected new admin accounts, IAM (identity and access management, the system controlling who can access what) role or policy changes nobody on the team made, SSH keys added across multiple hosts, processes or webshells (a script an attacker leaves behind for remote access) that survive a reboot.
- Extortion or ransomware: encrypted files, a ransom note, or backup systems suddenly and specifically inaccessible.
- A vendor or dependency failure that doesn't match any known bug or outage pattern, especially if the vendor confirms unauthorized access on their side.
The self-handle-vs-escalate test, three fast questions to ask before touching anything further:
- Is this reversible with a config change I already understand, or does it plausibly require forensics, legal, or vendor-security expertise I don't have?
- Could my next action (reboot, delete, revoke, patch) destroy evidence that security or legal will need?
- Is there a plausible person on the other side actively making decisions right now?
Any "yes" means escalate now, not after you've tried a few things yourself.
Worked example
An alert fires showing an outbound transfer of several gigabytes to an IP address with no history in that account's traffic baseline, immediately after a service account's credentials were used from an unfamiliar region. This hits two of the three escalate-now questions: it plausibly needs forensics expertise (a compromised credential's blast radius isn't something to guess at), and clumsy remediation could tip off an active attacker or destroy evidence. What not to do: reboot the host (destroys volatile memory that forensics needs) or immediately revoke every credential in sight without coordinating (can alert an attacker who's still active and cause them to accelerate or cover tracks). What to do instead: isolate the host at the network layer rather than powering it off, preserve logs and a memory snapshot, and page security immediately with the specific evidence (timestamps, transfer volume, destination IP, the credential and region involved) rather than a vague "something looks off." Security, not the on-call responder, then decides the sequencing of credential revocation, further isolation, and any legal or disclosure steps.
Trade-offs and pitfalls
- Over-escalating everything ambiguous as a security incident burns the security team's trust and slows down genuinely urgent calls later, the same alert-fatigue dynamic that applies to paging in general applies to security escalation specifically.
- Under-escalating to avoid looking alarmist risks destroying evidence or letting an active attacker continue while you troubleshoot as if it were an ordinary bug.
- The resolution isn't "use better judgment in the moment," because incident stress predictably degrades judgment on exactly these ambiguous calls; it's having a short, memorized decision list (the three questions above) that doesn't require clear thinking under pressure to apply correctly.
- Pitfall: treating "call security" as an admission that you failed. In a genuinely blameless culture, escalating on a plausible-but-unconfirmed signal is the correct, expected behavior, not something a responder should be second-guessed for after the fact if it turns out to be a false alarm.
Mid-incident during a Sev1, you discover the runbook you're following has outdated commands that don't work on the current cluster configuration. What do you do to keep the response moving, and how do you make sure the runbook gets fixed afterward?
Sample Answer
Direct answer
Keep the incident moving without trusting the stale command: switch to safe, read-only discovery to re-derive the actual current state instead of assuming the runbook's exact syntax still matches reality, get a second responder to sanity-check any ad-hoc workaround before running it, and narrate every command and its outcome in the incident channel as you go. Afterward, treat the correction as a first-class, owned follow-up: file it with the incident as evidence and require the same review the runbook normally gets, rather than merging a fix drafted under adrenaline with no second pair of eyes.
Structured elaboration
Keeping the response moving
- Don't keep retrying the stale command hoping it starts working; that burns clock on a Sev1.
- Fall back to read-only discovery to find what actually changed: list the current resource names or config instead of assuming the runbook's exact prior values, and check whether a known change (a migration, a rename, a tool upgrade) explains the mismatch.
- If a workaround command is genuinely needed, treat it like an experiment: scope it to the smallest possible blast radius (a single pod, host, or canary) if at all possible, and have a second responder review it before running anything destructive.
- Narrate in the incident channel as you go: exact command, who ran it, what happened. This live log is what makes the eventual runbook fix accurate instead of a reconstruction from memory the next morning.
Getting the runbook actually fixed afterward
- File the correction as its own owned action item, not "someone should update this."
- Attach evidence: the incident timeline entries showing what actually worked, captured live rather than recalled later.
- Route the fix through the same review the runbook would normally require. A fix drafted under incident pressure is exactly the kind of change that benefits from a second reviewer, not an exception to needing one.
- If the drift has a systemic cause, such as no process tying documentation updates to the change that invalidated it, raise that separately as its own postmortem action item, not just a one-off doc patch.
Worked example
kubectl rollout restart deployment/checkout -n prod fails with "deployment not found." The responder falls back to read-only discovery: kubectl get deploy -A | grep checkout shows the deployment now lives in namespace checkout-prod, after a namespace-per-service migration weeks earlier that never touched the runbook. The corrected command runs, the service recovers, and both the failed and working commands are logged with timestamps in the incident doc as they happen. The resulting follow-up action item reads: "update runbook RB-042's namespace reference and add a namespace-lookup step instead of a hardcoded name; owner: platform team; verified by a peer review plus a sandboxed dry-run before merge."
Trade-offs and pitfalls
- Fixing the runbook file directly, mid-incident, with no review is a common shortcut; the correct fix ships as a follow-up change through the normal review gate, informed by what was learned live.
- Treating the ad-hoc working command as tribal knowledge instead of writing it down immediately is how the same staleness reappears at the next incident; capture it in the channel the moment it works, not after the retro.
- Verifying a corrected command in a sandbox before trusting it live is safer, but a Sev1 usually doesn't have that time; this is really an argument for building runbooks with idempotent, safe-to-retry commands and dynamic lookups in the first place, so on-call isn't forced into that trade-off during the incident.
How would you decide whether a runbook is actually ready for on-call use, not just written? What would you check before trusting it during a real incident?
Sample Answer
A runbook being written isn't the same as it being trustworthy under pressure: writing tests whether the author understood the system, while readiness tests whether someone else, half-awake at 3am, can follow it and get the right outcome. Check readiness by having someone who didn't write it actually execute it against a real (or realistic) system, not by reading it for completeness.
What to check before trusting a runbook
- Has anyone other than the author run it? A runbook the author has never handed to someone else is unverified by definition; the author's own mental model fills gaps a stranger will trip on (an assumed tool is installed, an assumed permission is already granted, a step that says "check the dashboard" without saying which one).
- Are the steps executable as written, not just described? "Restart the service" is a description; "run
systemctl restart payments-apion each of the 3 hosts listed in the service registry" is executable. If a step requires judgment the runbook doesn't supply (how do you know which hosts?), that's a gap, not an acceptable level of abstraction. - Does it state what success looks like? A remediation step without a stated verification step (what metric or log line confirms this worked) leaves the responder guessing whether to move to the next step or escalate.
- Is it safe to run when the diagnosis is wrong? Incident responders under pressure sometimes run the wrong runbook, or run the right one when the actual cause differs from what it assumes. Check whether each destructive step is reversible, and whether the runbook states a precondition to verify before acting ("only run this if X").
- Is it current? Check for an owner and a last-verified date; a runbook referencing a deprecated tool, an old cluster name, or a rotation that no longer exists is worse than no runbook, because it costs time before the responder realizes it's wrong.
How to actually verify these, not just check for their presence
- Tabletop walkthrough: someone unfamiliar with the runbook reads it aloud, step by step, without help from the author, narrating what they'd actually type or click. Gaps surface immediately as "wait, what do I do here?" moments.
- Staging or canary drill: run the actual remediation against a staging environment or a single canary instance, and time it. This catches steps that look right on paper but fail against the real system (a command with an outdated flag, a permission the on-call role doesn't actually have).
- Cold-open test: hand it to someone with zero context on this specific service (not zero context on the systems generally) and see if they can act on it in under a target time, without pinging the original author. If they can't, the runbook is only usable by the person who wrote it, which defeats the point.
Worked example
A runbook for "database replica lag alert" says: "Check replica lag, if high, failover to standby." A cold-open drill immediately exposes three gaps: no link to where replica lag is displayed, no threshold for what counts as "high" (the alert already fired, so this should already be answered, but the runbook re-asks the question), and "failover to standby" doesn't say which standby if there are multiple, or what to verify afterward to confirm the failover succeeded rather than made things worse. Fixing it: link the specific dashboard panel, state the alert's own threshold so the runbook doesn't require re-deciding it, name the failover command with the specific standby-selection logic, and add a verification step ("confirm write latency on the new primary is under 50ms and replica lag on remaining replicas is decreasing"). Re-running the cold-open drill after the fix, the same tester completes it without asking a clarifying question, which is the actual pass condition.
Trade-offs and pitfalls
Running live drills has a real cost in engineering time and, for staging drills, some risk if the environment isn't well isolated from production; the return is worth it for any runbook covering a high-severity or destructive action, and can be scaled down to tabletop-only for low-risk, easily reversible ones. A common wrong turn is treating runbook review as a documentation-quality pass (is it well written, does it have headers) rather than an execution test; a beautifully formatted runbook that's never been run by anyone but its author is still unverified.
Sketch the shape of a runbook for a primary database that's become unresponsive while a replica is still healthy. What are the key decision points, like when do you fail over versus wait, what would you check first, and what does the rollback path look like if the failover goes wrong?
Sample Answer
Direct answer
A runbook for an unresponsive primary with a healthy replica has to separate "the primary looks dead" from "the primary is dead." Check whether it truly refuses writes versus is just slow or lock-contended, confirm at least one replica is caught up enough to safely promote, and only fail over with an explicit operator confirmation, since promotion is usually irreversible without a full topology rebuild. If no replica is safely caught up, the decision becomes an RTO (recovery time objective: how long the service can stay down) versus data-loss trade-off that gets escalated, not made unilaterally by whoever is holding the pager.
Structured elaboration
What to check first, before touching anything
- Confirm the primary is actually unresponsive: connectivity, a real write test, and whether this looks like a network partition, a true database hang, or lock contention.
- If a deadlock is suspected: stop new writes at the application or proxy layer, identify the blocking transaction(s) (via the database's active-session view), and safely terminate the offending transaction(s) before even considering failover. A large share of "unresponsive primary" pages turn out to be a stuck writer, not a dead node, and killing the offending transaction is far cheaper than a failover.
- Check replica health: replication lag, whether the replica process is actually running, and whether the replica's own health checks pass.
Decision tree
flowchart TD
A[Primary unresponsive alert fires] --> B{Primary accepting writes?}
B -->|Yes, just slow| C[No failover: investigate latency/locks]
B -->|No| D{Healthiest replica lag under 30s?}
D -->|Yes| E[Get operator confirmation for destructive failover]
E --> F[Promote healthiest replica]
F --> G[Repoint app connection string and DNS]
D -->|No, all replicas lagging| H[Escalate to DBA: weigh RTO vs data loss]
H --> I{Accept data loss to restore now?}
I -->|Yes| F
I -->|No| J[Wait, restore primary from backup]
Decision points explained
- Fail over only once the primary is confirmed non-writable, not just slow, and a replica exists with lag under an agreed threshold. Failing over while the primary is merely slow risks split-brain: two nodes both accepting writes.
- Wait and investigate when the primary is reachable and still landing writes, even slowly.
- A destructive failover needs an explicit human confirmation step, not silent automation, precisely because it is hard to reverse.
- Coordinate the failover live with the owning application team and a DBA before promoting: they know write patterns (in-flight jobs, batch writers) that a generic runbook can't encode, and the DBA can judge whether the replica is truly safe to promote.
Rollback path if the failover goes wrong
- If the promoted replica can't sustain traffic, or the DNS/connection-string cutover doesn't propagate cleanly, first check whether the original primary has since recovered and is not diverged. If it has and is clean, route traffic back to it.
- If the original primary is diverged or unclear, treat this as a second incident and run it through the same decision tree again, treating the newly promoted node as the current primary.
- Fence the old primary (block it from accepting writes) after promotion so it can't silently rejoin as a second writer.
Worked example
Two replicas exist when the primary stops accepting writes: replica A reports 3 seconds of replication lag, replica B reports 45 seconds. The runbook's threshold is "promote only if lag is under 30 seconds." Replica A clears the threshold and replica B does not, so the on-call engineer gets operator confirmation and promotes replica A, accepting up to 3 seconds of potential write loss rather than 45. If both replicas had shown 45 seconds of lag, the runbook routes to the RTO-versus-data-loss escalation instead of an automatic promotion.
Trade-offs and pitfalls
- Fully automating the failover removes the human check that prevents split-brain during a network partition, where the primary might be up but simply unreachable from the monitoring node. That ambiguity is exactly why the confirmation step exists.
- Waiting longer to confirm the primary is truly dead reduces the risk of an unnecessary failover but extends downtime; the health-check timeout and lag threshold are the levers that tune this trade-off.
- Forgetting to fence the old primary after promotion is the most common way a "successful" failover turns into a second, worse incident.
Unlock Full Question Bank
Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.