On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardTechnical
47 practiced

There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?

HardBehavioral
49 practiced

During a postmortem, the incident commander singles out one engineer as the cause of the outage. How do you respond in the moment to preserve a blameless culture, without letting accountability for the fix slide?

MediumTechnical
51 practiced

Walk me through how you'd run a postmortem after a Sev1 incident: what data you'd gather, how you separate contributing factors from the root cause, and how you turn it into action items that actually get done.

MediumTechnical
58 practiced

An automated remediation keeps firing and the service flips between healthy and unhealthy as a result, a feedback loop. How would you design the automation to avoid this kind of flapping?

MediumTechnical
40 practiced

You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.