On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

MediumTechnical
40 practiced

You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?

MediumTechnical
44 practiced

How do you make sure postmortem action items actually get done, and that lessons from one incident reach the teams who didn't experience it directly?

HardTechnical
46 practiced

How would you build a cost-benefit case for automating a recurring operational task, rather than continuing to have engineers handle it manually?

MediumTechnical
58 practiced

Some runbook steps involve sensitive actions, like production database admin commands or rotating credentials. How do you control who can run those steps and keep it auditable, without slowing a responder down during a real P1?

EasyTechnical
49 practiced

What's the difference between a runbook and a playbook, and when would you reach for one instead of the other?

Unlock Full Question Bank

Get access to all 41 On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.