On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardTechnical
46 practiced

The same incident keeps recurring every month despite repeated fixes. How would you run an RCA that surfaces the systemic process or tooling issue, rather than patching the same symptom again?

HardTechnical
49 practiced

During an incident, what would make you immediately escalate to the security team or an external vendor rather than continuing to handle it yourself? Give concrete examples of the signals that would trigger that call.

HardTechnical
56 practiced

Mid-incident during a Sev1, you discover the runbook you're following has outdated commands that don't work on the current cluster configuration. What do you do to keep the response moving, and how do you make sure the runbook gets fixed afterward?

MediumTechnical
40 practiced

How would you decide whether a runbook is actually ready for on-call use, not just written? What would you check before trusting it during a real incident?

MediumTechnical
52 practiced

Sketch the shape of a runbook for a primary database that's become unresponsive while a replica is still healthy. What are the key decision points, like when do you fail over versus wait, what would you check first, and what does the rollback path look like if the failover goes wrong?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.