On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardTechnical
47 practiced

There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?

MediumTechnical
55 practiced

Design a severity rubric, say P0 through P3, for a SaaS product. What determines the level, what SLA applies at each, and who has to be paged?

HardTechnical
56 practiced

Mid-incident during a Sev1, you discover the runbook you're following has outdated commands that don't work on the current cluster configuration. What do you do to keep the response moving, and how do you make sure the runbook gets fixed afterward?

HardTechnical
56 practiced

How would you actually validate that your runbooks work before you need them in a real incident? Describe a program for testing them under realistic conditions.

MediumTechnical
51 practiced

Walk me through how you'd run a postmortem after a Sev1 incident: what data you'd gather, how you separate contributing factors from the root cause, and how you turn it into action items that actually get done.

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.