On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

EasyTechnical
55 practiced

Why do runbooks tend to go stale in a large engineering org? What are the common root causes, and what would you actually do about each one?

HardTechnical
47 practiced

There's a real tension between making alerts more sensitive so you catch problems earlier, and suppressing alerts so responders aren't fatigued. How would you approach that trade-off, and how would you safely test a change to alert thresholds before rolling it out everywhere?

MediumTechnical
55 practiced

How would you keep an organization's runbooks accurate and useful over time instead of letting them rot? Talk through ownership and how you'd catch a runbook that's gone stale before someone relies on it during an incident.

MediumTechnical
48 practiced

A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?

MediumTechnical
58 practiced

How would you measure whether an on-call rotation is sustainable or quietly burning people out? What would you actually track?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.