On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

MediumTechnical
50 practiced

How do you hand off an on-call shift so nothing falls through the cracks? What does a good handoff actually need to include?

MediumTechnical
49 practiced

During a P1 outage, the first responder doesn't restore service within a few minutes and doesn't acknowledge the page. Walk through what happens next: escalation timeouts, who gets paged, which channels you use, and who ultimately declares a major incident.

HardTechnical
49 practiced

Design the guardrails for a system that lets on-call engineers trigger automated runbook actions directly from an alert. How do you prevent a misfire, or a compromised trigger, from causing a bigger outage than the one it was meant to fix?

EasyTechnical
41 practiced

What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?

MediumSystem Design
57 practiced

Design an on-call escalation system for an organization with multiple teams that need to coordinate coverage across time zones. How do you route pages, prevent alert-noise from cascading into unnecessary escalations, and decide who gets pulled in for a revenue-impacting versus a data-sensitive incident?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.