On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

EasyTechnical
41 practiced

What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?

MediumTechnical
84 practiced

How would you communicate about an ongoing outage differently to your own engineering team versus to non-technical stakeholders or customers? What changes, and what stays the same?

HardTechnical
49 practiced

A runbook's automated remediation step ran and it caused a partial outage instead of fixing anything. How would you investigate what went wrong, and what would you change to prevent it from happening again?

MediumTechnical
48 practiced

A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?

MediumTechnical
48 practiced

What would it look like to treat runbooks as code: version-controlled, linted, and tested in CI before a change can merge? Walk through how you'd implement that, and how you'd keep thousands of runbooks discoverable enough that an on-call engineer can find the right one within a couple of minutes.

Unlock Full Question Bank

Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.