InterviewStack.io LogoInterviewStack.io

On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardTechnical
40 practiced

You're asked to do a root-cause analysis for an incident where the telemetry is sparse. How would you reconstruct the timeline, and how would you turn the gaps you find into a prioritized plan for what to instrument next?

HardTechnical
53 practiced

One of your critical alerts fires constantly but turns out to be right only a fraction of the time. How would you redesign it so it's trustworthy again, and how would you decide which alerts generally deserve to page a human first?

MediumTechnical
84 practiced

How would you communicate about an ongoing outage differently to your own engineering team versus to non-technical stakeholders or customers? What changes, and what stays the same?

EasyTechnical
41 practiced

What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?

HardTechnical
56 practiced

Mid-incident during a Sev1, you discover the runbook you're following has outdated commands that don't work on the current cluster configuration. What do you do to keep the response moving, and how do you make sure the runbook gets fixed afterward?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.