InterviewStack.io LogoInterviewStack.io

On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

EasyTechnical
76 practiced

What does 'blameless' actually mean in a blameless postmortem, and why does it matter? What are the essential components of a good postmortem document?

MediumTechnical
46 practiced

Before a new service goes live and starts taking on-call pages, what would you want to see in place? Walk through what a production-readiness review should check.

MediumTechnical
40 practiced

You're deciding which of a few common runbook steps to automate: restarting a cached worker instance, reattaching a detached volume, and running a database schema migration. What criteria would you use to decide whether each should be fully automated, human-in-the-loop, or kept manual?

MediumTechnical
55 practiced

You get paged: p95 latency for a service has spiked and the error rate is climbing, starting a few minutes after a deploy went out. Walk through what you actually do in the first few minutes: what you check, how you decide on a mitigation, and when you'd escalate.

EasyTechnical
41 practiced

What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?

Unlock Full Question Bank

Get access to all 41 On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.