InterviewStack.io LogoInterviewStack.io

On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

MediumTechnical
48 practiced

A third-party vendor or SaaS dependency you don't control is down and it's affecting your customers. What do you do: what mitigations are actually available to you, how do you communicate about something you can't directly fix, and how do you escalate to the vendor?

EasyTechnical
89 practiced

How do you define severity levels for production incidents (say Sev1 through Sev4), and how does severity map to expected response time and who gets notified?

HardTechnical
46 practiced

How would you build a cost-benefit case for automating a recurring operational task, rather than continuing to have engineers handle it manually?

EasyTechnical
41 practiced

What's the difference between MTTD, MTTA, and MTTR? Given a short incident timeline, how would you calculate each, and what's a common mistake people make when interpreting these numbers?

MediumTechnical
42 practiced

How would you design a fair approach to compensating engineers for on-call work, balancing pay, time off in lieu, and rotation length?

Unlock Full Question Bank

Get access to all 42 On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.