InterviewStack.io LogoInterviewStack.io

On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardBehavioral
49 practiced

During a postmortem, the incident commander singles out one engineer as the cause of the outage. How do you respond in the moment to preserve a blameless culture, without letting accountability for the fix slide?

HardTechnical
49 practiced

A runbook's automated remediation step ran and it caused a partial outage instead of fixing anything. How would you investigate what went wrong, and what would you change to prevent it from happening again?

EasyTechnical
55 practiced

Why do runbooks tend to go stale in a large engineering org? What are the common root causes, and what would you actually do about each one?

HardTechnical
46 practiced

The same incident keeps recurring every month despite repeated fixes. How would you run an RCA that surfaces the systemic process or tooling issue, rather than patching the same symptom again?

MediumTechnical
55 practiced

Design a severity rubric, say P0 through P3, for a SaaS product. What determines the level, what SLA applies at each, and who has to be paged?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.