InterviewStack.io LogoInterviewStack.io

On-Call Practices and Runbook Design Questions

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

HardTechnical
49 practiced

Design the guardrails for a system that lets on-call engineers trigger automated runbook actions directly from an alert. How do you prevent a misfire, or a compromised trigger, from causing a bigger outage than the one it was meant to fix?

MediumTechnical
58 practiced

How would you measure whether an on-call rotation is sustainable or quietly burning people out? What would you actually track?

EasyBehavioral
46 practiced

Describe your approach and boundaries for being on-call. What kinds of alerts should page you versus just show up in Slack or email, and how do you protect your work-life balance while still being reliable?

MediumTechnical
55 practiced

You get paged: p95 latency for a service has spiked and the error rate is climbing, starting a few minutes after a deploy went out. Walk through what you actually do in the first few minutes: what you check, how you decide on a mitigation, and when you'd escalate.

MediumTechnical
54 practiced

Where's the line between an operational runbook and a security incident-response playbook, and how do you keep a responder from accidentally leaking something sensitive, like pasting a live credential, into a runbook they're editing during an incident?

Unlock Full Question Bank

Get access to all On-Call Practices and Runbook Design interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.