Situation: In my SRE role I’ve been on multiple incidents that lasted hours to days (e.g., a cross-region networking outage that spanned 36 hours). Sustaining focus and morale while avoiding burnout was critical to resolving the issue safely.
Task: My goal was to keep the team effective and healthy—ensure clear ownership, maintain throughput, and enable recovery afterward.
Action:
- Rotations & breaks: I enforce 2–4 hour responder shifts with a clear on-call lead and a relief schedule. Shifts are documented in a shared roster; handoffs use a short checklist (state, next steps, blockers). I pair inexperienced engineers with seniors during high-risk tasks.
- Communication: I set up a single source of truth (incident doc) and one async channel for updates; critical alerts go to voice/phone only. I run 15–20 minute syncs every 2–3 hours to recalibrate priorities and reduce context-switching.
- Psychological safety: I normalize asking for help, explicitly grant permission to step away, and call “time-outs” when frustration rises. I avoid public blame in chat and steer conversations to facts and hypotheses.
- Triage work: I separate investigatory work from low-risk remediation tasks so people can make measurable progress. I assign “ops” roles (monitoring, communication, rollback) so engineers can alternate between intense tasks and lighter coordination.
- Mitigations for fatigue: I provide food, quiet rest spaces, and encourage naps for long incidents. Managers rotate people off-call the next day and limit meetings for responders.
- Post-incident recovery: Mandatory “recovery day” (no on-call, light tasks) for primary responders, followed by a blameless postmortem within 3–5 days. We capture action items, reassign owners, and track them to completion. I also solicit feedback on what processes caused stress and adjust runbooks.
Result: Using these practices reduced error-prone handovers, kept uptime restoration times consistent, and reduced follow-up burnout—teams reported feeling more supported and incidents had clearer remediation paths. The recovery-day policy decreased post-incident sick days by measurable amounts on my team.