Situation: 48 hours before an SLA breach that impacts a strategic enterprise customer; engineering capacity is tied to roadmap work.
Task: Rally operations, engineering, account management, and the customer to prioritize a fix, implement short-term mitigations, and prevent business impact.
48-hour plan (chronological):
Hour 0–1: Convene an emergency triage call (PM leads)
- Attendees: Eng lead, Ops/SRE lead, Customer Success (CS)/AM, Legal/Contracts (if needed), CTO/VP Eng (decision authority).
- Output: Agreed scope (what causes breach), severity, estimated time-to-fix (TTF), and mitigation options. CTO/VP Eng granted authority to re-prioritize roadmap for 48 hours.
Hour 1–6: Rapid technical assessment + mitigation
- Engineering: produce a one-page runbook with root cause hypothesis, hotfix vs rollback decision, and required resources.
- Ops: deploy monitoring, throttles, circuit-breakers, or routing changes to reduce customer impact.
- AM/CS: draft a transparent customer message (agreed wording) offering mitigation, timeline, and compensation if needed.
Hour 6–24: Execute hotfix path and maintain communications
- Engineering works in focused war-room sprints with SRE support; product owner removes blockers (approvals, infra).
- Ops validates mitigations and increases alerting.
- AM/CS call the customer (named contact) within 6 hours to acknowledge issue, share mitigation steps, and set expectations. Follow-up every 6 hours.
Hour 24–48: Verify resolution, post-incident plan, customer close-out
- Run validation tests with customer if needed.
- Prepare RCA outline and proposed roadmap adjustments to prevent recurrence.
- Formalize compensation / SLA remediation with Legal and AM if breach occurs.
How I’d persuade each stakeholder:
- Engineering: Emphasize customer strategic value, show quantified business impact (ARR at risk, escalation risk), remove blockers (approve OT, allocate extra headcount, provide clear acceptance criteria). Offer to pause noncritical approvals and own cross-team coordination.
- Operations/SRE: Position mitigations as temporary protects for system health; commit to rollback plan and post-incident cleanup; provide executive backing for fast deploy windows.
- Account Management/CS: Give them precise messaging, timeline, and mitigation actions so they can retain trust with the customer; commit to transparent updates and compensation options.
- Customer: Be proactive, transparent, and accountable — explain what we know, what we’re doing now, expected timeline, and what protections we’ve put in place. Offer an immediate mitigation (traffic routing, usage credits, dedicated support) and a follow-up executive briefing and RCA.
Decision authority and escalation:
- Immediate re-prioritization authority: CTO/VP Eng or delegated engineering lead (pre-agreed in call).
- If estimated TTF > 48 hours: escalate to product exec and account exec to negotiate customer SLA terms/compensation and set executive touchpoint.
Outcome & follow-up:
- If fixed within 48 hours: confirm with customer, close incident, deliver RCA and roadmap changes, and restore paused roadmap work with documented trade-offs.
- If breach occurs: execute compensation, deliver RCA and remediation plan, and adjust KPIs to prevent recurrence.
This plan balances speed, clear authority, focused engineering effort, transparent customer communication, and concrete mitigations to minimize impact and preserve the strategic relationship.