🚨

Enterprise Operations & Incident Management Topics

Large-scale operational practices for enterprise systems including major incident response, crisis leadership, enterprise-scale troubleshooting, business continuity planning, and recovery. Covers coordination across teams during high-severity incidents, forensic investigation, decision-making under pressure, post-incident processes, and resilience architecture. Distinct from Security & Compliance in its focus on operational coordination and recovery rather than preventive security.

Incident Communication and Stakeholder Management

Communicating during and after an incident to internal stakeholders, executives, customers, partners, and regulators. Covers status-update content and cadence, status page and customer notifications, war-room and channel choices, translating technical state into business impact, and managing expectations under uncertainty. Also covers coordinating messages with legal and PR before disclosure, handling data exposure, vendor-caused and press-visible incidents, customer-facing post-incident summaries, and building the process: comms roles, pre-approved messaging, governance, drills and metrics for incident communication.

62 questions

Incident Severity Classification and Escalation

Assigning severity to incidents, deciding when and whom to page, and moving an issue up defined escalation paths. Covers severity matrices, escalation criteria, triage decisions, and scope-of-authority calls about when to pull in more responders or leadership. The routing and prioritization discipline that sits at the front of the incident lifecycle.

0 questions

Operational Risk Management

Identifying, assessing, and reducing operational risk before it becomes an incident. Covers operational risk categories (process, people, supplier, technology, execution), surfacing the risks in a large program such as a cloud migration, risk registers and ownership cadence, likelihood and impact scoring (heat maps, qualitative vs quantitative, expected loss, ranges and Monte Carlo, estimating with little history), scenario analysis and structured failure-mode review before a risky change, key risk indicators and early-warning signals, risk response strategies (avoid, reduce, transfer, accept) including contracts and insurance for supplier exposure, prioritizing mitigations by expected loss and cost per unit of risk reduced, residual risk reporting and escalation to leadership, key-person risk and single points of failure, risk appetite and risk-versus-speed trade-offs, systemic and recurring risk including human error, the three lines model (formerly three lines of defense), and building organizational resilience (resilience metrics, resilience testing programs, culture). Proactive risk reduction, not reactive incident handling. Disaster recovery and continuity planning, incident command, vendor due diligence, security and privacy risk, and project schedule risk are covered elsewhere.

6 questions

Disaster Recovery and Business Continuity

Planning and executing recovery from large-scale outages and disasters. Covers recovery objectives (RTO/RPO), failover and backup strategy, DR runbook automation, and business-continuity planning that keeps critical operations running through major disruptions. The resilience-and-recovery planning discipline for worst-case events.

0 questions

Postmortems, Root Cause Analysis, and Blameless Culture

Investigating what caused an incident and turning the lessons into lasting improvement. Covers root-cause techniques (five whys, causal chains, contributing-factor analysis), writing postmortem documents, and tracking follow-up action items to prevent recurrence, as well as facilitating those reviews blamelessly: building psychological safety, treating failures as learning opportunities rather than occasions for blame, and driving continuous-improvement loops across teams. The structured after-the-fact analysis discipline together with the organizational culture that makes it effective.

32 questions

On-Call Practices and Runbook Design

Running a sustainable on-call function: rotation design, production-readiness handoffs, and authoring runbooks that let responders act quickly. Covers runbook automation, on-call culture, escalation-ready documentation, and readiness reviews before a service takes production traffic. The operational-preparedness discipline that makes incidents survivable.

41 questions

Monitoring and Observability

Instrumenting systems so their internal state is visible: metrics, logs, and traces, dashboards, and operational health signals. Covers what to measure, how to build observability into services, and how to use telemetry to detect and diagnose problems. Distinct from alerting in that it focuses on visibility and instrumentation rather than notification.

0 questions

Incident Response and Management

The end-to-end operational lifecycle of a production incident: detection, triage, mitigation, resolution, and handoff. Covers response coordination, ownership of an active incident, and the mechanics of restoring service quickly. This is the generalist backbone topic that most operational roles are drilled on.

11 questions

Incident Command and Crisis Leadership

Leading people and decisions through high-severity operational events. Covers acting as incident commander (assigning roles like comms lead, ops lead, and scribe, driving the response tempo, and making authority calls while others execute) as well as the broader crisis-leadership dimension: making consequential decisions with incomplete information, rapid replanning as conditions change, and staying effective under acute time pressure including volume spikes and ambiguity. Focused on the command, coordination, and decision-quality layer rather than hands-on debugging.

0 questions