What problem does it solve?
Incident response mitigates the chaos and high business impact caused by production outages and degraded services by providing a clear, repeatable process for triage, communication, mitigation, and postmortem analysis to reduce mean time to resolution (MTTR).
Core Features & Use Cases
- Severity Classification: Clear SEV1–SEV4 criteria to quickly determine impact and required urgency.
- Response Framework: Step-by-step guidance for triage, assigning an incident commander, communicating status, mitigating immediate impact, and implementing a verified resolution.
- Postmortem & Continuous Improvement: Blameless postmortem format with timeline, root cause analysis (5 whys), and tracked action items with owners and due dates.
- Use Case: For a production database outage affecting all users, use this Skill to classify as SEV1, assign roles, coordinate mitigation steps, issue customer and internal communications, and run a postmortem after resolution.
Quick Start
Report the incident with severity and scope, for example: 'We have a SEV1: production API down in us-east-1; start incident response, assign an incident commander, notify SRE and product leads, and publish a status update'.