What problem does it solve?
When major outages occur, teams need a coordinated, blame-free process to triage, communicate, contain, and learn from incidents. This skill guides responders through dual modes (respond and postmortem) and provides a structured workflow to minimize downtime and accelerate remediation, while ensuring a thorough post-incident analysis.
Core Features & Use Cases
- Real-time incident coordination: assigns roles, collects essential incident details, and tracks actions across SRE, DevOps, and communications.
- Dual-mode operation: respond mode for active incidents and postmortem mode for blameless root-cause analysis after stability is restored.
- Artifacts & runbooks: generates incident body, timelines, contributing factors, remediation plans, and postmortem documentation to be stored under docs/incidents/{incident-id}.md.
- Outputs actionable guidance: escalation recommendations, SLO status checks, and structured incident reports for stakeholders.
Quick Start
Describe the incident details and initiate the respond workflow.