What problem does it solve?
This Skill provides a structured, blameless, multi-phase incident response process that reduces downtime and guides engineering teams through detection, hypothesis-driven investigation, mitigation, resolution, and postmortem creation.
Core Features & Use Cases
- Five-phase workflow: Detect, Investigate, Mitigate, Resolve, Postmortem with decision gates and severity classification.
- Hypothesis-driven diagnostics: Generate and prioritize multiple hypotheses, run targeted diagnostics, and converge on a root cause with evidence.
- Audit-ready artifacts: Create standardized incident records and postmortems in the local docs/ directory and optionally log a one-line summary to cloud memory via MCP.
- Use Case: For a P0 production outage, classify severity, consult PRODUCTS.md for tech stack context, run prioritized tests, apply mitigations, restore service, and produce a postmortem within 48 hours.
Quick Start
Start a P0 incident titled "Site outage" and run the incident commander to detect the issue, generate hypotheses, apply mitigation steps, and create a postmortem.