What problem does it solve?
This skill solves the chaos of production outages by providing a structured, mitigation-first playbook that ensures incidents are handled safely, logged accurately, and followed by blameless postmortems.
Core Features & Use Cases
- Mitigation-First Workflow: Prioritizes safe paths back to green (reverts, flag flips) before diagnosing root causes.
- Automated Timeline Logging: Automatically timestamps every action and decision into a dedicated incident directory.
- Postmortem Generation: Converts the incident timeline into a structured postmortem report that feeds directly into Learned Rules.
- Use Case: When a production service experiences a spike in error rates, use this skill to triage the severity, draft client communications, and track mitigation steps without losing focus on the incident clock.
Quick Start
Trigger the incident response workflow by stating the current issue such as production is down and we need to start the triage process.