What problem does it solve?
This Skill provides a structured response when production systems break, helping teams reduce user impact quickly, investigate causes calmly, and prevent the issue from happening again.
Core Features & Use Cases
- Incident Triage: Classify severity, assign responders, and open the incident workspace.
- Mitigation and Recovery: Roll back, disable features, scale services, and communicate status before deep debugging.
- Root-Cause Analysis and Follow-Up: Document the timeline, identify contributing factors, add fixes and tests, and produce a blameless post-incident review.
- Use Case: A payment API starts failing in production and customers cannot check out; this Skill guides the team from urgent mitigation through RCA, remediation, and lessons learned.
Quick Start
Use the incident skill to triage this production outage, reduce user impact, and produce the incident artifacts for root-cause analysis and postmortem review.