What problem does it solve?
This Skill helps teams respond when production is down or degraded by coordinating triage, diagnosis, fastest safe recovery, verification, and documentation so service disruption ends quickly and root causes are fixed.
Core Features & Use Cases
- Rapid triage and severity classification: Quickly assess symptoms, blast radius, start time, and recent changes to label the incident as SEV1–SEV4.
- Guided diagnosis and fastest safe remediation: Investigate logs, processes, resources, dependencies, and configuration to identify the root cause or best hypothesis, then apply an immediate recovery fix such as rollback, restore, restart, or failover.
- Recovery verification and permanent follow-up: Validate health and core functionality via monitoring, then escalate to a permanent fix plan using an architect when a bandaid was applied.
Quick Start
Use the incident skill when you see a production outage or major degradation and need an end-to-end response from triage through verification and incident documentation.