What problem does it solve?
This Skill helps engineering teams turn chaotic production incidents into structured, timely responses by providing severity classification, role-based coordination, runbooks, communication templates, and post-mortem processes so incidents are resolved faster and organizational learning is captured.
Core Features & Use Cases
- Structured Incident Command: Assign IC, comms lead, technical lead, and scribe with clear timeboxed decision steps and escalation triggers.
- Runbooks & Remediation Playbooks: Templates for detection, diagnosis, rollback, restart, scaling, and verification to reduce MTTR.
- Post-Mortem Facilitation: Blameless post-mortem templates, 5 Whys, action item tracking, and lessons learned to prevent repeats.
- SLO/SLI & On-Call Design: SLO definitions, burn rate policies, and on-call rotation designs to guide when to page and when to pause feature work.
- Use Case: Lead a SEV1 outage for a checkout API: declare severity, coordinate rollback or mitigation, communicate to stakeholders, verify SLIs, and produce a post-mortem with tracked actions.
Quick Start
Ask the agent to act as Incident Response Commander for a SEV2 outage on the checkout-api: assign roles, run diagnostics using runbook steps, propose immediate mitigations, and produce a timestamped timeline plus a post-mortem action list.