What problem does it solve?
This Skill helps you respond correctly during an active incident when a service is broken, degraded, or at risk right now. It prioritizes fast mitigation and stabilization over premature root-cause analysis, reducing user impact before deeper diagnosis begins.
Core Features & Use Cases
- Mitigate First: Recommends the fastest safe action such as rollback, failover, feature disablement, load shedding, or pausing risky jobs.
- Stabilize and Scope: Confirms whether impact improved, then narrows the blast radius across users, regions, services, and workflows.
- Preserve Evidence and Communicate: Captures a raw incident timeline, tracks actions and signals, and supports clear stakeholder updates throughout the event.
- Use Case: A checkout outage, payment failure, production degradation, or data-risk event is detected and you need immediate response steps and a structured handoff to RCA.
Quick Start
Use the incident-response skill to help me mitigate this live production incident, stabilize impact, and produce a raw evidence-backed timeline.