What problem does it solve?
This Skill provides a comprehensive framework to manage production incidents effectively, from initial detection and classification to resolution and post-mortem analysis, ensuring minimal disruption and rapid recovery.
Core Features & Use Cases
- Incident Classification: Quickly categorize incidents by severity (SEV1-SEV4) based on impact and user affected.
- Runbook Guidance: Step-by-step instructions for triage, mitigation, and resolution tailored for various platforms (Firebase, GCP, Mobile, Web).
- Communication Templates: Pre-written messages for internal stakeholders and external customers during and after an incident.
- Post-Mortem Framework: A structured template to conduct blameless post-mortems, identify root causes, and define preventative action items.
- Use Case: During a sudden spike in 5xx errors on your web application, use this Skill to classify the incident, follow the runbook to identify the cause (e.g., a recent deployment), mitigate the issue by rolling back, and then use the post-mortem template to document the event and create action items to prevent recurrence.
Quick Start
Use the incident-response skill to classify an incident with the following details: service is down, affecting 75% of users, and revenue is impacted.