What problem does it solve?
Incident investigation is often a chaotic, manual process involving multiple tools and a high risk of missing critical information. This skill provides a structured, automated workflow for incident response, guiding users through evidence gathering, root cause analysis, and documentation, significantly reducing resolution time and improving post-incident learning.
Core Features & Use Cases
- Guided Investigation: Follow alert-specific runbooks and systematic steps to diagnose issues.
- Automated Evidence Collection: Gather logs, metrics, events, and application status with pre-defined commands.
- Structured RCA Documentation: Generate comprehensive incident reports in
TODO_INCIDENTS.md for future learning.
- Use Case: An alert for "Prometheus Down" has fired. Use this skill to automatically follow the Prometheus Down runbook, gather relevant logs and metrics, and start documenting the incident for root cause analysis.
Quick Start
Use the investigate-incident skill to start an investigation for the "Prometheus Down" alert, following its runbook and gathering initial evidence.