What problem does it solve?
SRE teams often face slow, manual oncall response, missed early warning signs of outages, and inconsistent root cause analysis that leads to repeated incidents. This Skill automates the entire SRE operational workflow, from alert triage and oncall management to proactive patrol and continuous self-improvement, reducing mean time to resolution and preventing outages before they impact users.
Core Features & Use Cases
- 4 Operating Modes: Oncall (automated PagerDuty alert polling, triage, and notification), Diagnosis (multi-dimensional root cause analysis across Prometheus, Kubernetes, cloud CLIs, and logs), Patrol (proactive trend-based health checks to catch issues before they fire alerts), and Iteration (self-improvement based on incident retrospectives to boost diagnostic accuracy over time).
- PagerDuty Integration: Native support for alert querying, correlation, and (with explicit confirmation) acknowledge/resolve operations, with anti-hallucination confirmation loops to prevent accidental changes.
- Proactive Patrol: Trend analysis across 24h and 7d windows, fault tolerance verification, and resource limit checks to identify at-risk systems before they fail.
- Use Case: An SRE oncall engineer can use this Skill to automatically pull all triggered PagerDuty alerts, triage and correlate them, run parallel root cause analysis across multiple data sources, and send structured incident reports to team notification channels — all without manual intervention.
Quick Start
Use the sre-agent skill to start an oncall shift, automatically triage all current PagerDuty alerts, and run root cause analysis for any triggered incidents.