What problem does it solve?
This Skill helps engineering teams define measurable SLIs and SLOs, manage error budgets, reduce operational toil, and prepare robust incident response and chaos experiments to keep production services reliable and scalable.
Core Features & Use Cases
- SLO/SLI design and error budget management: templates, calculators, and policy examples for setting targets and computing remaining budget.
- Monitoring, alerting and dashboards: Prometheus rules, alerting patterns, and Grafana dashboard templates to detect and respond to violations.
- Automation and runbooks: self-healing patterns, automated runbook executors, capacity planning scripts, and chaos experiment runners to reduce manual toil.
- Use Case: Create a 99.9% availability SLO for a payments API, generate Prometheus queries and alerts, produce a runbook for fast remediation, and provide automation scripts to reduce repeat incidents.
Quick Start
Draft SLO definitions, Prometheus queries, an error budget policy, and a step-by-step runbook for the payment-api service targeting 99.9% availability.