What problem does it solve?
This Skill provides a practical framework to measure and improve service reliability by defining SLIs and SLOs, designing actionable alerts and dashboards, and tracking error budgets so teams can prioritize reliability work and reduce unplanned downtime.
Core Features & Use Cases
- SLI & SLO definitions: Standard formulas for availability, latency percentiles, error rates, throughput, and freshness with guidance on rolling windows and targets.
- Error budget policy & burn rate: Clear error budget states, burn rate thresholds, and recommended actions for green/yellow/red/exhausted states to drive decision-making.
- Alerting, dashboards & capacity planning: Principles for actionable alerts, USE/RED dashboard patterns, deployment annotations, and a capacity planning template for forecasting and scaling.
- Use Case: Create a 30-day SLO for an API, configure burn-rate alerts, add SLI panels to service dashboards, and run monthly error budget reviews.
Quick Start
Draft SLIs and a 30-day SLO for the payments service focusing on availability and P95 latency.