What problem does it solve?
Helps teams reduce production risk and operational toil by defining measurable SLOs/SLIs, enforcing error budgets, and running disciplined incident response so services remain available and performant.
Core Features & Use Cases
- SLO, SLI & SLA design: guidance on selecting metrics (availability, latency, correctness) and setting realistic targets.
- Error budget policy & burn-rate alerts: multi-window burn-rate thresholds and actions (freeze deploys, all-hands).
- On-call & runbooks: escalation timelines, minimum staffing, and runbook requirements for pageable alerts.
- Incident lifecycle & postmortems: detection, triage, mitigation, resolution, and prevention with assigned roles.
- Production readiness checklist: monitoring, tracing, dashboards, canaries, rollbacks, and automation to reduce toil.
- Use Case: Define a 99.9% availability SLO for a payments API, configure burn-rate alerts, create a runbook for the top pageable alert, and verify on-call escalation.
Quick Start
Use the reliability skill to define an SLO for a critical service, configure burn-rate alerts, and generate a runnable runbook for the highest-impact alert.