What problem does it solve?
SRE helps you prevent, detect, and recover from production issues by turning reliability goals into measurable SLIs/SLOs, actionable alerting, and repeatable incident processes.
Core Features & Use Cases
- SLO/SLI and error budget design: define service-level objectives with quantifiable SLIs, calculate error budgets, and set policies for how teams should act as budget consumption changes.
- Observability and alert engineering: implement the Four Golden Signals, create burn-rate alert rules (fast + slow), and produce runbook-ready alerts that page only when humans must act.
- Operational readiness and continuous improvement: generate production readiness review (PRR) reports, write blameless postmortems, run toil audits, and document infrastructure-as-code and chaos experiment plans with safety guardrails.
Quick Start
Ask the SRE skill to propose SLOs, SLIs, and burn-rate alerts for your checkout API in the production environment, and include a runbook outline for the top alert(s).