What problem does it solve?
This Skill helps teams define, review, and operate service level objectives in a way that reflects real user experience instead of arbitrary uptime numbers. It turns vague reliability goals into measurable SLI definitions, error budgets, and burn-rate alerts that guide engineering decisions.
Core Features & Use Cases
- SLO design: Create a structured SLO for a service or feature with a clear owner, window, target, and policy reference.
- Error budget math: Calculate allowed downtime and generate multi-window burn-rate thresholds for fast, slow, and ticket-only escalation.
- SLO review: Audit existing definitions for common mistakes such as missing SLI definitions, overly aggressive targets, short windows, or CPU-based proxies.
- Use case: A team launching checkout can define a request-success-rate SLO, compute its monthly budget, and wire alert thresholds into rollout and rollback decisions.
Quick Start
Ask this skill to design an SLO for your service, compute the error budget, and review the definition for common reliability mistakes before you put it into use.