What problem does it solve? Teams running production systems often lack measurable reliability targets, drown in operational toil, and discover weaknesses only after users are affected. This Skill provides a structured site reliability engineering approach covering SLO definition, error budget tracking, observability design, chaos engineering, and capacity planning. ## Core Features & Use Cases - SLO & Error Budget Framework: Define availability and latency SLOs with SLI formulas, measurement windows, and multi-window burn rate alerts. - Observability Design: Structure metrics, logs, and traces around the four golden signals (latency, traffic, errors, saturation) to debug incidents in minutes. - Toil Reduction & Chaos Engineering: Automate repetitive operational work and proactively inject failures to find weaknesses before users do. - Use Case: A payments team needs to know whether they can ship a risky feature. Use this Skill to check the remaining error budget for the payment-api availability SLO and get a data-driven ship-or-fix recommendation. ## Quick Start Ask the agent to define an availability SLO with burn rate alerts for your payment API service.