What problem does it solve?
This Skill provides comprehensive guidance for establishing and maintaining high reliability standards in production systems, ensuring minimal downtime and optimal performance.
Core Features & Use Cases
- Define SLIs/SLOs: Help teams set measurable service level objectives aligned with user expectations.
- Incident Management Procedures: Offer structured runbooks for detecting, responding to, and analyzing incidents efficiently.
- Automation & Monitoring: Assist in developing scripts and configurations for proactive alerts, automated toil reduction, and resilience testing.
- Use Case: An SRE team seeks to improve site reliability by implementing error budgets, chaos engineering experiments, and blameless postmortems across multiple services.
Quick Start
Describe your current system reliability goals, then review the provided scripts and documentation to start establishing SLIs, setting SLOs, and automating incident responses today.