What problem does it solve?
This Skill unit provides a comprehensive set of tools and best practices for Site Reliability Engineers (SREs) to manage production systems more effectively, reduce toil, and maintain high service reliability.
Core Features & Use Cases
- Service Level Objective (SLO) Management: Define and track SLOs for availability, latency, and error budgets.
- Error Budget Policies: Calculate and manage error budgets to plan for and respond to system failures.
- Monitoring and Alerting: Implement golden signals monitoring and configure alerting based on SLOs.
- Automation: Automate repetitive tasks and toil reduction with scripts and tools.
- Chaos Engineering: Design and execute chaos experiments to test system resilience and recovery.
- Incident Response: Develop incident response procedures, runbooks, and postmortems for improved MTTR.
Quick Start
Load the sre-engineer skill to start defining SLOs, creating error budget policies, and automating incident response procedures.