What problem does it solve?
This Skill provides expert guidance and tools to implement and manage Site Reliability Engineering (SRE) best practices, ensuring high availability, reliability, and performance of systems.
Core Features & Use Cases
- SLO/SLI Management: Define, track, and calculate compliance for Service Level Objectives and Indicators.
- Incident Management: Create, update, and report on incidents, including MTTR calculation.
- Monitoring & Alerting: Implement best practices for the four golden signals and define alert rules.
- Chaos Engineering: Design and run experiments to proactively identify system weaknesses.
- Use Case: A team can use this Skill to define SLOs for their API, track them against real-time metrics, and manage any incidents that arise, ensuring they meet their reliability targets.
Quick Start
Use the sre-expert skill to define standard SLOs for a web service.