What problem does it solve?
This Skill helps you define, implement, and maintain highly reliable production systems by focusing on Service Level Objectives (SLOs), error budgets, monitoring, and automation.
Core Features & Use Cases
- SLO/SLI Definition: Define clear, measurable Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to track system reliability.
- Error Budget Management: Implement policies for managing error budgets, guiding tradeoffs between reliability and feature velocity.
- Monitoring & Alerting: Configure robust monitoring for golden signals (latency, traffic, errors, saturation) and set up actionable alerts.
- Automation & Toil Reduction: Identify and automate repetitive operational tasks to reduce toil and improve efficiency.
- Incident Management & Chaos Engineering: Develop incident response procedures and proactively test system resilience through chaos engineering.
- Use Case: A team wants to ensure their e-commerce checkout service is available 99.95% of the time. This Skill can help them define the SLIs for availability, set up monitoring, create alerts for when the error budget is being consumed too quickly, and automate common remediation tasks.
Quick Start
Use the sre-engineer skill to define SLOs for the payment service, focusing on latency and availability.