What problem does it solve?
This Skill addresses the challenge of system monitoring, alerting, and reliability through the use of tools like Prometheus, Grafana, and logging practices, ensuring that you can proactively manage incidents and maintain system health.
Core Features & Use Cases
- System Monitoring: Monitor services using tools like Prometheus and Grafana.
- Alert Configuration: Set up alerts to notify you before potential issues arise.
- SLI/SLO Definition: Define Service Level Indicators and Objectives to measure service performance.
- Incident Response: Provide runbooks for common incidents and facilitate incident response.
- Use Case: If you need to monitor a service, configure alerts, define SLOs, or respond to an incident, this skill can assist you in achieving those goals.
Quick Start
Use the infra-sre skill to set up monitoring for a new service in your infrastructure.