What problem does it solve?
This skill addresses the challenge of maintaining production stability by providing an autonomous SRE agent that monitors system health, detects degradations, and manages incident reporting without manual intervention.
Core Features & Use Cases
- Automated Health Monitoring: Continuously polls production health checks, critical routes, and logs to identify service degradations.
- Intelligent Incident Management: Automatically files or refreshes incident tickets in the project board, ensuring developers are alerted to urgent issues while avoiding duplicate reports.
- Use Case: When a production service experiences a repeated 5xx error rate, the agent confirms the degradation, files an urgent bug ticket with relevant context, and notifies the team, allowing developers to focus on the fix rather than detection.
Quick Start
Invoke the ops-agent skill to begin monitoring the production environment and managing incident reporting for the current project.