What problem does it solve?
This Skill removes the guesswork from service observability by turning a vague production system into a clear monitoring plan with actionable signals, dashboards, and alerts.
Core Features & Use Cases
- SLI Definition: Identifies the most important service indicators such as latency, error rate, throughput, and saturation.
- Dashboard Planning: Outlines the key panels needed to understand health, load, and failure patterns at a glance.
- Alert Design: Separates symptom-based alerts from cause-based alerts and ties them to thresholds and escalation paths.
- Use Case: A team launching a payment webhook processor can use this Skill to define p99 latency goals, page-worthy failures, and dashboards for queue depth and delivery reliability.
Quick Start
Describe your service and any SLOs, and I will produce a monitoring plan with SLIs, dashboards, alerts, log-based metrics, and on-call guidance.