What problem does it solve?
This Skill helps you move from machine-focused monitoring to user-focused reliability by defining what healthy really means for a service. It replaces noisy, low-signal alerts with meaningful SLIs, SLOs, error budgets, dashboards, and runbooks tied to customer impact.
Core Features & Use Cases
- Monitoring inventory and gap analysis: Review existing dashboards, alerts, runbooks, and incident pain points to find missing signals and vanity metrics.
- Critical journey and SLO design: Identify the most important user journeys, choose the right SLIs, set targets, and document error-budget policy and ownership.
- Burn-rate alerting and operational readiness: Design fast-burn and slow-burn alerts, validate dashboards and runbooks, and ensure on-call responders can act from a fresh context.
Quick Start
Help me design or review the SLOs, alerts, dashboards, and runbooks for this service so we can focus on user-impacting reliability rather than noisy infrastructure metrics.