What problem does it solve?
Observability bridges the gap between raw telemetry and actionable incident response by helping teams detect, diagnose, and prevent production failures. It centralizes guidance for logs, metrics, traces, SLOs/SLIs/SLAs, dashboards, and alerting policies so engineers can answer critical questions during outages and reduce alert fatigue.
Core Features & Use Cases
- SLO / SLI design & error budgets: Templates and calculations for defining meaningful SLOs, converting targets into error budgets, and policies for burn-rate responses.
- Alerting & runbooks: Hierarchical alert taxonomy, burn-rate alerts, actionable runbooks with owners and escalation paths to avoid noisy pages.
- Instrumentation & dashboards: OpenTelemetry setup guidance, OTEL Collector configuration, Prometheus recording and alerting rule examples, and Grafana dashboard templates for golden signals.
- Use case: Audit a payment service to produce SLIs, a 30-day SLO policy with error budget rules, Prometheus recording rules and alerts, a Grafana overview dashboard, and an on-call runbook.
Quick Start
Ask the observability skill to audit the payment-service and generate SLIs, SLO targets with error budget policy, Prometheus recording and alerting rules, a Grafana dashboard template, and a runbook with owners and escalation steps.