What problem does it solve?
Builds a coherent monitoring and observability strategy so teams can detect, diagnose, and resolve production issues before customers are impacted. It replaces ad hoc or missing instrumentation with structured logs, meaningful metrics, distributed tracing, actionable alerts, and runbooks that fit the team's stage and budget.
Core Features & Use Cases
- Three pillars coverage: guidance for logs, metrics, and traces and how to instrument each layer.
- Alerting and SLOs: pragmatic alert thresholds, severity levels, and error budgets tailored to SaaS products.
- Operational practices: health checks, on-call rotations, incident response steps, status page recommendations, and synthetic monitoring.
- Use Case: For a startup migrating from no observability, provide a minimal-cost stack (Sentry + Better Stack), key metrics and alerts, a health endpoint plan, and an on-call/runbook checklist to reduce downtime.
Quick Start
Ask the monitoring skill to audit your current stack, recommend a logs/metrics/traces setup, propose 5 critical alerts with runbooks, and define SLO targets given your tech stack and team size.