What problem does it solve?
This Skill helps you prevent production surprises by setting up monitoring and alerting that reliably detects failures, performance regressions, and capacity risks.
Core Features & Use Cases
- Structured logging with correlation: Implement JSON logs with request IDs so events can be traced end-to-end during incidents.
- Metrics instrumentation and dashboards: Define counters, gauges, and histograms, then visualize RED/USE views in dashboards.
- Tracing and performance readiness: Configure OpenTelemetry-style spans and use profiling/performance testing references to find bottlenecks before they become outages.
- Alerting that balances signal and noise: Create actionable Prometheus alert rules with sensible thresholds, severities, and routing.
Quick Start
Use the monitoring-expert skill to design a complete observability plan for your service, including structured logs, Prometheus metrics, tracing instrumentation, dashboards, and alert rules tailored to critical paths.