What problem does it solve?
This Skill ensures production changes come with privacy-safe, correlated telemetry so operators can quickly detect user impact, diagnose root cause, and validate remediation without drowning in noisy alerts.
Core Features & Use Cases
- Production-grade signals: Design structured JSON logs, latency/traffic/error/saturation metrics, and distributed trace spans across boundaries.
- Actionable alerting: Define threshold-based, owned alerts linked to runbooks and tuned to avoid alert fatigue.
- SLI/SLO alignment: Establish user-impact SLI/SLO pairs and error budgets before release, with burn-rate alerting.
- Correlation and privacy guardrails: Enforce W3C traceparent propagation and prohibit logging sensitive data or high-cardinality metric labels.
Quick Start
Use the observability capability to design logs, metrics, traces, dashboards, and alerts for a production change so you can detect user impact within minutes and investigate by correlating evidence across systems.