What problem does it solve?
This Skill helps you install, validate, and troubleshoot the full monitoring and observability stack for an ARM64 K3s cluster so you can quickly confirm that metrics, logs, traces, and alerting are working end to end.
Core Features & Use Cases
- Stack validation: Check the health of Prometheus, Grafana, Alertmanager, Tempo, Loki, Alloy, and supporting services in the monitoring namespace.
- Observability troubleshooting: Diagnose readiness issues, missing logs, failed traces, broken HTTP routes, and Helm upgrade problems with practical cluster checks.
- Operational guidance: Verify routing, storage, datasources, and pipeline wiring for Gateway API, Loki multi-tenancy, Tempo metrics generation, and Alloy log collection.
- Use case: Use this Skill when a Grafana dashboard is empty, Loki is not receiving logs, Tempo is crashing after an upgrade, or you need a reliable health report before and after a rollout.
Quick Start
Ask for a concise health check of the monitoring stack and summarize any failing pods, Helm releases, routes, or datasources.