What problem does it solve?
This Skill provides comprehensive observability into your Kubernetes and OpenShift clusters, enabling proactive monitoring, rapid issue diagnosis, and efficient incident response.
Core Features & Use Cases
- Metrics & Logging: Query Prometheus/PromQL, Thanos, Loki, and ELK for deep insights into cluster performance and application behavior.
- Alert Management: Triage, tune, and manage alerts to reduce noise and ensure timely action.
- Incident Response: Follow playbooks for rapid detection, diagnosis, and resolution of incidents.
- SLO/SLI Tracking: Monitor Service Level Objectives and Indicators to ensure service health and reliability.
- Cloud-Specific Monitoring: Integrates with Azure Monitor (ARO) and AWS CloudWatch (ROSA).
- Use Case: When an alert fires indicating high error rates, use this Skill to immediately query metrics and logs, identify the root cause, and initiate a resolution process.
Quick Start
Analyze the current Prometheus metrics for high error rates across all services.