What problem does it solve?
This Skill helps you ask plain-English questions about a Kubernetes cluster and get actionable answers powered by HolmesGPT, so you can inspect health, logs, and metrics without manually stitching together multiple tools.
Core Features & Use Cases
- In-cluster troubleshooting: Query pods, deployments, and services in the ai namespace through HolmesGPT.
- Metrics-aware analysis: Use Prometheus-backed questions to investigate CPU, memory, latency, and error spikes.
- Operational support: Validate the LiteLLM OpenAI-compatible backend, test tool calling, and diagnose Gateway API exposure.
- Example use cases: Find unhealthy workloads, review recent failures, compare resource usage, or ask whether Grafana or other services are degraded.
Quick Start
Ask HolmesGPT a plain-English question about your cluster, such as which pods are unhealthy or which workloads are using the most CPU, and let it query the in-cluster backend and observability tools for you.