What problem does it solve?
GKE observability setup is complex and error-prone, with default configurations missing critical control plane metrics, leading to blind spots in cluster health, troubleshooting delays, and unplanned downtime for production workloads.
Core Features & Use Cases
- Golden Path Observability Configuration: Provides pre-defined, production-ready defaults for Cloud Logging, Cloud Monitoring, and managed Prometheus, including critical control plane metrics for API server, scheduler, and controller manager that are not enabled by default.
- Production Best Practices: Includes guidance for setting up actionable alerts, optimizing monitoring costs for non-production environments, and recommendations for distributed tracing and continuous profiling for microservice architectures.
- Use Case: A DevOps engineer managing a production GKE cluster can use this skill to enable full observability, set up alerts for pod crash loops and high API server latency, and reduce monitoring costs for development clusters by switching to system-only monitoring.
Quick Start
Use the gke-observability skill to configure full observability for your GKE cluster including Cloud Logging, managed Prometheus, and control plane monitoring metrics.