observability

Analyzes Prometheus/Grafana metrics and logs for Kubernetes and OpenShift incident response.

4|1|Updated Oct 27, 2025
One-click install
npx skills add https://github.com/kcns008/cluster-code --skill observability-kcns008
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability
Source: https://github.com/kcns008/cluster-code/tree/main/.claude/skills/observability
Command: npx skills add https://github.com/kcns008/cluster-code --skill observability-kcns008

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides comprehensive observability into your Kubernetes and OpenShift clusters, enabling proactive monitoring, rapid issue diagnosis, and efficient incident response.

Core Features & Use Cases

  • Metrics & Logging: Query Prometheus/PromQL, Thanos, Loki, and ELK for deep insights into cluster performance and application behavior.
  • Alert Management: Triage, tune, and manage alerts to reduce noise and ensure timely action.
  • Incident Response: Follow playbooks for rapid detection, diagnosis, and resolution of incidents.
  • SLO/SLI Tracking: Monitor Service Level Objectives and Indicators to ensure service health and reliability.
  • Cloud-Specific Monitoring: Integrates with Azure Monitor (ARO) and AWS CloudWatch (ROSA).
  • Use Case: When an alert fires indicating high error rates, use this Skill to immediately query metrics and logs, identify the root cause, and initiate a resolution process.

Quick Start

Analyze the current Prometheus metrics for high error rates across all services.

Frequently Asked Questions about observability

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I query Prometheus metrics and Loki logs to diagnose high error rates in Kubernetes?

Query Prometheus metrics and Loki logs to diagnose high error rates in Kubernetes by using PromQL and log aggregation to identify root causes and initiate rapid incident resolution processes.

What is the best way to triage and tune alerts in an OpenShift cluster?

Triage and tune alerts in an OpenShift cluster by managing alert noise and ensuring timely action through integrated incident response playbooks that facilitate rapid detection, diagnosis, and resolution.

Does this observability workflow support multi-cloud environments like AWS EKS and Azure AKS?

Yes, the observability workflow supports multi-cloud environments including AWS EKS, Azure AKS/ARO, and GCP GKE by integrating with cloud-specific monitoring services like AWS CloudWatch and Azure Monitor alongside Thanos and Grafana.

How do I track SLO and SLI compliance for services running on Kubernetes?

Track SLO and SLI compliance for Kubernetes services by monitoring Service Level Objectives and Indicators to ensure service health, executing automated data collection for SLO compliance checks and incident reports.

Can I use Grafana and Thanos together for metrics analysis in OpenShift?

Yes, you can use Grafana and Thanos together for metrics analysis in OpenShift to gain deep insights into cluster performance and application behavior through comprehensive observability and log aggregation.