holmesgpt-skill

Investigate Kubernetes and cloud-native infrastructure issues using live observability data.

7|1|Updated Jul 12, 2026
One-click install
npx skills add https://github.com/julianobarbosa/claude-code-skills --skill holmesgpt-skill
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: holmesgpt-skill
Source: https://github.com/julianobarbosa/claude-code-skills/tree/main/skills/holmesgpt-skill
Command: npx skills add https://github.com/julianobarbosa/claude-code-skills --skill holmesgpt-skill

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

HolmesGPT provides AI-powered troubleshooting across cloud-native ecosystems, integrating with observability data to perform root-cause analysis while operating with read-only RBAC.

Core Features & Use Cases

  • Root Cause Analysis: Investigates alerts across Kubernetes, Prometheus, and incident systems.
  • Multi-Source Integrations: 30+ toolsets for Kubernetes, Grafana, Loki, Tempo, etc.
  • Alert Integration: Integrates with AlertManager, PagerDuty, OpsGenie, Jira, Slack.
  • Use Case: Automatically investigate a paged incident and summarize root causes.

Quick Start

Install via Helm and connect providers; start an investigation with holmesgpt.

Frequently Asked Questions about holmesgpt-skill

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I investigate Kubernetes alerts automatically with AI?

AI-powered troubleshooting investigates Kubernetes alerts by analyzing live observability data from Prometheus, AlertManager, and PagerDuty to identify root causes and suggest remediation steps without manual log analysis.

Can I use AI troubleshooting with Prometheus and Grafana together?

Yes, multi-source integrations support simultaneous connections to Kubernetes, Prometheus, Grafana, Loki, Tempo, and DataDog, enabling root-cause analysis across your entire observability stack in a single investigation.

Does cloud-native troubleshooting work with read-only RBAC permissions?

Root-cause analysis operates under read-only RBAC constraints, allowing secure investigations without elevated cluster access or write permissions to your infrastructure.

How do I set up AI incident investigation with PagerDuty and Slack?

Connect via Helm or CLI installation, configure your AI provider (Anthropic, OpenAI, Azure, AWS Bedrock, Google Gemini, or Vertex AI), and integrate alert systems; investigations automatically trigger from PagerDuty and post summaries to Slack.

What observability platforms does cloud-native root-cause analysis support?

30+ toolsets integrate with Kubernetes, Prometheus, Grafana, Loki, Tempo, DataDog, AlertManager, OpsGenie, and Jira, supporting alert-driven investigations across heterogeneous cloud-native environments.

Can I customize troubleshooting runbooks and toolsets for my infrastructure?

YAML configuration allows custom toolset definitions and runbook templates, enabling tailored investigation workflows that align with your team's procedures and incident response patterns.