incident-analysis

Analyze production incidents using GCP Observability and gcloud CLI playbooks.

6|Updated Feb 13, 2026
One-click install
npx skills add https://github.com/damianpapadopoulos/auto-claude-skills --skill incident-analysis-damianpapadopoulos
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-analysis
Source: https://github.com/damianpapadopoulos/auto-claude-skills/tree/main/skills/incident-analysis
Command: npx skills add https://github.com/damianpapadopoulos/auto-claude-skills --skill incident-analysis-damianpapadopoulos

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires gcloud, google-cloud-observability-mcp, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill provides a structured approach to analyze and mitigate production incidents, helping you quickly identify root causes and apply the appropriate fixes.

Core Features & Use Cases

  • Tiered GCP Log Investigation: Utilizes playbooks-driven mitigation and structured validation for efficient incident analysis.
  • Staged Process: Guides through stages like intake, mitigation, classification, investigation, execution, validation, post-mortem, and optional reporting.
  • Integration with Observability Tools: Integrates with GCP Observability and gcloud CLI for detailed insights into service performance and resource utilization.
  • Investigation Modes: Offers default full investigation and opt-in live triage for active incidents.
  • Comprehensive Analysis: Covers various aspects including resource saturation, infrastructure, application layer, and dependency failures.

Quick Start

Use the incident-analysis skill to investigate a production incident and apply the appropriate playbook based on the symptoms observed.

Frequently Asked Questions about incident-analysis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I perform root cause analysis for a production incident in GCP?

Root cause analysis for a production incident in GCP is performed using a structured playbook that guides you through intake, mitigation, classification, investigation, and validation stages. It leverages GCP Observability and gcloud CLI to analyze logs, resource utilization, and dependency failures.

Does this incident analysis approach work with gcloud CLI and Google Cloud Observability tools?

Yes, this incident analysis approach directly integrates with gcloud CLI and Google Cloud Observability tools to gather detailed insights into service performance, resource saturation, and infrastructure issues during an active incident.

What's the best way to investigate application layer and dependency failures during an active incident?

The best way to investigate application layer and dependency failures is using an opt-in live triage mode, which applies playbook-driven mitigation to analyze production logs and resource saturation quickly during active incidents.

What stages are involved in playbook driven mitigation for production monitoring?

Playbook driven mitigation for production monitoring involves seven structured stages: intake, mitigation, classification, investigation, execution, validation, and post-mortem, with an optional reporting stage to finalize the analysis.

Do I need gcloud CLI access to analyze resource saturation and infrastructure incidents?

Yes, you need gcloud CLI access along with GCP Observability tools to effectively analyze resource saturation, infrastructure issues, and dependency failures when investigating production incidents using the structured playbooks.