incident-diagnosis

Automate diagnostic procedures for Kubernetes incidents and alerts.

91|16|Updated May 11, 2026
One-click install
npx skills add https://github.com/OpsinTech/opsintech-platform --skill incident-diagnosis
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-diagnosis
Source: https://github.com/OpsinTech/opsintech-platform/tree/main/skills/public/incident-diagnosis
Command: npx skills add https://github.com/OpsinTech/opsintech-platform --skill incident-diagnosis

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

The Incident Diagnosis Skill tackles the complex and time-consuming task of diagnosing system incidents, production alerts, and Kubernetes crashes, making it easier to quickly identify and resolve issues.

Core Features & Use Cases

  • Structured Diagnostic Workflows: Follow a defined process for incident resolution with step-by-step checklists.
  • SRE Expertise: Operates like a senior Site Reliability Engineer to provide detailed root cause analysis.
  • Customizable SOPs: Tailor workflows to specific alert types or service architectures.
  • Professional Reports: Generates detailed incident reports in a standardized format for auditing and documentation.
  • Use Case: Use the skill to troubleshoot a sudden CPU spike in your service. It will guide you through diagnostic steps, retrieve and analyze logs, and produce a professional report.

Quick Start

Invoke the 'incident-diagnosis' skill when a new system incident occurs.

Frequently Asked Questions about incident-diagnosis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I diagnose Kubernetes production incidents and alerts efficiently?

To diagnose Kubernetes production incidents, you can use a structured diagnostic workflow that guides you through step-by-step checklists, retrieves logs, and performs root cause analysis to resolve system alerts quickly.

What is the best way to troubleshoot a sudden CPU spike in a Kubernetes service?

The best way to troubleshoot a sudden CPU spike is to follow an SRE-grade diagnostic workflow that retrieves and analyzes service logs to identify the root cause and generate a professional incident report.

Can I customize incident resolution workflows for specific alert types?

Yes, you can customize incident resolution workflows by tailoring standard operating procedures to match specific alert types or service architectures, ensuring accurate system diagnostics for your environment.

Do I need scripting capabilities to run system diagnostics in Kubernetes?

Yes, scripting capabilities are required to automate diagnostic procedures and execute incident resolution workflows within a controlled operational environment for Kubernetes production support.

How does SRE expertise help with root cause analysis for system crashes?

SRE expertise applies a defined diagnostic process to system crashes, performing detailed root cause analysis and generating standardized incident reports for auditing and documentation.