DevOps Troubleshooter

Diagnose and resolve DevOps failures across CI/CD, infrastructure, runtime, and deployment pipelines.

1|12|Updated Feb 23, 2026
One-click install
npx skills add https://github.com/ChatAndBuild/chatchat-skills --skill devops-troubleshooter-chatandbuild
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: DevOps Troubleshooter
Source: https://github.com/ChatAndBuild/chatchat-skills/tree/main/skills/devops-troubleshooter
Command: npx skills add https://github.com/ChatAndBuild/chatchat-skills --skill devops-troubleshooter-chatandbuild

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill helps troubleshoot DevOps failures across CI/CD, infrastructure, runtime services, and deployment pipelines, providing evidence-driven incident triage and remediation steps.

Core Features & Use Cases

  • Evidence-Driven Triage: Provides systematic elimination, network partition testing, and resource contention analysis for diagnosing failures.
  • Domain-Specific Guides: Offers detailed guides for CI failures, deployment failures, runtime crashes, and networking issues.
  • Troubleshooting Workflow: Outlines a structured approach to troubleshooting, including capturing incident scope, verifying recent changes, and applying fixes with validation.

Quick Start

Use the DevOps Troubleshooter skill to diagnose a failure in your CI/CD pipeline by following the systematic elimination process.

Frequently Asked Questions about DevOps Troubleshooter

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I troubleshoot a failing CI/CD pipeline or deployment failure?

To troubleshoot a failing CI/CD pipeline, use evidence-driven incident triage to systematically eliminate potential causes. Capture the incident scope, verify recent configuration changes, and apply targeted remediation steps to validate the fix.

What is the best way to diagnose runtime crashes and infrastructure failures?

Diagnosing runtime crashes and infrastructure failures requires an evidence-driven troubleshooting workflow. Analyze system logs and resource contention systematically to identify the root cause, then apply validated remediation steps to restore runtime services.

How do I resolve networking issues and connectivity problems in my infrastructure?

Resolve networking issues by conducting network partition testing and applying domain-specific guides. Analyze network resources and system logs to isolate the connectivity failure, then apply remediation steps with validation to ensure resolution.

Do I need access to system logs and configuration files to use this incident management workflow?

Yes, resolving DevOps issues with this incident triage workflow requires access to system logs, configuration files, and network resources. This evidence-driven approach relies on verifying recent changes and analyzing these inputs to apply accurate fixes.

Why does my DevOps troubleshooting process fail to identify the root cause of incidents?

DevOps troubleshooting fails when it lacks a structured incident triage workflow. Capturing the incident scope, verifying recent changes, and systematically eliminating causes using domain-specific guides ensures accurate root cause identification and successful remediation.

Can I use this incident triage process for resource contention analysis across my deployment pipelines?

Yes, this incident triage process supports resource contention analysis across deployment pipelines. It systematically captures the incident scope and applies evidence-driven elimination to diagnose and resolve infrastructure resource conflicts.