incident-management

Coordinate production incident response through detection, triage, investigation, mitigation, resolution, and post-mortem analysis.

111|18|Updated Dec 17, 2025
One-click install
npx skills add https://github.com/dralgorhythm/claude-agentic-framework --skill incident-management
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-management
Source: https://github.com/dralgorhythm/claude-agentic-framework/tree/main/.claude/skills/operations/incident-management
Command: npx skills add https://github.com/dralgorhythm/claude-agentic-framework --skill incident-management

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Handles production incidents with triage, communications, and post-mortems for reliability.

Core Features & Use Cases

  • Incident Response: Detect, triage, investigate, mitigate, resolve, learn.
  • Post-Mortem: Structured reports and actionable items.
  • Use Case: Manage a Sev-1 outage and drive rapid restoration.

Quick Start

Initiate an incident response with a brief post-mortem plan.

Frequently Asked Questions about incident-management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response for production outages?

Incident response coordination structures detection, triage, investigation, mitigation, and resolution across a defined lifecycle. This Skill applies a blameless framework to outage management, moving from initial alert through post-mortem analysis to capture lessons and improve reliability.

What's the best way to conduct a blameless post-mortem after an outage?

Blameless post-mortems focus on systemic factors rather than individual blame. This Skill uses a standardized template to document incident timelines, root causes, and actionable items, creating institutional knowledge that prevents recurrence across your production services.

How do I triage and investigate production incidents systematically?

Systematic triage prioritizes incidents by severity and impact, then investigation follows structured steps to identify root causes. This Skill guides you through detect, investigate, and mitigation phases while maintaining clear communication and documentation for faster resolution.

Can I use incident management for Sev-1 outages?

Yes. This Skill handles high-severity production outages by applying the full incident lifecycle—from rapid detection and triage through coordinated mitigation and resolution—then capturing post-mortem insights to strengthen reliability.

What does a structured incident response lifecycle look like?

A structured lifecycle moves through Detect, Triage, Investigate, Mitigate, Resolve, and Learn phases. This progression ensures incidents are caught early, prioritized correctly, diagnosed thoroughly, fixed promptly, and analyzed post-incident to prevent future occurrences.

How do I maintain reliability across production services during incidents?

Reliability improves through coordinated incident response and post-mortem analysis. This Skill embeds learning into your process, converting each incident into actionable improvements that reduce future outage frequency and duration across your service infrastructure.