incident-response

Manage production incidents from detection through blameless postmortem documentation.

145|36|Updated Feb 26, 2026
One-click install
npx skills add https://github.com/w95/awesome-claude-corporate-skills --skill incident-response-w95
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/w95/awesome-claude-corporate-skills/tree/main/08-it-engineering/incident-response
Command: npx skills add https://github.com/w95/awesome-claude-corporate-skills --skill incident-response-w95

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill provides a structured framework to efficiently manage and resolve production incidents, minimizing downtime and impact.

Core Features & Use Cases

  • Severity Classification: Helps categorize incidents based on impact and urgency.
  • Response Framework: Guides users through the key stages of incident management: Triage, Communication, Mitigation, Resolution, and Postmortem.
  • Communication Templates: Offers guidance on crafting clear and consistent status updates.
  • Postmortem Format: Outlines a blameless approach to reviewing incidents and identifying action items.
  • Use Case: When a critical service goes down, this Skill can be invoked to ensure a rapid, organized, and effective response, from initial alert to final resolution and learning.

Quick Start

Initiate incident response by stating "we have an incident".

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage production incidents from detection to postmortem?

Production incidents are managed through a structured framework covering triage, communication, mitigation, resolution, and blameless postmortem analysis. This minimizes downtime by ensuring a rapid, organized response from the initial alert through final resolution and learning.

What is the best way to classify the severity of a system outage?

Classifying system outage severity involves categorizing incidents based on their impact and urgency. Proper severity classification helps prioritize response efforts and dictates the communication strategies and mitigation steps required to resolve the production issue.

How do I write a blameless postmortem after resolving a critical service outage?

Writing a blameless postmortem involves reviewing the incident timeline and identifying action items without attributing fault. It provides a structured format to document what happened, why, and how to prevent similar production incidents in the future.

How should I structure communication during an active production incident?

Incident communication should use clear, consistent templates to provide status updates throughout the outage management lifecycle. Structured communication strategies ensure stakeholders remain informed during triage, mitigation, and resolution of critical system degradations.

Can I use this incident response framework for performance degradations rather than just full outages?

Yes, the incident response framework facilitates structured management for both critical system outages and performance degradations. It supports the full lifecycle from severity classification through mitigation and postmortem analysis for any production issue.

When do I need a structured incident response process?

A structured incident response process is needed when a critical service goes down or degrades. It ensures efficient outage management, guiding teams through triage, communication, mitigation, and blameless postmortem documentation to minimize impact.