incident-response

Guide production incident response from detection to resolution with predefined roles and runbooks.

523|69|Updated Apr 28, 2026
One-click install
npx skills add https://github.com/rampstackco/claude-skills --skill incident-response-rampstackco
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/rampstackco/claude-skills/tree/main/skills/incident-response
Command: npx skills add https://github.com/rampstackco/claude-skills --skill incident-response-rampstackco

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Active production incidents disrupt service and erode trust. This skill provides a clear, repeatable framework with predefined roles, workflows, and templates to guide detection, triage, mitigation, communication, and resolution.

Core Features & Use Cases

  • Predefined incident roles (Incident Commander, Communications Lead, Operations Lead, Scribe, SMEs) to coordinate response.
  • Severity-guided decision framework and runbooks to determine escalation, actions, and communications.
  • End-to-end workflow covering detection, triage, mitigation, verification, and post-incident AAR.
  • Status-page and internal communication templates to keep stakeholders informed.
  • Reference playbooks and checklists to standardize responses and enable learning.

Quick Start

Describe an active incident to the skill and it will guide your team through detection, triage, mitigation, and resolution steps.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage an active production incident from detection to resolution?

Incident response is managed by assigning predefined roles like Incident Commander and Scribe, then following a severity-guided workflow for triage, mitigation, and verification. This framework standardizes the entire process from detection to resolution.

What roles do I need to define for effective incident management?

Effective incident management requires clear roles including Incident Commander, Communications Lead, Operations Lead, Scribe, and Subject Matter Experts. These roles coordinate triage, mitigation, and stakeholder communications during active production incidents.

How do I write a postmortem or after-action review after an outage?

An after-action review (AAR) is written by having the Scribe document timelines and actions during the incident, then using reference playbooks and checklists post-resolution to standardize learning and prepare the postmortem.

Can I use predefined runbooks to handle service degradations and security incidents?

Predefined runbooks can handle service degradations, security incidents, and outages across stack-agnostic environments. They provide severity-guided decision frameworks to determine escalation, mitigation actions, and necessary communications.

How do I keep stakeholders informed during a production incident?

Stakeholders are kept informed using predefined status-page and internal communication templates. The Communications Lead manages these updates based on the severity-guided decision framework throughout the incident lifecycle.

What is the best way to coordinate triage and mitigation for a production outage?

The best way to coordinate outage triage and mitigation is applying a repeatable framework with predefined incident roles and severity-guided runbooks. This ensures structured decision-making for safe, fast resolution across stack-agnostic environments.