incident-response

Coordinate incident response workflows for detection, communication, mitigation, and post-mortem analysis.

2|1|Updated Apr 13, 2026
One-click install
npx skills add https://github.com/dreamingechoes/dx-toolkit --skill incident-response-dreamingechoes
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/dreamingechoes/dx-toolkit/tree/main/templates/skills/incident-response
Command: npx skills add https://github.com/dreamingechoes/dx-toolkit --skill incident-response-dreamingechoes

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The Incident Response Skill provides a repeatable, blameless playbook to detect, classify, respond to, and learn from production incidents, reducing mean time to recovery and improving communication across teams.

Core Features & Use Cases

  • Detection and classification with severity levels to triage incidents
  • Assembling responders and escalation to minimize response time
  • Structured communication templates and real-time timeline documentation
  • Root cause investigation framework and post-mortem guidance

Quick Start

Describe your incident scenario and severity to generate an actionable runbook.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate an incident response workflow during a production outage?

An incident response workflow coordinates detection, communication, mitigation, and post-mortem analysis for on-call teams managing outages. It provides a structured framework with severity levels, responder roles, and escalation rules to reduce mean time to recovery.

What is a blameless post-mortem and how do I structure one after an incident?

A blameless post-mortem is a structured template for root cause investigation that documents timelines and analyzes systemic failures without assigning individual fault. It extracts learnings from degraded services and production outages to prevent recurrence.

How do I classify incident severity levels for degraded services?

Incident severity levels classify the impact of degraded services during detection and triage. This classification determines the appropriate responder roles, escalation rules, and communication templates to activate for the production incident.

Can I use this incident response runbook for on-call teams managing software systems?

Yes, the incident response runbook is applicable to on-call teams managing outages and degraded services across software systems. It provides structured communication templates, real-time timeline documentation, and escalation rules suited for these environments.

What is the best way to document a real-time timeline during an outage?

The best way to document a real-time timeline during an outage is using structured communication templates that log mitigation actions and status updates as they happen. This ensures accurate chronological records for blameless post-mortem analysis.