incident-response

Guide production incident handling across detection, response, resolution, and post-mortem phases.

1|Updated Jan 6, 2026
One-click install
npx skills add https://github.com/hyukudan/ai-skills --skill incident-response-hyukudan
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/hyukudan/ai-skills/tree/main/examples/skills/incident-response
Command: npx skills add https://github.com/hyukudan/ai-skills --skill incident-response-hyukudan

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a comprehensive guide and actionable steps for effectively managing and resolving production incidents, minimizing downtime and impact.

Core Features & Use Cases

  • Incident Lifecycle Management: Covers detection, response, resolution, and post-mortem phases.
  • Role Definition: Outlines responsibilities for solo responders and small/large teams.
  • Communication Templates: Provides standardized templates for internal and external updates.
  • Mitigation Strategies: Offers common technical approaches to resolve issues.
  • Post-Mortem Framework: Guides the creation of blameless post-mortems for continuous improvement.

Quick Start

Use the incident-response skill to guide me through the detection phase of a production outage.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage a production outage across detection, response, and resolution?

Production outage management requires a structured framework covering detection, response, resolution, and post-mortem phases. This skill provides mitigation strategies, role definitions for various team sizes, and debugging checklists to navigate incidents efficiently.

What is a blameless post-mortem and how do I write one after an incident?

A blameless post-mortem focuses on systemic issues rather than individual errors after an incident. This skill guides you through creating a standardized post-mortem framework to ensure continuous improvement and objective analysis of production failures.

Can I use this incident response framework for a solo on-call responder?

Yes, the incident response framework outlines specific responsibilities for solo on-call responders as well as small and large teams. It adjusts mitigation strategies and communication templates to fit your specific team size and incident severity.

How do I handle communication updates during an SRE incident?

SRE incident communication relies on standardized templates for both internal and external updates. This skill provides ready-to-use communication templates to ensure consistent, accurate information delivery during detection, response, and resolution phases.

What is the best way to structure debugging checklists for DevOps incident management?

Debugging checklists for DevOps incident management should follow structured mitigation strategies across the incident lifecycle. This skill provides actionable checklists and common technical approaches to systematically resolve production issues and minimize downtime.

When do I need a formal incident response process for production incidents?

A formal incident response process is needed when managing production incidents that require structured detection, mitigation, and resolution. This skill ensures efficient handling of various incident severities for SRE and DevOps teams to minimize downtime.