incident-handling

Automate incident response from detection through post-incident analysis.

1|1|Updated Nov 27, 2025
One-click install
npx skills add https://github.com/HikaruEgashira/agent-skills --skill incident-handling
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-handling
Source: https://github.com/HikaruEgashira/agent-skills/tree/main/meta/skills/incident
Command: npx skills add https://github.com/HikaruEgashira/agent-skills --skill incident-handling

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Autonomous incident handling coordinates the full lifecycle from detection to post-incident analysis, reducing response time and preventing recurrence.

Core Features & Use Cases

  • Structured lifecycle: detection, visualization of impact, root-cause analysis, recovery planning, and post-incident learning.
  • Automated guidance: coordinates containment, recovery, and reporting steps to minimize manual toil during incidents.
  • Use Case: when production systems experience degradation, this skill guides operators to assess impact, identify root causes, select recovery strategies, and document findings for the postmortem.

Quick Start

Initiate autonomous incident handling for the current incident and outline immediate containment and recovery steps.

Frequently Asked Questions about incident-handling

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate incident response for distributed systems from detection to postmortem?

Automated incident response coordinates the full lifecycle from detection to post-incident analysis across distributed systems, reducing response time and preventing recurrence by guiding containment, recovery, and root-cause analysis steps.

What is the best way to conduct root-cause analysis during an active production incident?

Root-cause analysis during an active incident is structured through automated guidance that visualizes impact, identifies underlying causes, and selects recovery strategies while documenting findings for the postmortem report.

Can I use autonomous incident handling to minimize manual toil during system degradation?

Autonomous incident handling minimizes manual toil during system degradation by coordinating containment, recovery, and reporting steps automatically, requiring structured guidance and explicit knowledge of internal services, logs, and runbooks.

Do I need to provide runbooks and internal service knowledge for automated incident management?

Automated incident management requires explicit knowledge of internal services, logs, and runbooks, along with structured guidance and guardrails to prevent unsafe actions during autonomous containment and recovery operations.

What are the limitations of autonomous incident handling in production environments?

Autonomous incident handling in production environments requires guardrails to prevent unsafe actions, limiting operations to explicitly defined internal services, logs, and runbooks without unstructured or unsupported recovery paths.

How does incident management automation handle the post-incident learning phase?

Incident management automation handles the post-incident learning phase by documenting root-cause analysis findings and recovery strategies into a structured postmortem report, preventing future recurrence of the same production issues.