incident-response

Triage production incidents and provide mitigation steps and post-mortem analysis.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/4asaanAI/Claude-patches --skill incident-response-4asaanai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/4asaanAI/Claude-patches/tree/main/framework-foundry/Claude%20Plugins/layaa-ai/skills/incident-response
Command: npx skills add https://github.com/4asaanAI/Claude-patches --skill incident-response-4asaanai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Provides a structured process to triage, diagnose, mitigate, communicate, and perform post-mortems for production incidents and outages, reducing downtime and coordinating clear, timely responses.

Core Features & Use Cases

  • Triage & Prioritization: Rapidly assess impact, scope, and severity to prioritize actions and stakeholders.
  • Diagnosis & Mitigation Guidance: Guide investigations, suggest immediate mitigations or rollbacks, and outline follow-up fixes.
  • Communication & Post-mortem: Draft incident communications for stakeholders and produce a post-mortem with root cause analysis and action items.
  • Use Case: An on-call engineer receives alerts about increased error rates; use this Skill to triage impact, propose immediate mitigations, coordinate on-call actions, and produce a post-mortem.

Quick Start

Triage the production outage for service X by summarizing impact, listing immediate mitigation steps, hypothesizing root causes, and drafting a stakeholder update.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage a production outage for a microservice?

To triage a production outage, assess the incident's impact, scope, and severity to prioritize mitigation actions. This structured approach guides on-call engineers through investigations and proposes immediate mitigations or rollbacks to restore service availability.

What is the best way to write a post-mortem after an incident?

Writing a post-mortem involves performing a root cause analysis and outlining specific remediation steps and action items. It provides a structured process to document the incident timeline, diagnose underlying causes, and prevent future outages.

Can I use this incident response process for cloud-based degraded performance?

Yes, this incident response process is explicitly applicable to cloud-based and microservice environments. It handles various service availability issues including degraded performance, data integrity problems, and emergency rollbacks.

How do I draft stakeholder communications during a service outage?

Drafting stakeholder communications during a service outage uses provided communication templates to deliver clear, timely updates. It ensures coordinated messaging regarding impact assessment, mitigation progress, and resolution expectations.

What immediate mitigation steps should I take for increased error rates?

For increased error rates, immediate mitigation steps include hypothesizing root causes and executing emergency rollbacks or configuration changes. The process guides investigations to propose quick fixes that stabilize service availability.