incident-response

Coordinate production incident response with mitigation-first protocols and structured communication.

25|3|Updated Jul 14, 2026
One-click install
npx skills add https://github.com/nimadorostkar/Claude-Skills-collection --skill incident-response-nimadorostkar
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/nimadorostkar/Claude-Skills-collection/tree/main/skills/devops/incident-response
Command: npx skills add https://github.com/nimadorostkar/Claude-Skills-collection --skill incident-response-nimadorostkar

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill prevents chaotic, ineffective incident management by enforcing a disciplined, priority-based workflow that focuses on immediate mitigation and blameless learning.

Core Features & Use Cases

  • Prioritized Mitigation: Enforces the critical rule of rolling back or mitigating before attempting to diagnose the root cause.
  • Structured Communication: Provides a framework for incident command roles and consistent status updates to prevent information silos.
  • Actionable Postmortems: Guides the creation of post-incident reports that result in concrete, owned, and dated action items rather than vague suggestions.

Quick Start

Use the incident-response skill to guide our team through the current production outage and draft a blameless postmortem.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage a production incident without chaotic responses?

Production incident management requires enforcing mitigation-first protocols and structured communication cadences to prevent chaos. This skill coordinates active service outages by prioritizing immediate rollback or mitigation before attempting root cause diagnosis.

What's the best way to structure on-call runbooks for production outages?

On-call runbooks should enforce incident command roles and consistent status updates to prevent information silos. This skill provides a framework for defining roles and producing verifiable, time-bound action items for reliable response.

How do I write a blameless postmortem after a service outage?

A blameless postmortem guides post-incident analysis by creating concrete, owned, and dated action items rather than vague suggestions. This skill ensures reports result in verifiable tasks that prevent future incidents.

Why should I mitigate a production incident before diagnosing the root cause?

Mitigating before diagnosing the root cause prevents extended downtime by enforcing the critical rule of rolling back or mitigating first. This priority-based workflow ensures immediate service restoration during active outages.

Can I use this incident response skill for post-incident analysis and on-call preparation?

Yes, this incident response skill applies to active service outages, post-incident analysis, and the development of reliable on-call runbooks. It requires adherence to defined incident command roles throughout each phase.

What incident command roles are needed for structured communication during outages?

Structured communication during outages requires defined incident command roles that enforce consistent status updates and prevent information silos. This skill coordinates these roles to maintain clear communication cadences.