sre-incident-response

Manage production incidents across detection, triage, mitigation, and postmortem workflows.

187|20|Updated Nov 20, 2025
One-click install
npx skills add https://github.com/TheBushidoCollective/han --skill sre-incident-response
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-incident-response
Source: https://github.com/TheBushidoCollective/han/tree/main/do/do-site-reliability-engineering/skills/sre-incident-response
Command: npx skills add https://github.com/TheBushidoCollective/han --skill sre-incident-response

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill guides you through effective incident response and postmortem processes, minimizing downtime and ensuring continuous learning from production issues. It automates the structured approach to incident management.

Core Features & Use Cases

  • Incident Severity & Process: Standardize incident classification (P0-P3) and follow a clear 5-step response process from detection to follow-up.
  • Role Definition & Communication: Assign clear roles (Incident Commander, Ops Lead, Comms Lead) and utilize templates for consistent internal and external communication.
  • Blameless Postmortems: Conduct structured postmortems to identify root causes, track action items, and foster a culture of continuous improvement without blame.
  • Use Case: During a critical P0 outage, activate this Skill to quickly assign roles, use the communication templates to keep stakeholders informed, and ensure all steps for mitigation and resolution are followed, leading to faster recovery and effective follow-up.

Quick Start

Use the sre-incident-response skill to draft an initial P1 incident notification for a service experiencing elevated error rates, including impact, IC, and next update time.

Frequently Asked Questions about sre-incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I respond to a production incident efficiently?

Incident response involves detecting, triaging, mitigating, and resolving production issues using a standardized process. This Skill guides you through a 5-step workflow for P0–P3 incidents, assigns clear roles like Incident Commander and Ops Lead, and provides communication templates to minimize downtime and keep stakeholders informed throughout recovery.

What should I include in an incident notification to stakeholders?

An incident notification must communicate impact, severity level, assigned Incident Commander, current status, and expected next update time. This Skill provides templates for consistent internal and external communication, ensuring all stakeholders receive timely, structured updates during active incidents.

How do I conduct a blameless postmortem after an outage?

A blameless postmortem is a structured review that identifies root causes and tracks action items without assigning individual blame. This Skill provides frameworks to document what happened, why it happened, and how to prevent recurrence, fostering continuous improvement and organizational learning from incidents.

What are the standard incident severity levels and when do I use them?

Incidents are classified into severity levels P0–P3, each triggering different response processes and escalation paths. This Skill defines each level's criteria and corresponding workflows, enabling teams to standardize incident classification and apply appropriate resources based on business impact.

What roles should I define for incident response?

Key incident roles include Incident Commander, who leads response; Ops Lead, who manages technical mitigation; and Comms Lead, who handles stakeholder communication. This Skill defines responsibilities for each role, ensuring clear ownership and coordinated action during production incidents.

Can I automate escalation paths for critical incidents?

Yes. This Skill supports automated escalation paths for incidents across severity levels, routing alerts to the right teams and decision-makers based on incident classification. Automation reduces response time and ensures no critical incident goes unaddressed.