incident-management

Implements structured incident management processes with severity assessment and response coordination.

46|4|Updated Jan 27, 2026
One-click install
npx skills add https://github.com/BagelHole/DevOps-Security-Agent-Skills --skill incident-management-bagelhole
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-management
Source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/compliance/continuity/incident-management
Command: npx skills add https://github.com/BagelHole/DevOps-Security-Agent-Skills --skill incident-management-bagelhole

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps organizations establish and follow structured processes for managing production incidents, ensuring timely resolution and minimizing downtime.

Core Features & Use Cases

  • Incident Response Workflow: Defines clear steps from detection to post-incident review.
  • Severity Levels: Establishes a common understanding of incident impact and urgency.
  • Incident Commander Role: Outlines the responsibilities of the person leading the response.
  • Post-Incident Reviews: Facilitates learning and improvement through blameless analysis.
  • Use Case: When a critical service experiences an outage, this Skill guides the team through the defined incident management process, from initial alert to root cause analysis and action item creation.

Quick Start

Use the incident-management skill to initiate a SEV1 incident response process.

Frequently Asked Questions about incident-management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage production incidents effectively during an outage?

To manage production incidents effectively, follow a structured workflow from initial detection to post-incident review. This includes severity assessment, response coordination, and establishing clear escalation paths to ensure timely resolution and minimize downtime.

What is a blameless post-mortem and when do I need one?

A blameless post-mortem is a post-incident review focused on learning and improvement through objective root cause analysis. You need this process after resolving any production incident to facilitate learning and create actionable improvement items.

How do I establish incident severity levels for production support?

Establish incident severity levels by assessing the impact and urgency of the production incident. Defining clear severity levels creates a common understanding across the team, guiding the appropriate response coordination and escalation paths required for efficient resolution.

What are the responsibilities of an incident commander during escalation?

The incident commander leads the incident response coordination and manages the escalation process during a production outage. Their responsibilities include guiding the team through the defined incident management process, from initial alert to resolution and post-incident review.

Can I use this approach for on-call scheduling and communication protocols?

Yes, this incident management approach supports on-call scheduling and defines communication protocols. It establishes clear escalation paths and incident response workflows, ensuring the on-call team can efficiently coordinate and resolve production incidents.