incident-management

Standardize production incident response workflows with severity levels, roles, and postmortems.

5|1|Updated Jun 17, 2026
One-click install
npx skills add https://github.com/roanbrasil/engineer-grade-agent-skills --skill incident-management-roanbrasil
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-management
Source: https://github.com/roanbrasil/engineer-grade-agent-skills/tree/main/skills/incident-management
Command: npx skills add https://github.com/roanbrasil/engineer-grade-agent-skills --skill incident-management-roanbrasil

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Unplanned production incidents cause extended downtime, revenue loss, and team burnout when response processes are uncoordinated or inconsistent. This Skill provides a standardized, battle-tested framework to respond to incidents quickly, minimize user impact, and turn outages into actionable reliability improvements.

Core Features & Use Cases

  • Standardized Incident Response: Clear severity levels, incident commander roles, and phased mitigation playbooks to cut mean time to recover (MTTR) for outages of any scale.
  • SLO & Error Budget Management: Burn rate alerting configurations and error budget policies to balance feature delivery with service reliability.
  • Blameless Postmortem Templates: Structured 5 Whys analysis and action item tracking to eliminate root causes and prevent incident recurrence.
  • Use Case: When your team experiences a SEV2 service degradation, use this Skill to follow the triage and mitigation steps, coordinate cross-team communication, and document a postmortem with clear action items assigned to owners.

Quick Start

Use the incident-management skill to guide your team through a full production outage response, from initial alert detection through postmortem documentation and action item tracking.

Frequently Asked Questions about incident-management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I standardize incident response workflows to reduce mean time to recover?

Standardize incident response workflows by implementing clear severity classification, incident commander roles, and phased mitigation playbooks to reduce mean time to recover and prevent repeated outages.

What is the best way to configure burn rate alerting for SLO management?

The best way to configure burn rate alerting for SLO management is applying error budget policies and burn rate configurations to dynamically balance feature delivery with service reliability.

How do I write a blameless postmortem to prevent incident recurrence?

Write a blameless postmortem by using structured 5 Whys analysis templates and tracking action items to assigned owners, eliminating root causes to prevent incident recurrence.

Can I use this framework for on-call rotation design to minimize team burnout?

Yes, you can use this framework for on-call rotation design to minimize team burnout by structuring incident commander roles and coordinating cross-team communication during SEV2 service degradations.

Does incident management cover chaos engineering game day planning?

Yes, incident management covers chaos engineering game day planning alongside SLO management and postmortem action item tracking to proactively test system reliability and prevent outages.