incident-management

Coordinate production incidents and post-mortems with structured playbooks and templates.

3|Updated Mar 5, 2026
One-click install
npx skills add https://github.com/bipinks/ghost-office --skill incident-management-bipinks
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-management
Source: https://github.com/bipinks/ghost-office/tree/main/.claude/skills/incident-management
Command: npx skills add https://github.com/bipinks/ghost-office --skill incident-management-bipinks

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Handling production incidents efficiently is challenging without a standardized process. This guide provides clear roles, timelines, and templates to reduce mean time to recovery and improve blameless post-mortems.

Core Features & Use Cases

  • Severity classification: predefined levels to triage and escalate.
  • Incident Commander (IC) duties: defined authority, communication cadence, and coordination.
  • Post-mortems & RCA templates: structured documentation to capture root cause and preventive actions.
  • Use cases include service outages, degraded performance, and critical incident runbooks across engineering and operations.

Quick Start

Start a simulated incident, identify IC, classify severity, and generate a blameless post-mortem using the provided templates.

Frequently Asked Questions about incident-management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response to reduce mean time to recovery during a production outage?

Incident response coordination reduces mean time to recovery by assigning predefined Incident Commander roles, establishing communication cadences, and applying severity classification to triage outages efficiently.

What is a blameless post-mortem and how does it document root cause analysis?

A blameless post-mortem documents root cause analysis (RCA) using structured templates to capture preventive actions and incident timelines without assigning individual blame, fostering an SRE culture of systemic improvement.

How do I classify incident severity for degraded performance or service outages?

Incident severity classification uses predefined levels to triage degraded performance and service outages, ensuring engineering and operations teams apply the correct escalation paths and runbooks for critical incidents.

What should be included in an SRE runbook for resolving critical incidents?

An SRE runbook for critical incidents should include structured playbooks, detection steps, resolution workflows, and communication templates to guide operations teams from incident detection through final post-mortem documentation.

Can I use these incident management templates for both engineering and operations teams?

Yes, incident management templates support both engineering and operations teams by standardizing communication cadences, IC duties, and RCA workflows across diverse service outage and degraded performance scenarios.