incident-response

Structure incident response with severity levels, runbook templates, and postmortem analysis.

Updated Feb 26, 2026
One-click install
npx skills add https://github.com/engineers-hub-ltd-in-house-project/eh-skills --skill incident-response-engineers-hub-ltd-in-house-project
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/engineers-hub-ltd-in-house-project/eh-skills/tree/main/skills/monitoring/incident-response
Command: npx skills add https://github.com/engineers-hub-ltd-in-house-project/eh-skills --skill incident-response-engineers-hub-ltd-in-house-project

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a comprehensive framework for managing incidents, from initial detection and triage to effective response, recovery, and thorough post-incident reviews, ensuring a structured and efficient approach to critical events.

Core Features & Use Cases

  • Incident Response Flow: Defines a clear process from detection to postmortem.
  • Severity Levels: Establishes clear definitions for incident severity and corresponding response times.
  • Runbook Design: Offers a template for creating detailed incident response playbooks.
  • On-call Rotation: Provides a structure for managing on-call schedules and escalations.
  • Postmortem Template: Includes a template for conducting blame-free incident reviews and identifying action items.
  • Use Case: When a critical service outage occurs (SEV1), this skill guides the team through the defined incident response flow, ensuring rapid triage, communication, and resolution, followed by a structured postmortem to prevent recurrence.

Quick Start

Use the incident-response skill to generate a postmortem template for a recent service outage.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I structure incident response procedures for critical system events?

Incident response procedures require a structured framework covering severity definitions, triage flows, on-call rotation, and postmortem templates to standardize handling of critical system events. This framework ensures efficient communication, rapid escalation, and root cause analysis during service outages.

What is the best way to define incident severity levels and response times?

Incident severity levels categorize critical system events by impact, establishing clear definitions for severity tiers and corresponding response times. Standardized severity definitions guide the incident response flow, ensuring rapid triage and appropriate resource allocation during outages.

How do I create a blame-free postmortem template for incident reviews?

A blame-free postmortem template structures incident reviews by documenting the timeline, identifying root causes, and tracking action items to prevent future occurrences. The template facilitates thorough postmortem analysis, ensuring standardized procedures are followed after critical system events.

How do I manage on-call rotation and escalation paths for DevOps teams?

On-call rotation management structures schedules and escalation paths for DevOps teams handling incident response. It provides a defined framework for on-call duties, ensuring efficient communication and rapid escalation during critical system events and service outages.

Can I use a runbook template to standardize incident response playbooks?

A runbook template offers a structured design for creating detailed incident response playbooks. It standardizes procedures for handling critical system events, ensuring the team follows defined incident response flows from initial detection through recovery and postmortem analysis.