incident-response-template

Configure production monitoring, alerting, and incident response runbooks.

69|5|Updated Nov 16, 2025
One-click install
npx skills add https://github.com/nahisaho/musubi --skill incident-response-template
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response-template
Source: https://github.com/nahisaho/musubi/tree/main/.claude/skills/site-reliability-engineer/incident-response-template.md
Command: npx skills add https://github.com/nahisaho/musubi --skill incident-response-template

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This document provides a reusable incident response runbook template to guide detection, triage, mitigation, and post-mortem processes.

Core Features & Use Cases

  • Phase-based incident lifecycle
  • Severity levels and escalation
  • Post-mortem structure and action items

Quick Start

Use the template to document a new incident with a clear timeline and responsibilities.

Frequently Asked Questions about incident-response-template

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an incident response runbook for my production system?

An incident response runbook is a documented procedure guiding your team through detection, triage, mitigation, and post-mortem phases. This template structures those phases with severity levels, escalation paths, and clear responsibilities so incidents are handled consistently and tracked for learning.

What should be included in a post-mortem after a production incident?

Post-mortems document the incident timeline, root cause, impact, and action items to prevent recurrence. This template provides a standardized structure capturing severity context, who was involved, what happened, and what the team will do differently—enabling systematic improvement across incidents.

How do I standardize incident response across my SRE team?

Standardization starts with a shared runbook template defining phases, severity criteria, and escalation procedures. This template gives your team a common framework for detection, communication, and resolution, reducing decision fatigue and ensuring consistent handling regardless of who responds.

What's the best way to structure incident severity levels and escalation?

Severity levels categorize incidents by business impact and guide escalation—determining who to notify, response time targets, and decision authority. This template includes severity definitions and escalation paths so your team knows exactly when and whom to involve at each stage.

Can I use this runbook template with Prometheus, Grafana, or other monitoring tools?

Yes. This template is monitoring-agnostic and works with any observability stack—Prometheus, Grafana, ELK, Datadog, New Relic, or others. It focuses on the incident lifecycle and response process rather than specific tooling, so you integrate it with whatever alerts and dashboards you use.