incident-response

Guide incident response with severity triage, blast radius scoping, and blameless postmortems.

13|3|Updated Mar 27, 2026
One-click install
npx skills add https://github.com/heaptracetechnology/heaptrace-skills --skill incident-response-heaptracetechnology
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/heaptracetechnology/heaptrace-skills/tree/main/lead-engineer/incident-response
Command: npx skills add https://github.com/heaptracetechnology/heaptrace-skills --skill incident-response-heaptracetechnology

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill guides teams through a blameless incident response workflow, turning chaotic outages into a repeatable, safe, and learnable process.

Core Features & Use Cases

  • Structured severity classification, blast radius scoping, and runbook templates for fast response.
  • End-to-end incident lifecycle: detect, triage, fix, and postmortem with actionable item management.
  • Use cases include production outages, security incidents, postmortem retros, and on-call playbook adoption.

Quick Start

Observe an outage, classify severity, scope the blast radius, and follow the triage-to-postmortem playbook to restore service and prevent recurrence.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I structure incident response for a production outage?

Incident response for a production outage is structured by triaging severity, scoping the blast radius, and implementing fixes followed by a blameless postmortem. This lifecycle approach turns chaotic outages into a repeatable, safe workflow.

What is a blameless postmortem and when should I use it?

A blameless postmortem is a retrospective analysis focusing on systemic root causes rather than individual fault. It is used after production outages or critical bugs to manage actionable items and prevent recurrence across the team.

How do I classify incident severity and scope the blast radius during triage?

Incident severity classification and blast radius scoping are handled during the triage phase using structured guidance and runbook templates. This ensures fast, disciplined response by defining impact levels before implementing fixes.

Can I use this incident response playbook for security incidents and on-call adoption?

Yes, the incident response playbook applies to security incidents, production outages, and on-call playbook adoption. It provides end-to-end lifecycle management from detection through validation and learning workflows.

What is the best way to run a root-cause analysis after a critical bug?

The best way to run root-cause analysis is through a structured post-incident template that documents triage decisions and remediation steps. This blameless approach captures actionable items to prevent future critical bugs.