incident-postmortem

Generate blameless incident postmortems and structured runbooks for production outages.

7|Updated Mar 19, 2026
One-click install
npx skills add https://github.com/camilooscargbaptista/cto-toolkit --skill incident-postmortem-camilooscargbaptista
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-postmortem
Source: https://github.com/camilooscargbaptista/cto-toolkit/tree/main/incident-postmortem
Command: npx skills add https://github.com/camilooscargbaptista/cto-toolkit --skill incident-postmortem-camilooscargbaptista

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Engineering teams struggle to capture clear, blameless incident reports and reliable runbooks under time pressure, which slows learning, increases repeat failures, and creates poor stakeholder communication. This Skill standardizes postmortem and runbook creation so teams can quickly document impact, timeline, root cause, and actionable prevention steps that are specific, assigned, and trackable.

Core Features & Use Cases

  • Blameless Postmortems: Structured templates for executive summaries, UTC timelines, detailed root cause analysis, contributing factors, and an appendix for logs and evidence.
  • Incident Runbooks: Step-by-step diagnostic procedures, exact commands to run, expected outputs, rollback options, and escalation matrices for on-call responders.
  • Severity & Communication: Severity matrix, escalation protocols, and Slack/status-page message templates to ensure consistent stakeholder updates.
  • Use case: After a production outage that increased 5xx errors, use this Skill to produce a postmortem with a precise timeline, identify missing monitoring, assign SMART action items, and publish an updated runbook to prevent recurrence.

Quick Start

Create a blameless postmortem for the recent production outage affecting the user-signup service with an executive summary, UTC timeline, root cause analysis, SMART action items with owners and due dates, and a prevention runbook.

Frequently Asked Questions about incident-postmortem

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I write a blameless postmortem for a production outage?

To write a blameless postmortem for a production outage, generate a structured report covering the executive summary, UTC timeline, root cause analysis, and SMART action items with specific owners and due dates to prevent recurrence.

What should be included in an on-call runbook for service incidents?

An on-call runbook for service incidents should include step-by-step diagnostic procedures, exact commands to run, expected outputs, rollback options, and an escalation matrix to guide responders during outages.

How do I document root cause analysis and SLA impact for incident escalation?

Documenting root cause analysis and SLA impact for incident escalation requires detailing contributing factors, assessing severity, and applying standardized templates for consistent stakeholder communication during post-incident analysis.

Can I generate Slack and status-page communication templates during an incident?

Yes, you can generate Slack and status-page communication templates during an incident by applying standardized severity matrices and escalation protocols to ensure consistent updates for stakeholders.

What is the best way to assign trackable action items after a post-incident review?

The best way to assign trackable action items after a post-incident review is to generate SMART prevention steps with designated owners and explicit due dates, ensuring accountability for engineering teams.