What problem does it solve?
Engineering teams struggle to capture clear, blameless incident reports and reliable runbooks under time pressure, which slows learning, increases repeat failures, and creates poor stakeholder communication. This Skill standardizes postmortem and runbook creation so teams can quickly document impact, timeline, root cause, and actionable prevention steps that are specific, assigned, and trackable.
Core Features & Use Cases
- Blameless Postmortems: Structured templates for executive summaries, UTC timelines, detailed root cause analysis, contributing factors, and an appendix for logs and evidence.
- Incident Runbooks: Step-by-step diagnostic procedures, exact commands to run, expected outputs, rollback options, and escalation matrices for on-call responders.
- Severity & Communication: Severity matrix, escalation protocols, and Slack/status-page message templates to ensure consistent stakeholder updates.
- Use case: After a production outage that increased 5xx errors, use this Skill to produce a postmortem with a precise timeline, identify missing monitoring, assign SMART action items, and publish an updated runbook to prevent recurrence.
Quick Start
Create a blameless postmortem for the recent production outage affecting the user-signup service with an executive summary, UTC timeline, root cause analysis, SMART action items with owners and due dates, and a prevention runbook.