What problem does it solve?
Provides a structured, repeatable process for responding to production outages and degradations, reducing time-to-detection and time-to-resolution while ensuring organizational learning through blameless post-mortems.
Core Features & Use Cases
- Severity classification: Classify incidents as SEV1–SEV4 with clear response time expectations and escalation rules.
- Timeline & evidence collection: Build UTC-timestamped timelines with linked logs, deploy records, and monitoring snapshots to avoid guessing.
- Root cause analysis & action tracking: Run guided 5-Whys analysis, produce a blameless post-mortem template, and track actionable items with owners and deadlines.
- Investigation loop & safety rules: Follow a guarded investigation loop that tests hypotheses, keeps/discards changes based on evidence, and enforces non-blaming and follow-up rules.
Quick Start
Invoke the incident skill by saying "/godmode:incident" and provide a concise incident summary including observed symptoms, time window, and any related deploys.