agency-incident-response-commander

Coordinate production incident response and post-mortem facilitation for distributed services.

Updated Apr 11, 2026
One-click install
npx skills add https://github.com/omeraltn/ice_cream_website_testing --skill agency-incident-response-commander-omeraltn
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agency-incident-response-commander
Source: https://github.com/omeraltn/ice_cream_website_testing/tree/main/.antigravity/agency-incident-response-commander
Command: npx skills add https://github.com/omeraltn/ice_cream_website_testing --skill agency-incident-response-commander-omeraltn

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps engineering teams turn chaotic production incidents into structured, timely responses by providing severity classification, role-based coordination, runbooks, communication templates, and post-mortem processes so incidents are resolved faster and organizational learning is captured.

Core Features & Use Cases

  • Structured Incident Command: Assign IC, comms lead, technical lead, and scribe with clear timeboxed decision steps and escalation triggers.
  • Runbooks & Remediation Playbooks: Templates for detection, diagnosis, rollback, restart, scaling, and verification to reduce MTTR.
  • Post-Mortem Facilitation: Blameless post-mortem templates, 5 Whys, action item tracking, and lessons learned to prevent repeats.
  • SLO/SLI & On-Call Design: SLO definitions, burn rate policies, and on-call rotation designs to guide when to page and when to pause feature work.
  • Use Case: Lead a SEV1 outage for a checkout API: declare severity, coordinate rollback or mitigation, communicate to stakeholders, verify SLIs, and produce a post-mortem with tracked actions.

Quick Start

Ask the agent to act as Incident Response Commander for a SEV2 outage on the checkout-api: assign roles, run diagnostics using runbook steps, propose immediate mitigations, and produce a timestamped timeline plus a post-mortem action list.

Frequently Asked Questions about agency-incident-response-commander

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response for a production outage in distributed services?

Incident response coordination assigns IC, comms lead, technical lead, and scribe roles with timeboxed decision steps and escalation triggers to resolve production outages faster. You get structured severity classification, runbook-driven diagnosis, and stakeholder communication templates.

What is a blameless post-mortem and how does it prevent repeat incidents?

A blameless post-mortem uses 5 Whys analysis and action item tracking to capture organizational learning without assigning fault. It produces templates and tracked remediation actions that prevent similar distributed service incidents from recurring.

How do I classify incident severity for a checkout API outage?

Incident severity classification uses severity matrices to evaluate impact on SLOs and user-facing functionality for distributed services. You apply SEV classification levels to trigger appropriate escalation, stakeholder communications, and runbook-driven mitigation steps.

Can I design on-call rotations and paging policies based on SLO burn rates?

On-call rotation designs integrate SLO definitions and burn rate policies to guide paging thresholds and feature work pauses. You get schedule designs and SLO/SLI frameworks that determine when to page engineers during distributed service degradation.

How do I run game-day exercises for incident response preparedness?

Game-day exercises simulate production incidents using runbook templates and severity matrices to test distributed service response workflows. You practice role-based coordination, runbook-driven diagnosis, and post-mortem facilitation before real outages occur.

What is the best way to reduce MTTR during a SEV1 incident?

Reducing MTTR during a SEV1 incident involves applying runbook templates for detection, diagnosis, rollback, restart, and scaling alongside structured incident command. You verify SLIs after mitigation and produce a timestamped timeline with tracked post-mortem actions.