engineering-incident-response-commander

Coordinate production incident response and facilitate blameless post-mortems.

Updated Feb 16, 2026
One-click install
npx skills add https://github.com/Adawodu/dynoclaw --skill engineering-incident-response-commander
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: engineering-incident-response-commander
Source: https://github.com/Adawodu/dynoclaw/tree/main/skills/engineering-incident-response-commander
Command: npx skills add https://github.com/Adawodu/dynoclaw --skill engineering-incident-response-commander

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides expert guidance and structured processes for managing production incidents, ensuring rapid resolution, minimizing impact, and fostering a culture of continuous improvement through blameless post-mortems.

Core Features & Use Cases

  • Incident Response Coordination: Guides teams through severity classification, role assignment, and real-time troubleshooting.
  • Post-Mortem Facilitation: Helps conduct blameless post-mortems to identify systemic causes and drive actionable improvements.
  • Readiness Building: Assists in designing on-call rotations, creating runbooks, and establishing SLO/SLI frameworks.
  • Use Case: When a critical service outage occurs, this Skill can be invoked to guide the on-call team through the established incident response playbook, ensuring all necessary steps are followed for swift resolution and thorough documentation.

Quick Start

Guide the team through a SEV1 incident using the incident response commander skill.

Frequently Asked Questions about engineering-incident-response-commander

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response during a production system outage?

Incident response coordination involves classifying severity, assigning roles, and guiding real-time troubleshooting to manage production outages. This ensures rapid resolution by following established playbooks for swift restoration and thorough documentation.

What is a blameless post-mortem and how does it improve system reliability?

A blameless post-mortem is a structured review process that identifies systemic causes rather than individuals. Conducting blameless post-mortems drives actionable improvements and fosters continuous improvement, ultimately enhancing overall system reliability.

How do I set up an on-call rotation and create effective runbooks?

Building operational readiness involves designing structured on-call rotations and creating detailed runbooks. This preparation establishes clear response procedures and SLO/SLI frameworks, ensuring engineering teams are equipped for reliable production management.

What is the best way to classify incident severity during an active outage?

The best way to classify incident severity is using a structured command framework that evaluates impact and urgency. Proper severity classification directs response coordination, stakeholder communication, and resource allocation during production incidents.

When do I need to implement SLO and SLI frameworks for production management?

You need to implement SLO and SLI frameworks when establishing measurable reliability targets for production systems. These frameworks support continuous improvement by providing clear indicators of system health and guiding on-call process design.

How do I facilitate stakeholder communication during a SEV1 incident?

Facilitating stakeholder communication during a SEV1 incident requires structured command processes that provide consistent updates. Incident response coordination manages these communications alongside severity classification and real-time troubleshooting efforts.