agency-incident-response-commander

Coordinate production incident response and facilitate blameless post-mortems.

Updated Jul 23, 2026
One-click install
npx skills add https://github.com/rajyeole6/AI-RECRUITER --skill agency-incident-response-commander-rajyeole6
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agency-incident-response-commander
Source: https://github.com/rajyeole6/AI-RECRUITER/tree/main/.agents/skills/engineering-incident-response-commander
Command: npx skills add https://github.com/rajyeole6/AI-RECRUITER --skill agency-incident-response-commander-rajyeole6

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill mitigates the chaos of production outages by providing a structured, blameless framework for incident management, ensuring rapid resolution and long-term system reliability.

Core Features & Use Cases

  • Structured Incident Command: Implements clear severity frameworks (SEV1-SEV4) and role assignments to prevent coordination breakdown during high-pressure events.
  • Blameless Post-Mortems: Facilitates deep-dive analysis into systemic failures using the 5 Whys, ensuring organizational learning rather than individual blame.
  • Operational Readiness: Provides templates for runbooks, SLO/SLI tracking, and on-call rotation design to proactively reduce system fragility.

Quick Start

Initiate the incident response commander to declare a new SEV2 incident and generate the initial communication template for the engineering team.

Frequently Asked Questions about agency-incident-response-commander

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate a production incident response to minimize system downtime?

Coordinate production incident response by establishing clear severity classifications from SEV1 to SEV4, assigning specific response roles, and maintaining structured cross-team communication to minimize downtime and prevent coordination breakdown.

How do I run a blameless post-mortem after a major outage?

Run a blameless post-mortem by facilitating deep-dive analysis into systemic failures using the 5 Whys methodology, ensuring organizational learning and long-term system reliability rather than focusing on individual blame.

What's the best way to structure on-call rotations and SLO tracking for SRE teams?

Structure on-call rotations by utilizing operational readiness templates for runbooks and SLO/SLI tracking, which proactively reduces system fragility and establishes clear escalation paths for production incidents.

When do I need to declare a formal incident severity level during an outage?

Declare a formal incident severity level immediately upon detecting a production outage, utilizing a structured SEV1 through SEV4 framework to dictate the required response scale, communication frequency, and remediation efforts.

Does this incident management approach work without established SRE methodologies?

This incident management approach requires adherence to established SRE methodologies and blameless culture principles to effectively coordinate technical remediation, root cause analysis, and action item tracking across teams.

How do I track action items and root causes after resolving an incident?

Track action items and root causes by operating across the entire incident lifecycle from initial detection to post-mortem facilitation, ensuring systemic failures are documented and remediation steps are monitored for long-term reliability.