incident-responder

Coordinate incident command, investigation, and post-incident analysis for production outages.

Updated Feb 24, 2026
One-click install
npx skills add https://github.com/chicanoandres702/SentientAIBrowser --skill incident-responder-chicanoandres702
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-responder
Source: https://github.com/chicanoandres702/SentientAIBrowser/tree/main/.agents/workflows/incident-responder
Command: npx skills add https://github.com/chicanoandres702/SentientAIBrowser --skill incident-responder-chicanoandres702

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Expert incident response and management are often manual and error-prone under outages. This skill provides a structured, repeatable workflow for rapid problem resolution, effective communication, and post-incident learning to restore service quickly and safely.

Core Features & Use Cases

  • Establishes an on-call Incident Command structure and role delegation for quick decision-making
  • Observability-driven investigation using logs, metrics, traces, and incident telemetry
  • Blameless post-mortems, root-cause analysis, and action-item tracking to prevent recurrence
  • Stakeholder and customer communications templates and status updates during outages
  • Stabilization, rollback evaluation, and staged recovery guidance

Quick Start

Initiate incident response by assuming command and kicking off stabilization and detailed investigation workflows.

Frequently Asked Questions about incident-responder

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response during a production outage?

Incident response coordination during a production outage requires establishing an on-call Incident Command structure for quick decision-making, assigning predefined roles, and executing immediate stabilization workflows to restore service safely.

What is a blameless post-mortem in SRE?

A blameless post-mortem in SRE is a post-incident analysis process focused on root-cause analysis and action-item tracking to prevent recurrence, without attributing fault to individual responders.

Can I use this for reliability investigations across distributed systems?

Yes, you can use this for reliability investigations across distributed systems, as it supports observability-driven investigation using logs, metrics, traces, and incident telemetry to diagnose degraded services.

How do I manage stakeholder communications during service degradation?

Managing stakeholder communications during service degradation involves using predefined communication templates and sending regular status updates to stakeholders and customers throughout the outage lifecycle.

What's the best way to evaluate rollback and staged recovery for outages?

Evaluating rollback and staged recovery for outages requires following structured stabilization guidance that assesses deployment reversions and coordinates staged restoration of affected production services.

Does this incident management workflow require external dependencies?

No, this incident management workflow operates without external dependencies, providing a structured, repeatable framework for rapid problem resolution, communication, and post-incident learning.