incident-responder

Manages incident lifecycle from detection to blameless post-mortems using diagnostic scripts and predefined protocols.

Updated Feb 6, 2026
One-click install
npx skills add https://github.com/ntuan2502/piggy --skill incident-responder-ntuan2502
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-responder
Source: https://github.com/ntuan2502/piggy/tree/main/.agent/skills/incident-responder
Command: npx skills add https://github.com/ntuan2502/piggy --skill incident-responder-ntuan2502

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides an expert SRE and Incident Commander to rapidly restore service, maintain communication, and prevent future failures during critical incidents.

Core Features & Use Cases

  • Incident Management: Guides through detection, triage, declaration, and resolution.
  • Diagnosis & Fix: Uses logs, traces, and metrics for hypothesis-driven problem-solving and safe remediation.
  • Communication: Ensures transparent updates to internal teams and external stakeholders.
  • Post-Mortems: Facilitates blameless analysis and action items to prevent recurrence.

Quick Start

Run the health check script for the specified service URL.

Frequently Asked Questions about incident-responder

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I manage incident response for rapid service restoration?

Incident response for rapid service restoration is managed by guiding teams through detection, triage, declaration, and resolution. The process uses diagnostic scripts and predefined protocols to execute health checks and analyze system metrics for effective mitigation.

What is a blameless post-mortem and when do I need one for SRE?

A blameless post-mortem is a facilitated analysis that identifies action items to prevent future failures after an incident. You need this SRE process following any critical service disruption to maintain transparent communication and ensure problem prevention without assigning individual blame.

How do I diagnose system failures using logs, traces, and metrics?

Diagnose system failures by using logs, traces, and metrics for hypothesis-driven problem-solving and safe remediation. This approach analyzes system metrics and diagnostic data to identify root causes during the incident lifecycle and execute effective mitigation.

Can I use diagnostic scripts to run health checks on a service URL?

Yes, you can run diagnostic scripts to execute health checks on a specified service URL. This provides expert SRE incident response capabilities, checking system metrics and service status to support rapid service restoration and incident resolution.

What is the best way to maintain communication with stakeholders during an on-call incident?

The best way to maintain communication during an on-call incident is to ensure transparent updates to both internal teams and external stakeholders. An Incident Commander manages these communication protocols alongside service restoration efforts throughout the incident lifecycle.

Does incident response work without predefined protocols and system metrics analysis?

No, effective incident response requires predefined protocols and system metrics analysis for effective mitigation. Without executing health checks and analyzing diagnostic data from logs and traces, the SRE incident resolution process cannot safely remediate critical failures.