incident-responder

Guide triage, signal collection, and remediation planning for production incidents.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/matt-metivier/zk-hub --skill incident-responder-matt-metivier
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-responder
Source: https://github.com/matt-metivier/zk-hub/tree/main/skills/general/tools/incident-responder
Command: npx skills add https://github.com/matt-metivier/zk-hub --skill incident-responder-matt-metivier

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill helps engineering teams rapidly investigate and remediate incidents by unifying observability data, runbooks, and past incident patterns to reduce MTTR.

Core Features & Use Cases

  • Triage and classification, signal gathering, and correlation across deployments and services
  • Guided root-cause analysis using logs, traces, and metrics with actionable remediation steps
  • Post-incident documentation and knowledge capture for future prevention

Quick Start

Provide an incident title and affected service to trigger triage and remediation guidance.

Frequently Asked Questions about incident-responder

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I structure incident triage to reduce MTTR during a production outage?

Structured incident triage reduces MTTR by guiding signal collection and correlation across logs, metrics, and traces to classify the event and reference runbooks for actionable remediation steps.

What is the best way to correlate observability signals during an on-call event?

Correlating observability signals during an on-call event involves unifying logs, traces, and metrics to identify root-cause patterns across affected microservices and guide structured remediation planning.

How do I conduct root-cause analysis for degraded microservices?

Root-cause analysis for degraded microservices is conducted by correlating observability data with deployment events and referencing runbooks to determine actionable remediation steps.

Can I use this incident workflow for post-incident reviews and knowledge capture?

This incident workflow can be used for post-incident reviews to document summaries, capture knowledge of past incident patterns, and support future prevention across microservices.

Does this incident response process work for customer-reported issues?

This incident response process works for customer-reported issues by triggering triage and remediation guidance using an incident title and affected service to investigate degraded services.