incident-triage

Frame production incidents to guide containment and recovery.

3|Updated Mar 24, 2026
One-click install
npx skills add https://github.com/ldilov/harness-forge --skill incident-triage-ldilov
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-triage
Source: https://github.com/ldilov/harness-forge/tree/main/skills/incident-triage
Command: npx skills add https://github.com/ldilov/harness-forge --skill incident-triage-ldilov

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Rapid incident framing helps teams quickly understand the scope of outages, regressions, and severe failures to accelerate containment and recovery.

Core Features & Use Cases

  • Identify blast radius and failure signature
  • Isolate the most likely failing boundary
  • Define contain, mitigate, and verify steps; produce a concise status summary

Quick Start

Describe an incident triage plan for a current outage and outline steps to contain, mitigate, and verify recovery.

Frequently Asked Questions about incident-triage

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I determine the blast radius of a production outage?

To determine the blast radius of a production outage, you analyze service logs, telemetry, and recent changes to identify the failure signature. This isolates failing boundaries and scopes affected services for rapid containment and recovery.

What is the best way to frame an incident for rapid root-cause analysis?

Framing an incident for root-cause analysis involves assessing the failure signature across logs and telemetry. This isolates the most likely failing boundary and defines steps to contain, mitigate, and verify recovery actions.

How do I create an incident triage plan to contain and mitigate severe failures?

You create an incident triage plan by evaluating recent changes and telemetry to define containment and mitigation steps. This isolates the failing boundary and produces a concise incident status summary to guide recovery verification efforts.

Can I use this incident triage approach for service regressions and outages?

Yes, you can use this incident triage approach for service regressions and outages. It applies to severe failures across services by utilizing available logs and telemetry to define verification plans and produce a status summary.

What steps are needed to isolate failing boundaries during an incident?

Steps to isolate failing boundaries include analyzing the failure signature, blast radius, and recent changes. This frames the incident to define containment and verification plans, producing a concise incident status summary for recovery.