incident-triage-harness

Coordinate production incident triage workflows with evidence-first hypothesis testing.

124|10|Updated Dec 8, 2025
One-click install
npx skills add https://github.com/madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-triage-harness
Source: https://github.com/madebyaris/advance-minimax-m3-cursor-rules/tree/main/.cursor/skills/incident-triage-harness
Command: npx skills add https://github.com/madebyaris/advance-minimax-m3-cursor-rules --skill incident-triage-harness

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production-grade incident triage workflows that guide teams through evidence-first analysis to confirm symptoms, assess blast radius, correlate signals, inspect code, and apply safe mitigations during outages, regressions, or suspicious behavior.

Core Features & Use Cases

  • Evidence-loop with Step 0 to Step 4 guidance, including triage, inspection order, and mitigation bias
  • Structured prompts and templates for logs, traces, dashboards, and multimodal evidence
  • Clear closeout expectations and verification steps to validate mitigation and capture remaining unknowns

Quick Start

Begin an evidence-first incident triage by outlining the symptom, collecting evidence (logs, metrics, traces, and visuals), and selecting the smallest safe mitigation.

Frequently Asked Questions about incident-triage-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is an evidence-first incident triage workflow?

Evidence-first incident triage is a structured process that guides teams to confirm symptoms, assess blast radius, correlate signals, and apply safe mitigations during outages or regressions.

How do I triage production outages using logs, metrics, and traces?

Coordinate production incident triage by collecting multimodal evidence like logs, metrics, and traces, then applying structured hypothesis testing from Step 0 to Step 4 to select safe mitigations.

Can I use structured triage workflows for suspicious runtime behavior and regressions?

Yes, you can apply structured incident triage workflows to outages, alerts, regressions, or suspicious runtime behavior across services to ensure reliable incident resolution.

What is the best way to apply safe mitigations during an incident?

The best way to apply safe mitigations is by following an evidence-first loop to inspect code and correlate signals, ensuring you select the smallest safe mitigation to resolve the incident.

How do I validate incident mitigation and capture remaining unknowns?

Validate incident mitigation by following defined closeout expectations and verification steps, which ensure the mitigation is effective and capture any remaining unknowns for resolution.