incident-diagnosis-review

Evaluate AI agent incident diagnoses for accuracy, efficiency, and escalation.

Updated May 11, 2026
One-click install
npx skills add https://github.com/rafalwizen/plugin-architecture-test --skill incident-diagnosis-review-rafalwizen
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-diagnosis-review
Source: https://github.com/rafalwizen/plugin-architecture-test/tree/main/tools/kg-incidents/evaluator_skills/incident-diagnosis-review
Command: npx skills add https://github.com/rafalwizen/plugin-architecture-test --skill incident-diagnosis-review-rafalwizen

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Manually reviewing AI-generated incident diagnoses for root cause accuracy, diagnostic efficiency, and correct ownership escalation is inconsistent and time-consuming, leading to missed issues or misrouted incidents. This Skill standardizes the evaluation process with evidence-based scoring rules to ensure fair, consistent assessments of AI incident response performance.

Core Features & Use Cases

  • Structured 4-Dimension Scoring: Evaluates root cause accuracy, diagnostic path efficiency, fix scoping, and escalation awareness with clear, calibrated rubrics.
  • Evidence-Based Evaluation: Requires direct quotes from agent trajectories and workspace artifacts for all scores, eliminating guesswork and bias.
  • Ground Truth Calibration: Anchors scoring against predefined ground truth decisions to ensure consistency across different evaluators and incidents.
  • Use Case: SRE teams and platform engineers can use this Skill to fairly assess AI incident response agents, identify performance gaps, and ensure agents are correctly routing incidents to owning teams.

Quick Start

Invoke the incident-diagnosis-review skill to evaluate the AI agent's incident diagnosis output for the ai-description latency degradation using the provided trajectory and workspace artifacts.

Frequently Asked Questions about incident-diagnosis-review

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI agent incident diagnosis accuracy in SRE workflows?

You evaluate AI incident diagnosis accuracy by scoring the agent's execution trajectory and workspace artifacts against ground truth decision files, measuring root cause precision, diagnostic efficiency, fix scoping, and correct ownership escalation.

What is incident diagnosis review for AI agent performance?

Incident diagnosis review is an evaluation process that applies evidence-based, calibrated rubrics to assess AI agent responses during production incidents, ensuring fair and consistent measurement of root cause accuracy and escalation routing.

How do I score AI incident response escalation awareness?

You score escalation awareness by comparing the agent's routing decisions against predefined ground truth ownership files, requiring direct quotes from the agent's trajectory to validate that incidents are escalated to the correct owning teams.

Can I use ground truth calibration for root cause analysis evaluation?

Yes, ground truth calibration anchors incident diagnosis scoring against predefined decision files, ensuring consistent root cause analysis evaluation across different evaluators and incidents by eliminating guesswork and bias.

What artifacts do I need to review AI incident response performance?

You need the agent's execution trajectory, workspace artifacts including commit history and incident reports, and ground truth decision files to produce evidence-based, calibrated scoring across the four evaluation dimensions.

Does manual AI incident diagnosis review cause inconsistent escalation routing?

Manual AI incident diagnosis review causes inconsistent escalation routing due to time constraints and human bias, which standardized evidence-based scoring rules solve by requiring direct trajectory quotes and ground truth calibration.