What problem does it solve?
Manually reviewing AI-generated incident diagnoses for root cause accuracy, diagnostic efficiency, and correct ownership escalation is inconsistent and time-consuming, leading to missed issues or misrouted incidents. This Skill standardizes the evaluation process with evidence-based scoring rules to ensure fair, consistent assessments of AI incident response performance.
Core Features & Use Cases
- Structured 4-Dimension Scoring: Evaluates root cause accuracy, diagnostic path efficiency, fix scoping, and escalation awareness with clear, calibrated rubrics.
- Evidence-Based Evaluation: Requires direct quotes from agent trajectories and workspace artifacts for all scores, eliminating guesswork and bias.
- Ground Truth Calibration: Anchors scoring against predefined ground truth decisions to ensure consistency across different evaluators and incidents.
- Use Case: SRE teams and platform engineers can use this Skill to fairly assess AI incident response agents, identify performance gaps, and ensure agents are correctly routing incidents to owning teams.
Quick Start
Invoke the incident-diagnosis-review skill to evaluate the AI agent's incident diagnosis output for the ai-description latency degradation using the provided trajectory and workspace artifacts.