What problem does it solve?
This Skill helps you evaluate competing skill outputs objectively by providing purpose-built agents that analyze results after comparisons, compare outputs without bias, and grade executions against expectations.
Core Features & Use Cases
- Analyzer (Post-hoc): Explains why a winner beat a loser by unblinding comparison inputs, inspecting the relevant skills, and reading execution transcripts to produce actionable improvement guidance.
- Grader: Grades expectations as PASS/FAIL using evidence from transcripts and output files, including checking for superficial or unverifiable claims.
- Comparator (Blind): Compares two outputs without knowing which skill produced them by generating a rubric, scoring structure and content, and selecting a winner (or tie) based on rubric outcomes.
Quick Start
Use the comparator agent to judge which of two outputs is better for a specific eval task, then run the analyzer to generate improvement suggestions for the losing skill.