What problem does it solve? Comparing candidate outputs or picking the best final answer is error-prone when judgments rely on gut feel, presentation polish, or invented reasoning about the process. This Skill enforces a disciplined, evidence-anchored method for judging end-state quality without needing full execution trajectories. ## Core Features & Use Cases - Locked Outcome Rubrics: Define 3-7 task-specific scoring dimensions with explicit pass/partial/fail boundaries before judging any candidate. - Bias-Aware Comparison: Run order-robustness checks, distrust unsupported chain-of-thought, and calibrate confidence when evidence is thin. - Decision Packages: Produce a concise verdict with the winner, margin, decisive evidence, rejected-candidate failure modes, and uncertainty notes. - Use Case: Given two candidate support replies, a reference answer, and explicit criteria, the Skill locks a shared rubric, scores both replies on end-state evidence, and reports the stronger outcome with a calibrated confidence level. ## Quick Start Compare these two candidate responses using only the final outputs and the provided reference criteria, then pick the stronger outcome with a confidence note.