What problem does it solve?
Single-pass evaluations of code or documents suffer from individual bias and shallow analysis. This Skill runs a structured multi-agent debate where independent judges challenge each other's assessments until they converge on accurate, evidence-based scores.
Core Features & Use Cases
- Meta-Judge Specification: A meta-judge generates tailored evaluation rubrics, checklists, and scoring criteria once, shared by all judges across all rounds.
- Independent Parallel Analysis: Three judges independently score the solution against the specification, preventing groupthink.
- Iterative Debate Rounds: Up to 3 debate rounds where judges defend positions with evidence, challenge disagreements, and revise scores until consensus.
- Use Case: After implementing a REST API, run the debate evaluation to get consensus scores on correctness, design, security, performance, and documentation, with a final pass/fail recommendation.
Quick Start
Ask the agent to evaluate your solution file with judge-with-debate, providing the solution path and the task it was supposed to accomplish.