judge

Score test-kitchen implementations using a fixed five-criteria rubric.

90|6|Updated Oct 15, 2025
One-click install
npx skills add https://github.com/2389-research/claude-plugins --skill judge
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: judge
Source: https://github.com/2389-research/claude-plugins/tree/main/test-kitchen/skills/judge
Command: npx skills add https://github.com/2389-research/claude-plugins --skill judge

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Scores test-kitchen implementations using a fixed five-criteria rubric to provide objective, comparable assessments across cookoff and omakase-off events.

Core Features & Use Cases

  • Structured Gate Check and Scoring Worksheet generation for each candidate implementation (impl-1, impl-2, impl-3; variant-a, variant-b).
  • Enforces consistent evaluation through the required Gate Check, Feasibility, Scoring Worksheet, Judge Scorecard, and Hard Gates sections.
  • Use Case: Run the judge on multiple implementations and obtain a complete, comparable scoring report suitable for decision-making.

Quick Start

Run the judge on your candidate implementations to produce a complete scoring report in the prescribed format.

Frequently Asked Questions about judge

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I score multiple implementations in a cookoff fairly?

Score multiple implementations by applying a fixed five-criteria rubric to generate objective, comparable assessments across cookoff and omakase-off events. This structured approach enforces consistent evaluation through Gate Check and scoring worksheet sections.

What is a rubric-based scoring worksheet for test-kitchen implementations?

A rubric-based scoring worksheet is a structured format for evaluating test-kitchen implementations using fixed criteria including Gate Check, Feasibility, Scoring Worksheet, Judge Scorecard, and Hard Gates sections to ensure objective comparisons across candidates.

How do I generate a comparable scoring report for candidate implementations?

Generate a comparable scoring report by running the judge on candidate implementations like impl-1, impl-2, impl-3 or variant-a, variant-b to produce structured Gate Check and scoring worksheet outputs suitable for decision-making.

Can I use a fixed five-criteria rubric for omakase-off events?

Yes, the fixed five-criteria rubric applies to both cookoff and omakase-off scenarios to compare multiple implementations fairly. It enforces the required output format, structured worksheets, and rubric-based scoring rules for objective assessment.

What sections are required in a structured scoring report for test-kitchen implementations?

Required sections include Gate Check, Feasibility, Scoring Worksheet, Judge Scorecard, and Hard Gates. These sections enforce consistent evaluation and produce a complete scoring report in the prescribed format for comparing candidate implementations.