run-judges

Orchestrate parallel AI judge agents to evaluate implementation plans and code artifacts.

102|10|Updated Mar 4, 2026
One-click install
npx skills add https://github.com/closedloop-ai/claude-plugins --skill run-judges
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: run-judges
Source: https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges
Command: npx skills add https://github.com/closedloop-ai/claude-plugins --skill run-judges

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the rigorous quality assessment of implementation plans and code artifacts by orchestrating specialized AI judge agents, ensuring consistent and objective evaluation.

Core Features & Use Cases

  • Parallel Judge Execution: Runs multiple AI judges concurrently to evaluate plans (16 judges) or code (11 judges).
  • Automated Reporting: Aggregates judge outputs into a structured judges.json or code-judges.json report.
  • Configurable Thresholds: Allows customization of evaluation strictness via override files.
  • Use Case: After generating a code implementation for a new feature, use this Skill to automatically run a suite of judges that assess code quality, adherence to best practices, and test coverage, providing a detailed quality report.

Quick Start

Run the judges skill to evaluate the current implementation plan.

Frequently Asked Questions about run-judges

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate code quality evaluation for my implementation plans?

Automated code quality evaluation orchestrates parallel AI judge agents to assess plans and code artifacts, aggregating results into structured JSON reports. You execute specialized judges concurrently to review implementation plans or code for adherence to best practices and test coverage.

What is LLM evaluation in automated testing and how does it work?

LLM evaluation in automated testing uses specialized AI judge agents to review software engineering artifacts against predefined schemas. It runs multiple judges concurrently, validating implementation plans and code outputs to ensure consistent and objective quality assessment.

How do I configure evaluation thresholds for AI judges assessing code artifacts?

Configure AI judge evaluation thresholds by using override files to customize assessment strictness. This allows you to adjust automated evaluation criteria for both implementation plans and code artifact types according to your quality requirements.

Can I use AI judges to evaluate software engineering plans before coding starts?

Yes, you can use AI judges to evaluate software engineering plans before coding starts by running 16 specialized plan evaluation agents concurrently. The judges assess plan quality and validate outputs against predefined schemas into a structured report.

What is the best way to validate code artifacts against predefined schemas?

The best way to validate code artifacts against predefined schemas is running parallel AI judge agents that automatically aggregate and validate outputs. This automated testing approach ensures consistent evaluation of code quality and artifact compliance using configurable thresholds.

Does automated plan evaluation support configurable thresholds for different artifact types?

Automated plan evaluation supports configurable thresholds for different artifact types via override files. You can customize evaluation strictness for both plan and code artifact types, validating outputs against predefined schemas before generating structured JSON reports.