judge

Score AI agent responses against standardized rubrics with weighted totals.

2|Updated Mar 25, 2026
One-click install
npx skills add https://github.com/slabgorb/sidequest --skill judge-slabgorb
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: judge
Source: https://github.com/slabgorb/sidequest/tree/main/.pennyfarthing/skills/pf-judge
Command: npx skills add https://github.com/slabgorb/sidequest --skill judge-slabgorb

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The Judge Skill provides a standardized framework to evaluate AI responses, ensuring objective, repeatable benchmarking across tasks and teams.

Core Features & Use Cases

  • Unified rubric evaluation across dimensions: Correctness, Depth, Quality, and Persona, for solo, comparison, and phase evaluations.
  • Deterministic scoring and transparent outputs: weighted totals, rubric anchors, and detailed breakdowns per mode for reproducibility and auditability.
  • Phase and comparison use cases: benchmark evaluation pipelines, side-by-side agent comparisons, and facilitator-mode evaluations for experiments and product tests.

Quick Start

Run an evaluation scenario by supplying contestant specs, a challenge description, and a sample response to generate scores and a structured evaluation narrative.

Frequently Asked Questions about judge

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI agent responses against standardized rubrics?

To evaluate AI agent responses against standardized rubrics, supply contestant specs, a challenge description, and a sample response to generate scores and a structured evaluation narrative. This enforces consistent dimensions like correctness, depth, quality, and persona for objective benchmarking.

What is the best way to benchmark AI responses for reproducibility?

The best way to benchmark AI responses for reproducibility is using a framework that outputs deterministic weighted totals, rubric anchors, and detailed per-mode breakdowns. This ensures repeatable and auditable evaluation across tasks and teams.

Can I compare two AI agents side-by-side using a standardized scoring system?

You can compare two AI agents side-by-side using the comparison evaluation mode, which scores responses against shared rubric dimensions. It generates weighted totals and detailed breakdowns to objectively highlight performance differences.

What dimensions should I score when evaluating prompt-engineering outputs?

When evaluating prompt-engineering outputs, you should score the dimensions of correctness, depth, quality, and persona. These standardized dimensions ensure comprehensive objective evaluation across benchmark and research workflows.

Does the AI evaluation framework support phase-based testing for experiments?

The AI evaluation framework supports phase-based testing for experiments through its phase evaluation mode. This mode facilitates facilitator-mode evaluations and product tests, generating structured narratives and weighted totals across multiple testing stages.