evaluation

Assess agent performance with weighted rubrics and pass/fail thresholds.

3|Updated Mar 22, 2026
One-click install
npx skills add https://github.com/0xharryriddle/codex-field-kit --skill evaluation-0xharryriddle
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/0xharryriddle/codex-field-kit/tree/main/archive/upstream/chasebuild-agent-skills/context-engineering/skills/evaluation
Command: npx skills add https://github.com/0xharryriddle/codex-field-kit --skill evaluation-0xharryriddle

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Evaluation of agent systems is complex due to non-determinism and multi-model interactions; this skill provides a repeatable framework for measuring performance, validity, and safety.

Core Features & Use Cases

  • Multi-dimensional rubrics covering factual accuracy, completeness, citation accuracy, source quality, and tool efficiency.
  • Test-set management and production monitoring to track improvements and detect regressions.
  • Flexible scoring and reporting with weighted dimensions and pass/fail thresholds.

Quick Start

Run a quick evaluation by feeding a sample task and its expected outcome into the framework to obtain a baseline score.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent performance systematically in production?

Agent evaluation uses structured rubrics and multi-dimensional scoring to measure factual accuracy, completeness, citation accuracy, source quality, and tool efficiency across test sets and production monitoring, ensuring repeatable scoring with weighted dimensions and pass/fail thresholds.

What metrics should I track for multi-model agent evaluation?

Multi-dimensional agent evaluation tracks factual accuracy, completeness, citation accuracy, source quality, and tool efficiency. It applies weighted rubrics across test sets and production monitoring to detect regressions and ensure repeatable scoring with pass/fail thresholds.

How do I create a baseline score for my AI agent test sets?

To establish a baseline score for agent test sets, feed a sample task and its expected outcome into the evaluation framework to obtain a baseline score using multi-dimensional rubrics covering factual accuracy, completeness, and tool efficiency.

Can I use weighted rubrics for production monitoring of non-deterministic agents?

Weighted rubrics support production monitoring for non-deterministic agents by providing a repeatable framework to measure performance, validity, and safety. They apply multi-dimensional scoring across test sets to track improvements and detect regressions with pass/fail thresholds.

What is the best way to measure citation accuracy and source quality in agent responses?

Measuring citation accuracy and source quality in agent responses requires multi-dimensional evaluation rubrics. These rubrics apply weighted scoring and pass/fail thresholds across test sets and production monitoring to ensure repeatable scoring for continuous improvement.