eval-harness

Benchmark pixl-crew skills, agents, and prompts with configurable evaluation runs.

2|Updated Mar 16, 2026
One-click install
npx skills add https://github.com/hamzaPixl/pixl-ai --skill eval-harness-hamzapixl
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-harness
Source: https://github.com/hamzaPixl/pixl-ai/tree/main/packages/crew/skills/eval-harness
Command: npx skills add https://github.com/hamzaPixl/pixl-ai --skill eval-harness-hamzapixl

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This evaluation harness provides a structured, repeatable framework to assess pixl-crew skills, agents, and prompts, enabling objective quality metrics and regression tracking.

Core Features & Use Cases

  • Automated, configurable evaluations (capability and regression) across versions.
  • Generation of structured eval reports and dashboards to surface regressions and performance trends.
  • Test-case discovery and rubric-driven scoring to guide ongoing skill/agent/prompt improvement.

Quick Start

Run pixl with the eval-harness target to start a capability or regression evaluation.

Frequently Asked Questions about eval-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent and prompt quality for automated workflows?

You benchmark agent and prompt quality by running an evaluation harness that applies structured criteria and rubric-driven scoring to generate per-test reports tracking regressions over time.

What is regression testing for AI prompts and how does it work?

Regression testing for AI prompts runs configurable evaluations across versions to validate capability, generating per-test reports that track performance trends and surface regressions over time using structured criteria.

How do I benchmark skills across multiple versions to track performance?

Benchmark skills across versions by executing a regression evaluation with a configurable run count, enforcing structured criteria to generate per-test reports that track performance trends over time.

Can I export evaluation reports for skills and prompts?

Yes, the evaluation harness supports results export, generating structured per-test reports that surface regressions and performance trends from your configurable capability and regression test runs.

Do I need to define test cases manually to measure prompt quality?

No, the evaluation harness features automated test-case discovery, using rubric-driven scoring and structured criteria to measure prompt quality and guide ongoing skill, agent, and prompt improvement.

When should I use a repeatable evaluation harness for testing agents?

Use a repeatable evaluation harness when you need objective quality metrics and regression tracking for agents and prompts, enabling capability and regression testing to surface regressions over time.