eval-harness

Evaluate agent accuracy, efficiency, alignment, and quality across sessions.

7|1|Updated Feb 5, 2026
One-click install
npx skills add https://github.com/besync-labs/antigravity-ai-kit --skill eval-harness-besync-labs
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-harness
Source: https://github.com/besync-labs/antigravity-ai-kit/tree/main/.agent/skills/eval-harness
Command: npx skills add https://github.com/besync-labs/antigravity-ai-kit --skill eval-harness-besync-labs

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluation of agent performance across accuracy, efficiency, alignment, and quality to guide improvements.

Core Features & Use Cases

  • Evaluation Dimensions: Accuracy, Efficiency, Alignment, and Quality with concrete questions to assess behavior.
  • Evaluation Metrics: Target thresholds (e.g., first-time success rate >80%, iterations <3, test coverage >80%, build 100%).
  • Report Format: Generates structured evaluation reports and learnings for future sessions.
  • Integration: Run at session end for learning and continuous improvement.

Quick Start

Run the evaluation harness at the end of a session to measure agent performance on the selected task.

Frequently Asked Questions about eval-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I measure AI agent performance across different tasks?

Measuring AI agent performance involves applying an evaluation harness to quantify accuracy, efficiency, alignment, and quality across standardized testing scenarios, enabling repeatable metrics for cross-task comparisons and continuous improvement.

What metrics should I use for agent evaluation and benchmarking?

Agent evaluation metrics should include target thresholds like first-time success rate over 80%, iterations under 3, test coverage over 80%, and successful builds. These standardized benchmarks quantify performance across accuracy, efficiency, alignment, and quality dimensions.

How do I generate standardized reports for AI testing sessions?

Generate standardized reports for AI testing by running an evaluation harness at the end of a session. This produces structured evaluation reports capturing performance learnings and metrics to guide future improvements and continuous quality assurance.

Can I use an evaluation harness to compare agent performance across sessions?

Yes, you can use an evaluation harness to compare agent performance across sessions. It standardizes testing scenarios with defined dimensions and target metrics, producing repeatable results that enable direct cross-task comparisons and track continuous improvement over time.

When should I run agent performance evaluation in my workflow?

Run agent performance evaluation at the end of a session. Integrating the evaluation harness at this stage captures structured learnings and generates reports for continuous improvement without disrupting active task execution.