agent-evaluation

Evaluate AI agents using a 5-dimension rubric and multiple assessment methods.

145|28|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/guia-matthieu/clawfu-skills --skill agent-evaluation-guia-matthieu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agent-evaluation
Source: https://github.com/guia-matthieu/clawfu-skills/tree/main/skills/meta/agent-evaluation
Command: npx skills add https://github.com/guia-matthieu/clawfu-skills --skill agent-evaluation-guia-matthieu

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a structured framework for evaluating the quality, consistency, and reliability of AI agents and skills, ensuring they meet performance standards before deployment and during operation.

Core Features & Use Cases

  • Systematic Evaluation: Offers multiple methods (Direct Scoring, LLM-as-Judge, Pairwise Comparison, Pressure Testing) for assessing AI performance.
  • Bias Detection & Mitigation: Identifies and provides strategies for common AI evaluation biases like position bias and verbosity bias.
  • Use Case: Before deploying a new customer support agent, use this skill to run a battery of tests against predefined criteria, ensuring it follows instructions accurately and provides coherent, complete responses.

Quick Start

Use the agent-evaluation skill to evaluate the 'customer-support-agent' using the 5-dimension rubric.

Frequently Asked Questions about agent-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI agent performance before deployment?

You can evaluate AI agent performance using a 5-dimension rubric and multiple assessment methodologies. This framework requires structured test cases and clear evaluation criteria to systematically measure quality, consistency, and reliability before operation.

What is the best way to test AI skills for instruction accuracy?

The best way to test AI skills for instruction accuracy is through pressure testing and direct scoring against a predefined rubric. This identifies whether the agent follows instructions accurately and provides coherent, complete responses.

How does LLM-as-Judge work for quality assurance?

LLM-as-Judge works for quality assurance by using an AI model to evaluate the outputs of another agent. It serves as one of multiple systematic assessment methodologies available to measure performance and ensure reliability.

Can I detect position bias and verbosity bias in AI agents?

Yes, you can detect position bias and verbosity bias in AI agents. The evaluation framework specifically identifies these common AI biases and provides mitigation strategies to ensure accurate and fair performance monitoring.

Do I need structured test cases to run agent evaluation?

Yes, structured test cases are required to run agent evaluation. Accurate assessment of AI agents demands structured test cases and clear evaluation criteria to properly measure performance against the 5-dimension rubric.

How do I compare two AI agents using pairwise comparison?

Pairwise comparison evaluates two AI agents by directly comparing their outputs against each other. This methodology helps determine which agent provides higher quality, more consistent, and reliable responses within the evaluation framework.