evaluation

Evaluate agent performance using multi-dimensional rubrics for accuracy and tool efficiency.

Updated Apr 13, 2026
One-click install
npx skills add https://github.com/Syedyasir001/rvu-LIBFLOW --skill evaluation-syedyasir001
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation
Source: https://github.com/Syedyasir001/rvu-LIBFLOW/tree/main/.agent/skills/library/evaluation
Command: npx skills add https://github.com/Syedyasir001/rvu-LIBFLOW --skill evaluation-syedyasir001

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables comprehensive evaluation of agent performance, ensuring quality, identifying issues, and improving agent systems.

Core Features & Use Cases

  • Multi-Dimensional Evaluation: Assess factual accuracy, completeness, citation accuracy, source quality, and tool efficiency.
  • Performance Drivers: Apply BrowseComp research findings for effective evaluation budgets and model upgrades.
  • Evaluation Rubric Design: Build multi-dimensional rubrics and map dimensions to numeric scores.
  • Continuous Evaluation: Integrate evaluation into the development workflow and monitor production quality.
  • Use Case: Create a rubric to evaluate the performance of an agent across various dimensions, continuously monitoring and improving its quality.

Quick Start

Use the evaluation skill to assess the performance of your agent system and compare different configurations or models.

Frequently Asked Questions about evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent performance across multiple dimensions like factual accuracy and tool efficiency?

To evaluate agent performance, you build a multi-dimensional rubric that scores factual accuracy, completeness, citation accuracy, source quality, and tool efficiency. This approach requires your agent outputs and test sets to conduct a thorough analysis.

What is multi-dimensional evaluation for AI agents?

Multi-dimensional evaluation is the process of assessing agent outputs using a rubric with specific quality dimensions. It maps dimensions like factual accuracy and tool efficiency to numeric scores, ensuring quality and identifying issues in agent systems.

How do I set up continuous evaluation to monitor agent quality in production?

You set up continuous evaluation by integrating the evaluation rubric into your development workflow. This monitors production quality continuously, applying performance drivers to optimize evaluation budgets and support model upgrades.

Can I compare different agent configurations or models using an evaluation rubric?

Yes, you can compare different configurations or models by applying a consistent multi-dimensional rubric to their outputs. This assesses factual accuracy, completeness, and tool efficiency across test sets to identify the optimal agent setup.

What do I need to assess factual accuracy and tool efficiency in my agent system?

You need your agent outputs and corresponding test sets to assess factual accuracy and tool efficiency. The evaluation applies a multi-dimensional rubric to map these quality dimensions to numeric scores for analysis.

When should I use a multi-dimensional rubric instead of basic agent testing?

Use a multi-dimensional rubric when basic agent testing cannot capture nuanced quality issues. It is needed when you must evaluate factual accuracy, citation accuracy, and tool efficiency simultaneously to identify performance drivers and optimize budgets.