eval

Evaluate and rank agent performance metrics in AgentHub sessions.

Updated Apr 2, 2026
One-click install
npx skills add https://github.com/4lerman/text_evaluator --skill eval-4lerman
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval
Source: https://github.com/4lerman/text_evaluator/tree/main/.agents/skills/engineering-advanced-skills/agenthub/skills/eval
Command: npx skills add https://github.com/4lerman/text_evaluator --skill eval-4lerman

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

The eval Skill solves the problem of efficiently ranking and evaluating agent results in an AgentHub session, ensuring clear and fair performance assessments.

Core Features & Use Cases

  • Multi-Mode Evaluation: Offers both metric-based evaluation and LLM judge mode to suit different evaluation needs.
  • Automated Ranking: Provides automated ranking of agents based on their performance metrics or the outcomes of an LLM judge comparison.
  • Hybrid Mode: Offers a hybrid approach for more comprehensive evaluation by combining both metric and LLM judge rankings.
  • Session-Specific: Allows evaluation of specific sessions within the AgentHub platform, enabling precise performance analysis.

Quick Start

Run the eval Skill by entering the command /hub:eval to evaluate the latest session.

Frequently Asked Questions about eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent performance in an AgentHub session?

You can evaluate agent performance in an AgentHub session by running the eval Skill, which automates performance evaluation and ranking based on your session metrics.

Can I use an LLM judge to rank agent results instead of metrics?

Yes, you can use the LLM judge mode to rank agent results. This Skill supports both metric-based evaluation and LLM judge comparisons to suit different performance assessment needs.

How do I combine metric analysis and LLM judgment for agent ranking?

To combine metric analysis and LLM judgment for agent ranking, use the hybrid evaluation mode. This approach integrates both metric-based and LLM judge rankings to provide a more comprehensive evaluation.

Does the eval Skill require any specific external dependencies?

No, the eval Skill does not require any external dependencies. It operates entirely within the AgentHub environment to provide an efficient and accurate evaluation process.

What is the best way to automate agent ranking for my latest session?

The best way to automate agent ranking for your latest session is to enter the command `/hub:eval`. This triggers the evaluation process and generates a clear ranking system automatically.