researcher-evaluation

Evaluate GenAI agents with G-Eval methodology and structured performance reports.

1|Updated Jan 23, 2026
One-click install
npx skills add https://github.com/VibeTechnologies/VibeTeam --skill researcher-evaluation
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: researcher-evaluation
Source: https://github.com/VibeTechnologies/VibeTeam/tree/main/.opencode/skills/researcher-evaluation
Command: npx skills add https://github.com/VibeTechnologies/VibeTeam --skill researcher-evaluation

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluates GenAI agents using the G-Eval methodology to produce structured performance reports.

Core Features & Use Cases

  • Provides a standardized Required Output Format for evaluating multiple frameworks (AutoGen, CrewAI, OpenHands)
  • Includes a comprehensive methodology (G-Eval) that scores accuracy, reasoning, and actionability across tasks
  • Offers a CLI workflow to run benchmarks and compare frameworks, with clear evaluation dimensions and a summary of results

Quick Start

Run the evaluation workflow on a sample task using the G-Eval methodology to compare agent outputs.

Frequently Asked Questions about researcher-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate GenAI agent performance across different frameworks?

You can evaluate GenAI agents by applying the G-Eval methodology to score accuracy, reasoning, and actionability. This Skill benchmarks agent performance across frameworks like AutoGen, CrewAI, and OpenHands to generate structured reports.

What is the G-Eval methodology for benchmarking GenAI agents?

The G-Eval methodology is a standardized evaluation process that scores GenAI agents on accuracy, reasoning, and feedback. It uses a defined scoring scale and evaluation dimensions to produce structured performance reports.

Can I use this to benchmark AutoGen against CrewAI?

Yes, you can benchmark AutoGen against CrewAI. The Skill provides a CLI workflow to run benchmarks and compare agent outputs across multiple frameworks using a standardized required output format.

How do I run a GenAI agent benchmark using a CLI workflow?

You can run a GenAI agent benchmark by using the provided CLI workflow to execute the G-Eval methodology on sample tasks. This process compares framework outputs and summarizes results based on defined evaluation dimensions.

What evaluation dimensions are scored when assessing GenAI agents?

The GenAI agent evaluation scores dimensions including accuracy, reasoning, and actionability. These metrics are calculated using the G-Eval methodology to provide a comprehensive summary of agent performance.

Do I need any external dependencies to generate GenAI agent evaluation reports?

No external dependencies are required to generate GenAI agent evaluation reports. The Skill operates independently to apply the G-Eval methodology and output structured performance summaries.