google-agents-cli-eval

Evaluate ADK agent conversations, tool usage, and performance with custom metrics.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/aranlucas/agents --skill google-agents-cli-eval-aranlucas
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: google-agents-cli-eval
Source: https://github.com/aranlucas/agents/tree/main/.agents/skills/google-agents-cli-eval
Command: npx skills add https://github.com/aranlucas/agents --skill google-agents-cli-eval-aranlucas

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires agents-cli, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This skill provides a comprehensive toolset for evaluating the performance of ADK agents, addressing challenges related to agent quality, evaluation methodologies, and performance optimization.

Core Features & Use Cases

  • Agent Evaluation: Evaluates multi-turn agent conversations and responses.
  • Custom Metrics: Allows users to define custom metrics for specific evaluation scenarios.
  • User Simulation: Generates dynamic user scenarios for more realistic evaluations.
  • Optimization Tools: Offers optimization tools for enhancing agent performance.
  • Use Case: Evaluate the effectiveness of an agent in handling customer service inquiries by comparing its responses to a set of predefined scenarios.

Quick Start

Run the 'agents-cli eval generate' command to create evaluation datasets for the agent.

Frequently Asked Questions about google-agents-cli-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate multi-turn agent conversations for response quality?

Evaluating multi-turn agent conversations requires a framework that analyzes conversation quality, tool usage, and overall performance. You can assess agents against predefined scenarios using relevant evaluation datasets to measure effectiveness.

Can I define custom metrics for agent evaluation scenarios?

Yes, defining custom metrics for agent evaluation allows you to tailor assessments to specific scenarios. This capability helps measure unique conversation qualities and tool usage patterns relevant to your ADK agents.

How do I generate realistic user simulation scenarios for agent testing?

Generating realistic user simulation scenarios involves creating dynamic interactions for agent testing. This approach provides in-depth assessments by mimicking real user behaviors during the evaluation of ADK agents.

Do I need the agents-cli tool to run agent performance analysis?

Yes, you need the agents-cli tool to run agent performance analysis. It is a required dependency for executing evaluation commands and generating the necessary evaluation datasets for your ADK agents.

How do I create evaluation datasets for ADK agents?

Creating evaluation datasets for ADK agents involves running the agents-cli eval generate command. This produces the foundational data needed to compare agent responses against predefined scenarios.

What is the best way to optimize ADK agent performance after evaluation?

Optimizing ADK agent performance after evaluation involves using dedicated optimization tools. These tools help enhance overall agent quality based on the detailed analysis of conversation and tool usage metrics.