google-agents-cli-eval

Evaluate and optimize AI agents with Python-based metrics and the Agent Platform.

Updated Jun 15, 2026
One-click install
npx skills add https://github.com/ironkid90/lucky5-v8 --skill google-agents-cli-eval-ironkid90
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: google-agents-cli-eval
Source: https://github.com/ironkid90/lucky5-v8/tree/main/plugins/skills/google-agents-cli-eval
Command: npx skills add https://github.com/ironkid90/lucky5-v8 --skill google-agents-cli-eval-ironkid90

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires agents-cli, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a comprehensive set of tools and metrics for evaluating and optimizing agent performance, enabling detailed analysis and iterative improvement.

Core Features & Use Cases

  • Evaluation Metrics: Access a wide range of built-in and custom metrics to measure agent performance across various dimensions.
  • Data Synthesis: Generate synthetic user scenarios for testing agents in different conversational contexts.
  • Optimization Tools: Utilize tools like prompt optimization to enhance agent performance based on specific metrics.
  • Use Case: Use this Skill to analyze the performance of an agent in handling customer inquiries, ensuring it meets quality standards and efficiency goals.

Quick Start

Run 'agents-cli eval generate' to generate traces of your agent's performance on a given dataset.

Frequently Asked Questions about google-agents-cli-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate AI agent performance using detailed metrics?

To evaluate AI agent performance, you can use built-in and custom evaluation metrics to measure behavior across various dimensions. This Skill provides a suite of Python-based tools to generate traces and analyze your agent's performance on a given dataset.

Can I generate synthetic datasets to test my LLM agent?

Yes, you can generate synthetic datasets to test your LLM agent. The data synthesis feature creates synthetic user scenarios, allowing you to test agents in different conversational contexts before deployment.

What is the best way to optimize agent prompts based on evaluation results?

The best way to optimize agent prompts is by using dedicated optimization tools that enhance performance based on specific metrics. This allows for iterative agent development and targeted performance analysis after running initial evaluations.

Do I need google-agents-cli to run agent evaluations?

Yes, you need the google-agents-cli dependency installed to execute agent evaluations and optimizations. This Skill requires the agents-cli environment to generate traces and apply performance metrics to your AI agents.

What are the limitations when analyzing agent performance with synthetic data?

When analyzing agent performance with synthetic data, limitations may include scenarios that do not fully capture real-world conversational complexity. Use the optimization tools iteratively to refine metrics and ensure the agent meets actual quality standards.