What problem does it solve?
This Skill provides a structured, multi-phase framework for rigorously evaluating AI agents, ensuring comprehensive testing and evidence-based reporting.
Core Features & Use Cases
- Structured Evaluation Pipeline: Guides users through planning, test data generation, execution, analysis, and documentation.
- Automated Test Execution: Implements and runs evaluation pipelines with real agent interactions.
- Evidence-Based Reporting: Generates detailed reports with actionable insights and prioritized recommendations.
- Use Case: A team developing a new customer service AI agent can use this Skill to systematically test its performance against predefined metrics, identify critical issues, and document improvements before deployment.
Quick Start
Use the eval skill to start the evaluation process for the agent located at './my-agent'.