What problem does it solve?
This Skill provides a comprehensive framework for evaluating AI agents, allowing users to conduct rigorous assessments of agent capabilities, quality, and reliability.
Core Features & Use Cases
- Multi-Grader Evaluation: Supports three types of graders: code-based, model-based, and human.
- Evaluation Types: Offers capability and regression evaluations with specific pass targets.
- Workflows: Integrates with various workflows such as running evaluations, comparing models and prompts, creating judges and use cases, running scenarios, and viewing results.
- Domain Patterns: Pre-configured for coding, conversational, research, and computer-use agent types.
- Integration: Works with THE ALGORITHM ISC rows for automated verification.
Quick Start
Run a comprehensive evaluation of the 'Research' skill using the following command:
bun run ~/.claude/skills/Evals/EvalServer/cli-run.ts --use-case Research