What problem does it solve?
This Skill helps build and evaluate agent evaluation systems, addressing non-determinism, multi-dimensional rubrics, and continuous monitoring for agent pipelines.
Core Features & Use Cases
- Deterministic Checks: Validates schema, duplicates, and rubric math before LLM judgment.
- Multi-Dimensional Rubrics: Scores factual accuracy, completeness, citation accuracy, source quality, and tool efficiency.
- Evaluation Methodologies: Provides guidelines for LLM-as-judge, human evaluation, and end-state evaluation.
- Test Set Design: Offers guidance on selecting representative samples and stratifying by complexity.
- Context Engineering Evaluation: Validates context strategies and runs degradation tests.
- Continuous Evaluation: Integrates evaluation into the development workflow and monitors production quality.
Quick Start
Run the evaluate skill on the agent response and compare it to the expected output to determine the quality of the response.