What problem does it solve?
This Skill helps you reliably evaluate agent performance by scoring outcomes across multiple quality dimensions, so you can detect regressions and improve pipelines instead of guessing.
Core Features & Use Cases
- Multi-dimensional evaluation: Score factual accuracy, completeness, citation accuracy, source quality, and tool efficiency separately instead of relying on a single number.
- Non-determinism-aware testing: Evaluate results against outcome criteria, not fixed execution paths, so agents can vary their routes while still being judged fairly.
- LLM-as-judge and human sampling: Use structured rubric prompts for scale, then supplement edge cases and subtle failures with human review.
- Continuous monitoring and quality gates: Track pass rates over time and block or alert on quality drops using thresholds.
Quick Start
Use this skill when you want to evaluate whether a new agent version meets your acceptance thresholds for quality and efficiency before deploying it.