What problem does it solve?
ACE's LLM-as-Judge evaluation system suffers from a systematic generosity bias that scores flawed artifacts 8–9/10 even when they contain confirmed real-world issues like factual errors, missing citations, or unmeasured adversarial gaps, producing untrustworthy quality assessments that fail to catch deployability problems.
Core Features & Use Cases
- Calibration Methodology: Provides a repeatable, auditable 6-step process for calibrating ACE's per-skill evaluation rubrics, including ground-truth catalogue building, detection rate tracking, multi-run variance testing, and dimension coverage checks.
- Bias Mitigation: Includes safeguards against common calibration failures like score anchoring, missing fitness dimensions, and rubric inflation that lead to false-pass quality gates.
- Use Case: If your ACE evaluation rubrics are scoring hollow or flawed app builds, chatbot transcripts, or other artifacts too high, use this skill to calibrate them to catch real issues and produce trustworthy, production-ready quality scores.
Quick Start
Use the eval-calibration skill to calibrate your ACE per-skill evaluation rubric against a ground-truth catalogue of known artifact flaws to eliminate LLM-as-Judge scoring bias and produce accurate quality assessments.