What problem does it solve? Teams shipping prompts, agents, and workflows often lack a disciplined way to compare candidates, detect regressions, and decide release readiness, leading to overstated claims and unsafe promotions. ## Core Features & Use Cases - Evaluation Design: Defines evaluation surfaces, metric taxonomies, scorecards, and rubrics tied to actual decisions and gates. - Regression & Drift Review: Compares baseline vs candidate under comparable conditions, detects semantic drift, coverage regression, and anomalies. - Release Gating: Enforces gates with owner, threshold, and action, including harness-grade labeling from harness-designed to production-monitored. - Use Case: When reviewing a new coding prompt package, run the required case mix (features, bug fixes, security, prompt injection), keep executed-vs-unexecuted status explicit, and produce a pass/hold/reject verdict with evidence. ## Quick Start Ask the agent to evaluate whether the attached candidate prompt package is release-ready compared to the current baseline, with a scorecard and gate recommendation.