What problem does it solve? Teams shipping prompts, agents, or workflows often lack a disciplined way to compare versions, detect regressions, and decide whether a candidate is release-ready, leading to decisions based on vibes rather than evidence. ## Core Features & Use Cases - Evaluation Design: Select the right evaluation surface, define contract-first acceptance criteria, and build metric taxonomies tied to actual decisions. - Regression and Drift Analysis: Compare baseline versus candidate behavior, detect semantic drift, coverage regression, and anomalies without being misled by noisy single points. - Release Gating: Define gates with owners, thresholds, and actions for rollout, canary continuation, and prompt-version promotion. - Use Case: When comparing two prompt versions before rollout, use this Skill to build a scorecard with weighted criteria, identify critical-cohort failures, and produce a release recommendation with explicit evidence and unresolved risks. ## Quick Start Use the v14-eval-ops skill to compare the baseline and candidate prompt versions and produce a release-readiness scorecard with regression findings.