What problem does it solve?
Changing instruction wording (constitutions, skills, agent guidelines) without evidence is guesswork. This Skill provides a controlled, blinded evaluation workflow that measures whether instruction variants actually change agent behavior, producing evidence for a human decision instead of relying on prose review or self-reported compliance.
Core Features & Use Cases
- Blinded A/B Evaluation: Runs instruction variants under neutral labels with identical tasks, tools, and context so neither runner nor judge sees which variant is baseline or preferred.
- Behavior-Based Scoring: Scores observable artifacts (sources read, edits made, tests run, claims reported) against a precommitted rubric rather than trusting self-reported compliance.
- Advisory Synthesis: Produces an evidence-backed recommendation (retain, revise, combine, reject, or run a follow-up) without automatically mutating canonical instruction files.
- Use Case: You want to know whether a new wording in your agent's constitution actually reduces scope drift. This Skill designs an organic task, runs both variants blind, judges the outputs under neutral labels, and reports which wording earned its place.
Quick Start
Ask the agent to evaluate two instruction variants using the happier-instruction-eval workflow with a controlled blinded comparison.