What problem does it solve?
This Skill eliminates the guesswork of creating consistent, reviewable evaluation frameworks for OPL Foundry Lab candidates (agents, skills, prompts, work orders) by providing a structured, authority-respecting design workflow that avoids overstepping ownership boundaries.
Core Features & Use Cases
- Structured Task Case Development: Build standardized test case sets covering happy path, boundary, negative, and regression scenarios to validate candidate behavior against expected outcomes.
- Scorecard & Failure Taxonomy Definition: Create clear pass/hold criteria and classify failures into standardized categories for consistent, unbiased review.
- Promotion/Hold Evidence Packaging: Generate standardized evidence packages that real owners can inspect to make informed promotion or hold decisions.
- Use Case: When testing a new grant writing agent, use this Skill to design test cases that verify it respects source boundaries, produces correctly shaped outputs, and does not make unauthorized readiness claims.
Quick Start
Use the opl-eval-harness-designer skill to design a complete evaluation harness for the new patent drafting agent, including all required task cases, scorecard, and failure taxonomy.