opl-eval-harness-designer

Design structured evaluation harnesses for OPL Foundry Lab candidates.

8|5|Updated Apr 2, 2026
One-click install
npx skills add https://github.com/gaofeng21cn/one-person-lab --skill opl-eval-harness-designer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: opl-eval-harness-designer
Source: https://github.com/gaofeng21cn/one-person-lab/tree/main/plugins/opl-foundation-skills/skills/opl-eval-harness-designer
Command: npx skills add https://github.com/gaofeng21cn/one-person-lab --skill opl-eval-harness-designer

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill eliminates the guesswork of creating consistent, reviewable evaluation frameworks for OPL Foundry Lab candidates (agents, skills, prompts, work orders) by providing a structured, authority-respecting design workflow that avoids overstepping ownership boundaries.

Core Features & Use Cases

  • Structured Task Case Development: Build standardized test case sets covering happy path, boundary, negative, and regression scenarios to validate candidate behavior against expected outcomes.
  • Scorecard & Failure Taxonomy Definition: Create clear pass/hold criteria and classify failures into standardized categories for consistent, unbiased review.
  • Promotion/Hold Evidence Packaging: Generate standardized evidence packages that real owners can inspect to make informed promotion or hold decisions.
  • Use Case: When testing a new grant writing agent, use this Skill to design test cases that verify it respects source boundaries, produces correctly shaped outputs, and does not make unauthorized readiness claims.

Quick Start

Use the opl-eval-harness-designer skill to design a complete evaluation harness for the new patent drafting agent, including all required task cases, scorecard, and failure taxonomy.

Frequently Asked Questions about opl-eval-harness-designer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design an evaluation harness for AI agents?

Design an evaluation harness for AI agents by building structured task cases, defining scorecards with pass/hold criteria, and creating a failure taxonomy to classify behavioral outcomes for review.

What is a failure taxonomy in agent testing?

A failure taxonomy in agent testing is a standardized classification system that categorizes candidate failures into consistent groups, enabling unbiased review and clear promotion or hold decisions.

How do I create pass/hold criteria for skill promotion?

Create pass/hold criteria for skill promotion by defining a scorecard that validates candidate behavior against expected outcomes across happy path, boundary, negative, and regression test scenarios.

Can I use this to package promotion evidence for OPL Foundry candidates?

Yes, you can generate standardized promotion evidence packages for OPL Foundry candidates that real owners inspect to make informed readiness decisions without overstepping authority boundaries.

Does the evaluation harness design work for grant writing agents?

Yes, the evaluation harness design works for grant writing agents by verifying they respect source boundaries, produce correctly shaped outputs, and avoid making unauthorized readiness claims.

What are the limitations of designing agent test cases with this approach?

This approach is limited to source-only assessment and does not overstep owner authority or make readiness claims, meaning it produces reviewable evidence packages rather than final promotion decisions.