diyu-eval-harness

Evaluate AI behavior across governance, capability, and regression for eight registered objects.

Updated Jul 11, 2026
One-click install
npx skills add https://github.com/andyan77/diyu-agent --skill diyu-eval-harness
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: diyu-eval-harness
Source: https://github.com/andyan77/diyu-agent/tree/main/.claude/skills/diyu-eval-harness
Command: npx skills add https://github.com/andyan77/diyu-agent --skill diyu-eval-harness

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill provides a structured framework to identify and quantify AI behavior across governance, capability, and regression for eight registered objects, enabling objective evaluation and traceable decision-making.

Core Features & Use Cases

  • Governance evaluation: verify presence and integrity of governance artifacts and ensure alignment with upstream references.
  • Capability evaluation: assess declared capabilities against actual behavior for active objects.
  • Regression evaluation: compare current results against a baseline to detect degradation or improvement.
  • Scorecard generation: compute a 6-dimension, 10-point scorecard and emit both human-readable reports and machine-readable artifacts.
  • Evidence and auditing: produce Markdown reports and YAML scorecards suitable for governance reviews and archives.

Quick Start

Invoke the evaluation harness to run governance, capability, and regression checks across all registered objects and generate the 6-dimension scorecard outputs.

Frequently Asked Questions about diyu-eval-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate AI behavior evaluation for governance and capability checks?

AI behavior evaluation is automated by running a harness that identifies and quantifies governance, capability, and regression metrics across registered objects, producing scorecards and traceable decision artifacts.

What is a governance-focused AI capability evaluation and when do I need it?

A governance-focused AI capability evaluation verifies the presence and integrity of governance artifacts and ensures alignment with upstream references. You need it when assessing declared capabilities against actual behavior for objective auditing.

How do I generate a 10-point scorecard for AI capability and regression assessments?

To generate a 10-point scorecard, invoke the evaluation harness to run checks across all registered objects, computing a six-dimension scorecard that emits both Markdown reports and YAML artifacts.

Does the evaluation harness export machine-readable artifacts for compliance review?

Yes, the evaluation harness exports machine-readable YAML scorecards alongside human-readable Markdown reports, providing evidence and auditing artifacts suitable for governance reviews and archival compliance.

Can I detect AI behavior degradation using regression evaluation against a baseline?

Yes, you can detect degradation or improvement by running a regression evaluation that compares current AI behavior results against an established baseline across your registered objects.

What is the best way to evaluate eight registered AI objects for governance alignment?

The best way to evaluate eight registered AI objects is to invoke the end-to-end evaluation harness, which executes governance, capability, and regression workflows and generates a six-dimension, 10-point scorecard for review.