opi-eval

Run end-to-end regression tests on the opi runtime and analyze fidelity.

5|2|Updated May 19, 2026
One-click install
npx skills add https://github.com/OdradekAI/opi --skill opi-eval
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: opi-eval
Source: https://github.com/OdradekAI/opi/tree/main/.claude/skills/opi-eval
Command: npx skills add https://github.com/OdradekAI/opi --skill opi-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill assesses the fidelity and performance of the opi runtime by executing comprehensive regression tests against real LLM providers, collecting runtime traces, and dispatching an independent evaluator subagent for detailed analysis.

Core Features & Use Cases

  • Regression Testing: Compiles the opi binary and runs structured test cases to detect fidelity degradation.
  • Independent Evaluation: An evaluator subagent analyzes test case definitions, runtime signals, and evaluation dimensions.
  • Reporting: Generates structured extraction and detailed report files with pass rates, metrics, and trends.

Quick Start

Run the opi-eval skill with the default configuration to evaluate the opi runtime and detect fidelity degradation.

Frequently Asked Questions about opi-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run regression testing to evaluate AI runtime fidelity?

Regression testing for runtime fidelity requires compiling the opi binary and executing structured test cases against real LLM providers to detect performance degradation. An independent evaluator subagent then analyzes runtime signals and test definitions across multiple dimensions.

What is end-to-end runtime assessment for AI tools?

End-to-end runtime assessment executes structured test cases against real LLM providers to measure fidelity. It collects runtime traces and dispatches an evaluator subagent to analyze performance metrics, generating structured reports with pass rates and trends.

How do I detect fidelity degradation in an AI runtime?

Detect fidelity degradation by running structured regression tests against real LLM providers. The evaluation process collects runtime traces and generates detailed report files containing pass rates and metrics to identify performance drops.

Do I need an evaluator subagent for runtime assessment?

Yes, comprehensive runtime assessment requires an evaluator subagent to analyze test case definitions, runtime signals, and evaluation dimensions. The opi binary and structured test cases are also necessary to execute the regression testing process.

What does an AI tool regression test report include?

A regression test report includes structured extraction files with pass rates, metrics, and performance trends. The evaluation results document fidelity assessment across multiple dimensions based on runtime signals collected during test case execution.

Does runtime fidelity evaluation work with real LLM providers?

Yes, runtime fidelity evaluation executes structured test cases directly against real LLM providers. This approach captures authentic runtime traces and signals for the evaluator subagent to analyze, ensuring accurate regression and fidelity assessment.