cross-harness

Automate LongMemEval benchmarking across Codex, Gemini, and OpenCode harnesses.

3|1|Updated Apr 11, 2026
One-click install
npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill cross-harness
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: cross-harness
Source: https://github.com/tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/cross-harness
Command: npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill cross-harness

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill automates the benchmarking of LongMemEval across different harnesses (Codex, Gemini, OpenCode) and integrates their results into a single scorecard.

Core Features & Use Cases

  • Cross-Harness Benchmarking: Execute LongMemEval tasks on multiple harnesses simultaneously.
  • Result Aggregation: Combine and compare results from different harnesses in a unified format.
  • Use Case: If you want to compare the performance of your model across Codex, Gemini, and OpenCode, this Skill can help you automate the process and generate a comprehensive report.

Quick Start

Use the cross-harness skill to benchmark your model across different harnesses with the command: /cross-harness:run <action> --harness <harness> --run-id <run-id>.

Frequently Asked Questions about cross-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is the best way to generate a single scorecard for cross-harness benchmarking?

The best way to generate a single scorecard for cross-harness benchmarking is to use the cross-harness Skill, which simultaneously executes LongMemEval tasks across Codex, Gemini, and OpenCode and integrates their results.

How do I benchmark LongMemEval tasks across Codex, Gemini, and OpenCode?

To benchmark LongMemEval across Codex, Gemini, and OpenCode, you can use the cross-harness Skill to automate task execution and aggregate the results into a single comprehensive scorecard for comparison.

How do I automate model performance testing across multiple harnesses?

Automate model performance testing by running the cross-harness Skill with a command specifying your action, target harness, and run identifier to execute and compare NLP evaluation results automatically.

Can I compare QA system results from different harnesses in a unified format?

Yes, you can compare QA system results in a unified format by using the cross-harness Skill to aggregate outputs from Codex, Gemini, and OpenCode into a single integrated scorecard.

Do I need Python scripts to run LongMemEval benchmark comparisons?

Yes, you need Python scripts to execute the cross-harness Skill and carry out the LongMemEval benchmark comparisons across the different Codex, Gemini, and OpenCode harnesses.

What is the best way to generate a single scorecard for cross-harness benchmarking?

The best way to generate a single scorecard for cross-harness benchmarking is to use the cross-harness Skill, which simultaneously executes LongMemEval tasks across Codex, Gemini, and OpenCode and integrates their results.