compare-runs

Compare two LongMemEval runs by scorecards, harness configurations, and per-question-type accuracy.

3|1|Updated Apr 11, 2026
One-click install
npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill compare-runs-tmuskal
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: compare-runs
Source: https://github.com/tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/compare-runs
Command: npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill compare-runs-tmuskal

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill allows users to easily compare two LongMemEval runs, providing insights into the differences in scorecards, harness configurations, and per-question-type accuracy.

Core Features & Use Cases

  • Side-by-Side Comparison: Visualize the differences between two runs at a glance.
  • Detailed Metrics: Analyze metrics like overall accuracy, user and assistant performance, and temporal reasoning.
  • Configurations and Scores: Compare harness configurations and scorecards for a comprehensive view.
  • Use Case: If you're experimenting with different models or settings, this Skill helps you identify which configuration yields the best results.

Quick Start

Compare two LongMemEval runs with the command: /compare-runs <runIdA> <runIdB>

Frequently Asked Questions about compare-runs

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I compare two LongMemEval runs side by side?

You can compare two LongMemEval runs by passing their run IDs to the comparison command, which analyzes scorecards, harness configurations, and per-question-type accuracy to highlight performance differences.

What metrics are included in a LongMemEval run comparison?

A LongMemEval run comparison analyzes overall accuracy, user and assistant performance, temporal reasoning, and per-question-type accuracy, alongside the harness configurations and scorecards from both runs.

How do I identify which model configuration yields the best results in LongMemEval?

By running a side-by-side comparison of different LongMemEval runs, you can analyze their respective scorecards and harness configurations to identify which model settings yield the highest accuracy.

Do I need access to run directories to compare LongMemEval scorecards?

Yes, comparing LongMemEval runs requires access to the respective run directories and a working knowledge of LongMemEval's data structure to successfully analyze the scorecards and harness configurations.

Can I compare harness configurations across different LongMemEval benchmarking runs?

Yes, the comparison process evaluates harness configurations alongside scorecards and accuracy metrics, providing a comprehensive view of how different settings affect LongMemEval run performance.

What is the best way to analyze per-question-type accuracy differences between LongMemEval runs?

The best way to analyze per-question-type accuracy is to use a side-by-side comparison tool that evaluates the scorecards from both LongMemEval runs to pinpoint specific performance variations.