plan-mode-official-leaderboard-abctest

Compare three models on GSM8K, MATH, and ARC-Challenge with statistical tests.

Updated Oct 28, 2025
One-click install
npx skills add https://github.com/zapabob/SO8T --skill plan-mode-official-leaderboard-abctest
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: plan-mode-official-leaderboard-abctest
Source: https://github.com/zapabob/SO8T/tree/main/OpenCode_src/skills/plan_mode_official_leaderboard_abctest
Command: npx skills add https://github.com/zapabob/SO8T --skill plan-mode-official-leaderboard-abctest

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Plan-mode skill executes official leaderboard-compliant A/B/C tests to compare Phi-3.5-mini-instruct, Borea-phi3.5-instinct-jp, and AEGIS-Phi3.5mini-jpv2.4 on GSM8K, MATH, and ARC-Challenge, providing statistically validated results.

Core Features & Use Cases

  • Official benchmark evaluation across GSM8K, MATH, and ARC-Challenge for three models.
  • Statistical analysis including t-tests, effect sizes (Cohen's d), confidence intervals, and multiple-comparison corrections.
  • Result aggregation and leaderboard-style reporting for reproducible comparisons.
  • SO8T workflow integration and automated reporting.

Quick Start

Configure the test with your models and benchmarks, then run the OfficialABCTestPlan to generate a statistically validated leaderboard.

Frequently Asked Questions about plan-mode-official-leaderboard-abctest

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run an A/B/C test on GSM8K, MATH, and ARC-Challenge benchmarks?

Run an A/B/C test on GSM8K, MATH, and ARC-Challenge benchmarks by configuring your three target models and executing the OfficialABCTestPlan to generate a statistically validated leaderboard report automatically.

What statistical methods are needed to compare three language models on official leaderboards?

Comparing three models on official leaderboards requires t-tests, effect sizes (Cohen's d), confidence intervals, and multiple-comparison corrections to ensure the performance differences are statistically significant and reproducible.

Can I evaluate Phi-3.5-mini-instruct variants using a standardized benchmark harness?

Yes, you can evaluate Phi-3.5-mini-instruct variants like Borea-phi3.5 and AEGIS-Phi3.5mini using the official evaluation harness with predefined sample sizes and parallel evaluation across mathematical and reasoning datasets.

How do I generate a reproducible leaderboard report with confidence intervals and effect sizes?

Generate a reproducible leaderboard report by applying standardized benchmark evaluations and aggregating the results with automated reporting that includes confidence intervals, effect sizes, and multiple-comparison corrections.

Does the official benchmark evaluation support parallel evaluation across multiple datasets?

Yes, the official benchmark evaluation supports parallel evaluation across GSM8K, MATH, and ARC-Challenge using predefined sample sizes to produce statistically validated model comparisons efficiently.