plan-mode_arc_gsm8k_improvement

Analyze ARC-Challenge and GSM8K evaluation failures for timeout, extraction, contamination, and seed stability.

Updated Oct 28, 2025
One-click install
npx skills add https://github.com/zapabob/SO8T --skill plan-mode-arc-gsm8k-improvement
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: plan-mode_arc_gsm8k_improvement
Source: https://github.com/zapabob/SO8T/tree/main/OpenCode_src/skills/plan_mode_arc_gsm8k_improvement
Command: npx skills add https://github.com/zapabob/SO8T --skill plan-mode-arc-gsm8k-improvement

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Plan mode identifies and fixes weaknesses in ARC-Challenge and GSM8K evaluations for the AEGIS model by analyzing timeout rates, extraction failures, data contamination, and seed stability.

Core Features & Use Cases

  • ARC-Challenge improvement analysis: timeout rate, extraction failure analysis, response-pattern analysis, robust extraction implementation.
  • GSM8K sanity checks: data contamination detection, multi-seed evaluation, zero-shot evaluation, scoring logic validation.
  • Multi-objective evaluation workflow: parallel evaluations, statistical validation, comparative analysis, automated report generation.
  • SO8T integration optimizations: existing ABC test integration, checkpoint management, resource optimization, automatic improvement proposals.

Quick Start

Run the ARC/GSM8K improvement Plan with your model path and seeds to generate analysis and immediately review and apply the top recommended improvements.

Frequently Asked Questions about plan-mode_arc_gsm8k_improvement

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze data contamination and timeout failures in ARC-Challenge evaluations?

You can analyze data contamination and timeout failures in ARC-Challenge evaluations by running a plan-mode workflow that inspects response patterns, extraction failures, and timeout rates to generate robust extraction improvements.

How do I run multi-seed GSM8K sanity checks for zero-shot evaluation?

To run multi-seed GSM8K sanity checks for zero-shot evaluation, execute the improvement plan with your model path and seeds to validate scoring logic and detect data contamination across parallel runs.

What is the best way to ensure reproducible benchmark improvements across ARC and GSM8K datasets?

The best way to ensure reproducible benchmark improvements across ARC and GSM8K datasets is to apply plan-mode execution with parallel multi-seed testing, statistical validation, and automated report generation.

Can I automate report generation for multi-metric benchmark evaluation analysis?

Yes, you can automate report generation for multi-metric benchmark evaluation analysis by utilizing the integrated workflow that performs comparative analysis, statistical validation, and automatic improvement proposals.

Why does my benchmark evaluation pipeline suffer from seed instability and data drift?

Your benchmark evaluation pipeline suffers from seed instability and data drift when multi-seed configurations lack statistical validation, requiring automated data-drift checks and seed stability analysis to correct.

Does this evaluation workflow support existing test integration and checkpoint management?

Yes, this evaluation workflow supports existing ABC test integration, checkpoint management, and resource optimization to provide SO8T integration optimizations and automatic improvement proposals.