skill-evals-optimize

Triage failing OpenCode skill-loading eval cases with minimal reversible fixes.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/chandima/opencode-config --skill skill-evals-optimize
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-evals-optimize
Source: https://github.com/chandima/opencode-config/tree/main/.codex/skills/skill-evals-optimize
Command: npx skills add https://github.com/chandima/opencode-config --skill skill-evals-optimize

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Triages failing OpenCode skill-loading eval cases and applies minimal fixes to restore progress.

Core Features & Use Cases

  • Identify the latest failed evals and isolate root causes.
  • Apply small, reversible changes to prompts, descriptions, or tests to improve pass rates.
  • Retest affected cases with a strict 2-iteration cap and report outcomes.

Quick Start

Identify the latest failed evals, implement a minimal, reversible fix, and re-run targeted evals up to two iterations.

Frequently Asked Questions about skill-evals-optimize

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage failing OpenCode eval cases automatically?

Triage failing OpenCode eval cases by identifying the latest failures, isolating root causes, and applying minimal, reversible changes to prompts or tests.

What is the iteration limit for fixing failing eval datasets?

The iteration limit for fixing failing eval datasets is capped at two iterations to enforce strict safeguards and prevent broad, uncontrolled permission changes.

How do I apply minimal fixes to restore progress on failed skill-loading evals?

Apply small, reversible changes to prompts, descriptions, or tests to restore progress, then retest affected cases and report clear pass/fail summaries.

Can I use this triage workflow across multiple eval datasets?

Yes, the triage workflow applies across multiple eval datasets, focusing on isolating root causes and implementing targeted fixes for skill-loading evaluations.

Why does the eval optimization process prevent broad permission changes?

The eval optimization process prevents broad permission changes to enforce safeguards, ensuring only minimal, reversible modifications are applied during triage.

What safeguards are in place when optimizing failing OpenCode evals?

Safeguards include capping iterations at two, preventing broad permission changes, requiring minimal reversible modifications, and reporting clear pass/fail summaries.