skill-eval

Run automated benchmarks to measure pass rates, token usage, and timing for skills.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/reiserwang/Coding_Agent --skill skill-eval-reiserwang
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-eval
Source: https://github.com/reiserwang/Coding_Agent/tree/main/.gemini/skills/skill-eval
Command: npx skills add https://github.com/reiserwang/Coding_Agent --skill skill-eval-reiserwang

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

Evaluate, benchmark, compare, and recursively improve any skill across .gemini/skills/ and .claude/skills/. This ensures you can measure quality (pass rate, token usage, execution time), run blind A/B comparisons between skill versions, and progressively optimize instructions for better triggering.

Core Features & Use Cases

  • Automated evaluation workflow: generate test cases, run with skill and baseline, and collect metrics.
  • Benchmarking and comparison: aggregate pass rates, tokens, and durations; supports blind comparator and iteration logs.
  • Iterative improvement: optimize skill descriptions and prompts based on feedback; persistence of lessons in a memory workspace.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Quick Start

Run skill-eval:generate to create test cases, then run skill-eval:run to execute them and gather metrics.

Frequently Asked Questions about skill-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark and compare AI agent skills across different versions?

To benchmark AI agent skills, run automated workflows that record pass rates, token usage, and execution timing. You can execute blind A/B comparisons between skill versions to measure quality and progressively optimize instructions for better triggering.

How do I evaluate pass rates and token usage for multi-agent workflows?

Evaluating pass rates and token usage for multi-agent workflows involves running automated benchmarks that capture structured artifacts like evals.json. This enables deterministic evaluation by quantifying skill quality and aggregating performance metrics for analysis.

Can I use automated evaluation workflows to improve prompts in .gemini/skills/ and .claude/skills/?

Yes, you can evaluate, compare, and iteratively improve any skill across .gemini/skills/ and .claude/skills/. The workflow generates test cases, runs baselines, collects metrics, and optimizes skill descriptions based on feedback.

What is the best way to generate test cases and run benchmarks for skill evaluation?

The best way to generate test cases and run benchmarks is to execute a two-step workflow: first generate test cases, then run them against the skill and baseline to gather metrics like execution time and pass rates for iterative improvement.

How does iterative improvement work when optimizing skill descriptions and prompts?

Iterative improvement works by applying automated benchmarking feedback to optimize skill descriptions and prompts. It persists lessons in a memory workspace, allowing you to recursively enhance instruction triggering and compare aggregated metrics across versions.

What metrics are captured when running blind A/B comparisons for skill quality?

When running blind A/B comparisons for skill quality, the workflow captures pass rates, token usage, and execution durations. These metrics are aggregated into structured artifacts like evals.json and timing benchmarks to enable deterministic evaluation.