What problem does it solve?
Helps users create, test, and iteratively improve Claude Code skills by providing a structured authoring workflow, automated evaluation/benchmarking tools, and utilities to optimize skill descriptions so the right skill triggers in real sessions. It removes the repetitive, error-prone steps of drafting SKILL.md, designing evals, running with-skill vs baseline comparisons, grading results, and packaging a final skill.
Core Features & Use Cases
- Skill authoring: Guided process for drafting SKILL.md, capturing intent, and writing examples and test prompts.
- Automated evaluation & benchmarking: Run parallel with-skill and baseline runs, capture timing/tokens, grade expectations, and aggregate benchmark summaries.
- Analysis & improvement: Tools to generate a reviewer UI, run blind comparisons, analyze why one version won, and iteratively improve descriptions to boost triggering accuracy.
- Packaging & deployment: Validate and package skills for distribution and include helper scripts for common workflows.
Quick Start
Ask the skill-creator to draft a SKILL.md for a new skill, produce 2–3 realistic test prompts, run the eval loop, and return a benchmark summary and suggested description improvements.