What problem does it solve?
Reduces the friction of authoring, testing, and iterating Claude skills by providing a repeatable workflow for drafting SKILL.md, generating evals, running with-skill and baseline comparisons, grading outputs, and optimizing the skill description for reliable triggering.
Core Features & Use Cases
- Drafting & Authoring: Guided structure and examples for writing SKILL.md frontmatter and operational instructions.
- Eval & Benchmarking: Create test prompts, run parallel with-skill and baseline runs, capture timing/tokens, grade outputs, and aggregate benchmark statistics.
- Iterate & Improve: Automate an improve loop that drafts assertions, analyzes results, updates the skill, and optimizes the skill description for better trigger accuracy.
- Utilities: Includes scripts for aggregating benchmarks, generating a review viewer, packaging skills, and description-optimization helpers.
Quick Start
Tell the assistant the task you want automated and ask it to draft SKILL.md, propose 2–3 realistic test prompts, run an initial evaluation comparing with-skill and without-skill, and summarize the results and next improvements.