What problem does it solve?
Helps developers and product teams turn informal workflows or ideas into robust Claude-compatible skills by guiding skill drafting, test creation, automated evaluations, and iterative improvements so skills trigger reliably and produce measurable outcomes.
Core Features & Use Cases
- Skill authoring: Draft SKILL.md with clear metadata, triggering guidance, and progressive-disclosure instructions.
- Evaluation & benchmarking: Generate eval sets, run parallel with-skill and baseline comparisons, grade outputs, aggregate metrics, and produce human-friendly benchmark reports.
- Iteration & optimization: Automate description optimization, propose revisions based on quantitative results, package validated skills, and surface analyst suggestions for improvements.
- Use Case: Create a new code-generation helper, generate 3 realistic test prompts, run with/without-skill baselines, view results in the eval viewer, and iterate descriptions until trigger accuracy improves.
Quick Start
Tell the assistant what you want the skill to accomplish, ask it to draft SKILL.md and a small eval set, then run an evaluation loop to compare with a baseline and produce a review.