What problem does it solve?
This Skill provides a structured workflow to create, refine, evaluate, and optimize Claude Skills. It helps users design new skills from scratch, iterate on existing skills, run evals, benchmark performance, and tune triggering descriptions for better discoverability.
Core Features & Use Cases
- End-to-end skill lifecycle: creation, modification, evaluation, benchmarking, and description optimization.
- Evaluation orchestration: run trigger evaluations across train/test sets, track results, and generate actionable feedback.
- Iterative improvement: propose description improvements based on results and history, with run-loop guidance.
- Artifacts & governance: maintain evals.json, history, and benchmark outputs to ensure reproducibility.
- Skill packaging and validation helpers: includes tooling to validate, package, and deploy skills.
Quick Start
Draft a new skill, write test prompts, run the evaluation loop, and iterate until you are satisfied, then optionally optimize the triggering description.