What problem does it solve?
Create, refine, and measure the effectiveness of Claude skills. This Skill provides a structured workflow to draft new skills, iterate on their descriptions, run evaluation loops, and benchmark performance to improve triggering accuracy and reliability.
Core Features & Use Cases
- Skill creation: Initialize new skills from scratch based on user needs.
- Iterative improvement: Edit SKILL.md, run evals, and apply feedback to descriptions and prompts.
- Evaluation & benchmarking: Automate trigger evaluations, collect metrics, and compare configurations (with_skill vs without_skill).
- Documentation tooling: Leverage bundled scripts for report generation, description improvement, and packaging.
- Use case: A team wants to prototype a new skill and rapidly iterate on its trigger phrasing to maximize activation in real conversations.
Quick Start
Draft a new skill, run an evaluation loop, review results, and iterate until the description triggers reliably for your target prompts.