What problem does it solve?
Building effective AI skills requires more than just writing instructions — it demands systematic testing, iteration, and optimization to ensure reliable triggering and consistent performance. Without structured evaluation, skills often undertrigger or produce inconsistent results.
Core Features & Use Cases
- Skill Creation: Guide users through capturing intent, writing SKILL.md files, and organizing bundled resources like scripts and references.
- Evaluation & Benchmarking: Run quantitative tests with baseline comparisons, grade outputs against assertions, and analyze performance metrics like pass rate, timing, and token usage.
- Description Optimization: Iteratively improve skill descriptions for better triggering accuracy using train/test splits and automated optimization loops.
Quick Start
Use the skill-creator skill to design a new skill for summarizing meeting transcripts, write test cases for it, and run a benchmark to compare its performance against the baseline.