What problem does it solve?
Building effective AI agent skills requires more than just writing instructions—it demands systematic testing, measurement, and iterative refinement to ensure reliable triggering and consistent performance across diverse user queries.
Core Features & Use Cases
- Skill Creation & Editing: Draft new skills from user intent or improve existing ones with structured workflows.
- Quantitative Evaluation: Run evals with subagent comparisons, grading, and variance analysis to measure skill impact.
- Description Optimization: Automatically improve skill triggering accuracy using train/test splits and Claude-powered iteration loops.
- Packaging & Distribution: Validate and package skills into distributable .skill files for installation across agents.
Quick Start
Use the skill-creator skill to build a new skill for extracting invoice data from PDFs by describing the workflow and test cases you want to validate.