What problem does it solve?
Streamlines the creation, evaluation, and iterative improvement of skills and system prompts for multiple AI platforms so teams can produce reliable, testable skills without reinventing infrastructure for every iteration.
Core Features & Use Cases
- Multi-platform authoring: Draft SKILL.md frontmatter and instruction bodies for Claude Code, Claude.ai, OpenAI/ChatGPT, Cursor, Windsurf, and generic .md skill formats.
- Evaluation & benchmarking: Generate eval sets, run parallel with-skill and baseline runs, grade outputs, aggregate statistics, and produce human-facing benchmark reports.
- Improve & package: Automatically propose description improvements, run optimization loops, and package validated skills into distributable .skill archives using included scripts.
- Use Case: Rapidly create a new Claude Code skill, run trigger evaluations versus baselines, iterate on descriptions and assertions, and produce a benchmarked, packaged skill ready for deployment.
Quick Start
Ask the assistant to "Help me create a Claude Code SKILL.md for [task], draft 3 test prompts, and run a baseline vs with-skill evaluation."