skill-creator

Orchestrate end-to-end skill creation, evaluation, and description optimization.

Updated Apr 4, 2026
One-click install
npx skills add https://github.com/PheonixCodder/Beeclean-Production --skill skill-creator-pheonixcodder
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-creator
Source: https://github.com/PheonixCodder/Beeclean-Production/tree/main/.agents/skills/skill-creator
Command: npx skills add https://github.com/PheonixCodder/Beeclean-Production --skill skill-creator-pheonixcodder

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires anthropic, pyyaml, and includes scripts (resource) components.

What problem does it solve?

Automates the end-to-end process of creating, evaluating, and refining AI skills, reducing manual guesswork and accelerating improvement cycles.

Core Features & Use Cases

  • Facilitates rapid skill drafting, evaluation, and iteration using eval prompts, transcripts, and dashboards.
  • Automates comparison via blind evaluation and post-hoc analysis to determine better skill variants.
  • Supports packaging, testing, and updating skill descriptions to improve triggering accuracy.

Quick Start

Run an initial evaluation loop to start refining the skill: draft the SKILL.md, run evaluators, review results, and iterate until the trigger metrics stabilize.

Frequently Asked Questions about skill-creator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate AI skill creation and evaluation workflows?

Automating AI skill creation involves orchestrating end-to-end workflows that draft, validate, and iterate on skills using eval prompts and transcripts. This process reduces manual guesswork by coordinating benchmarking and description optimization until trigger metrics stabilize.

What is the best way to iterate on AI skills using benchmarking?

The best way to iterate on AI skills is using a looped workflow that compares variants through blind evaluation and post-hoc analysis. This benchmarking method determines better skill variants by reviewing evaluation transcripts and dashboards to refine performance.

How do I evaluate and optimize skill descriptions for better triggering?

Evaluating and optimizing skill descriptions requires running automated testing and updating packaging to improve triggering accuracy. By running an initial evaluation loop and reviewing benchmark results, you can stabilize trigger metrics for accurate deployment.

Can I use Anthropic and PyYAML dependencies for skill testing automation?

Yes, you can use Anthropic and PyYAML dependencies to support skill testing automation. These components enable the scripts to orchestrate end-to-end skill creation, validate drafts, and iterate using evaluation transcripts within the workflow.

Does blind evaluation work for comparing AI skill variants?

Blind evaluation works effectively for comparing AI skill variants by automating post-hoc analysis. This mechanism evaluates the variants without bias, determining which skill performs better before packaging and testing the final description optimization.

Why do my skill trigger metrics fluctuate during iteration?

Skill trigger metrics fluctuate during iteration when the evaluation loop has not sufficiently refined the skill drafts and descriptions. Consistent progress requires guiding the workflow through validation, benchmarking, and optimization until the metrics stabilize.