skill-creator

Automate skill creation, iteration, and benchmarking with SKILL.md generation and evaluation loops.

Updated Mar 31, 2026
One-click install
npx skills add https://github.com/alex-robert-fr/better-skill-creator --skill skill-creator-alex-robert-fr
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-creator
Source: https://github.com/alex-robert-fr/better-skill-creator/tree/main/skills/skill-creator
Command: npx skills add https://github.com/alex-robert-fr/better-skill-creator --skill skill-creator-alex-robert-fr

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires anthropic.

What problem does it solve?

Create, iterate on, and benchmark skills with a structured workflow to reduce subjective guesswork and accelerate decision-making.

Core Features & Use Cases

  • Co-ordinate the entire skill lifecycle: creation, iteration, evals, and benchmarking.
  • Compare multiple skill versions (including vanilla) and produce decision briefs.
  • Generate SKILL.md content and orchestration artifacts for testing and packaging.

Quick Start

Write a SKILL.md, assemble an eval_set, run parallel evals, review the benchmark, and decide the best iteration with the dashboard.

Frequently Asked Questions about skill-creator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark and iterate on skills to drive evidence-based decisions?

To benchmark and iterate on skills, this tool coordinates interview prompts, evaluation loops, and benchmark analysis to surface actionable insights, enforcing a deterministic workflow for skill creation, testing, iteration, and packaging.

What is the best way to coordinate the lifecycle of skills from creation to packaging?

The best way to coordinate the skill lifecycle is to write a SKILL.md, assemble an eval_set, run parallel evals, review the benchmark, and decide the best iteration using the dashboard to reduce subjective guesswork.

Can I compare multiple skill versions and generate decision briefs?

Yes, you can compare multiple skill versions, including vanilla baselines, to produce decision briefs that automate the end-to-end process of creating, iterating on, and benchmarking skills.

Do I need the anthropic dependency to run evaluation loops for SKILL.md generation?

Yes, the anthropic dependency is required to run the evaluation loops and orchestration artifacts that generate SKILL.md content and enforce the deterministic workflow for testing and packaging skills.

Why does manual skill testing produce subjective guesswork instead of actionable insights?

Manual skill testing lacks a structured workflow, whereas automating the end-to-end process of creation, iteration, and benchmarking enforces deterministic evaluation loops that surface actionable insights for evidence-based decisions.

What are the limitations of using a deterministic workflow for skill benchmarking?

A deterministic workflow for skill benchmarking enforces a strict process for SKILL.md generation and evaluation loops, which may reduce subjective guesswork but requires assembling a proper eval_set to run parallel evals effectively.