skill-creator

Develop, test, and optimize custom AI skills with Python-based benchmarking.

12|3|Updated Jun 17, 2026
One-click install
npx skills add https://github.com/phuhao00/bony-agent --skill skill-creator-phuhao00
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: skill-creator
Source: https://github.com/phuhao00/bony-agent/tree/main/.agent/skills/skill-creator
Command: npx skills add https://github.com/phuhao00/bony-agent --skill skill-creator-phuhao00

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This skill streamlines the complex lifecycle of creating, testing, and refining custom AI skills, ensuring they are robust, accurate, and performant.

Core Features & Use Cases

  • Iterative Development: Provides a structured loop for drafting, evaluating, and rewriting skill instructions based on real-world performance.
  • Quantitative Benchmarking: Automates the execution of test cases and generates comparative metrics to validate improvements.
  • Trigger Optimization: Uses automated loops to refine skill descriptions, ensuring the AI triggers the skill exactly when needed.
  • Use Case: If you are building a specialized research agent, use this skill to create the initial prompt, run it against a set of test queries, analyze the results, and automatically optimize the description to prevent under-triggering.

Quick Start

Use the skill-creator to draft a new skill for summarizing technical documentation and set up an initial test suite.

Frequently Asked Questions about skill-creator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build and test custom AI skills iteratively?

To build and test custom AI skills iteratively, you can use a structured loop to draft, evaluate, and rewrite prompt instructions based on real-world performance and quantitative benchmarking.

What is the best way to optimize AI skill descriptions for accurate triggering?

The best way to optimize AI skill descriptions for accurate triggering is to run automated loops that refine the description, ensuring the AI triggers the skill exactly when needed and prevents under-triggering.

How does quantitative benchmarking work for prompt optimization?

Quantitative benchmarking for prompt optimization works by automating the execution of test cases against your AI skills, generating comparative metrics and HTML reports to validate performance improvements.

Do I need a Python environment to run automated test suites for AI development?

Yes, you need a Python-based execution environment to run automated test suites for AI development, as it is required to manage subagent workflows, benchmark aggregation, and HTML report generation.

Can I evaluate research agent performance using assertion grading?

Yes, you can evaluate research agent performance using quantitative assertion grading, which automates test case execution and generates comparative metrics to validate improvements.

When should I not use automated skill engineering for my AI workflows?

You should not use automated skill engineering if you lack a Python-based execution environment, as it is strictly required to manage subagent workflows and benchmark aggregation.