autorefine

Automates eval-guided improvement of SKILL.md through audit and mutation.

9|Updated Mar 22, 2026
One-click install
npx skills add https://github.com/surahli123/autorefine-skill-improvement --skill autorefine
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: autorefine
Source: https://github.com/surahli123/autorefine-skill-improvement/tree/main/autorefine
Command: npx skills add https://github.com/surahli123/autorefine-skill-improvement --skill autorefine

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

AutoRefine provides a guided, eval-grounded workflow to iteratively improve a SKILL.md by surfacing misalignments, validating improvements with judges and domain tests, and mutating toward verifiable gains.

Core Features & Use Cases

  • Structured preflight audits, Gulf 1–3 evaluation pipeline, and a disciplined mutation loop to refine SKILL.md with eval-grounded evidence.
  • Pattern-aware routing and adapter-aware domain evaluation, ambient learning, and domain-metric integration to ensure changes are verifiable.
  • Works across a broad spectrum of skills, from automation to agent workflows, with full audit trails and a clear session-close apply-back.

Quick Start

Run AutoRefine on a target skill path to start the guided improvement workflow.

Frequently Asked Questions about autorefine

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I iteratively improve a SKILL.md file with eval-grounded evidence?

AutoRefine improves a SKILL.md by automating eval-grounded autoresearch, combining a design-audit with Hamel's Three Gulfs and a mutation loop to validate changes through judge validation and domain tests.

What is the mutation loop process for skill improvement?

The mutation loop process for skill improvement applies phase-based mutations to a SKILL.md, routing downstream work through pattern-aware strategies and contract integrity checks to ensure verifiable, eval-grounded gains before session-close cleanup.

How does a design-audit using Hamel's Three Gulfs evaluate skill workflows?

A design-audit using Hamel's Three Gulfs evaluates skill workflows by surfacing misalignments between intent and execution, providing structured preflight checks before entering the mutation loop to refine the SKILL.md.

Can I apply automated evaluation and mutation to any SKILL.md?

Yes, you can apply automated evaluation and mutation to any skill defined by a SKILL.md, spanning automation to agent workflows, using adapter-domain evaluation and ambient learning to verify improvements across diverse domains.

What is the best way to validate SKILL.md improvements with domain tests?

The best way to validate SKILL.md improvements with domain tests is running a disciplined mutation loop that uses judge validation and domain-metric integration, ensuring changes are verifiable and backed by eval-grounded evidence.

Why does my skill workflow need pattern-aware routing and contract integrity checks?

Your skill workflow needs pattern-aware routing and contract integrity checks to maintain verifiable improvements during mutation, ensuring adapter-domain evaluation aligns with ambient learning and preventing structural breakdowns across the SKILL.md.