autoresearch

Run autonomous experimentation loops to mutate prompts and improve Claude Code skills.

2|Updated May 1, 2026
One-click install
npx skills add https://github.com/nikkogibler/MarsFounderIO --skill autoresearch-nikkogibler
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: autoresearch
Source: https://github.com/nikkogibler/MarsFounderIO/tree/main/my-instructions/skills/autoresearch
Command: npx skills add https://github.com/nikkogibler/MarsFounderIO --skill autoresearch-nikkogibler

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Autoresearch dramatically reduces the manual effort required to improve Claude Code skills by automating iterative testing, evaluation, and prompt mutation, delivering repeatable improvements over time.

Core Features & Use Cases

  • Autonomous experimentation loops: generate skill outputs, apply binary eval criteria, and mutate prompts to improve performance.
  • Versioned optimization: create and compare successive SKILL.md revisions with logs and changelogs.
  • Transparent governance: maintain a baseline, dashboard, and auditable mutation history for safe, trackable improvement.
  • General applicability: works on any Claude Code skill you want to optimize, not just a single target.

Quick Start

Run baseline autoresearch on a target skill to generate an improved SKILL.md, a results log, and a dashboard.

Frequently Asked Questions about autoresearch

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate prompt mutation for Claude Code skills using binary evals?

Automate prompt mutation by running autonomous experimentation loops that score outputs against binary evals and mutate prompts to improve performance. This process generates an improved SKILL.md, a results log, and a dashboard.

What is autonomous experimentation for AI skill optimization?

Autonomous experimentation for AI skill optimization is an automated process that iteratively generates outputs, evaluates them against binary criteria, and mutates prompts to deliver repeatable improvements over time.

Do I need defined binary evaluators to run autonomous prompt mutation loops?

Yes, you need clearly defined binary evaluators to run autonomous prompt mutation loops. The optimization process requires baseline establishment, binary eval criteria, and structured logging to deliver trackable improvements.

How does versioned optimization track changes to mutated SKILL.md files?

Versioned optimization tracks changes by creating and comparing successive SKILL.md revisions. It produces an auditable mutation history, a changelog, and a results log to ensure safe and trackable improvement.

Can I use autonomous experimentation loops to optimize any Claude Code skill?

Yes, you can use autonomous experimentation loops to optimize any Claude Code skill. The framework has general applicability and works on any target skill you want to improve, not just a single specific target.

What is the best way to track AI prompt optimization results in a dashboard?

The best way to track AI prompt optimization results is to generate a live dashboard from test runs. The dashboard maintains transparent governance by displaying baseline data, structured logs, and auditable mutation history.