autoresearch

Benchmark and improve existing repo skills through repeated evals and score tracking.

55|8|Updated Mar 29, 2026
One-click install
npx skills add https://github.com/j1ngg/tech-marketing-framework --skill autoresearch-j1ngg
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: autoresearch
Source: https://github.com/j1ngg/tech-marketing-framework/tree/main/.agents/skills/autoresearch
Command: npx skills add https://github.com/j1ngg/tech-marketing-framework --skill autoresearch-j1ngg

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill provides an autonomous framework to benchmark and improve an existing repo skill through repeated evaluations, prompt mutations, and score tracking.

Core Features & Use Cases

  • Autonomous evaluation loops: Run repeated tests across multiple prompts and configurations to identify superior variants.
  • Score tracking and comparison: Collect metrics over iterations to guide optimization and future refinements.
  • Use Case: Teams can iteratively enhance a Codex/CUDA-like skill by continuously evaluating variants and selecting the best-performing prompts.

Quick Start

Start an autoresearch cycle against an existing skill, then review the score history and mutate prompts for improvement.

Frequently Asked Questions about autoresearch

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up autonomous evaluation loops for prompt mutations?

Yes, you can benchmark and improve an existing repo skill by running repeated evaluations against it. The autonomous evaluation loop tracks scores iteratively to guide optimization and select the best-performing prompt mutations for your skill.

What is autonomous evaluation for prompt mutations and score tracking?

Autonomous evaluation is a framework that benchmarks and improves existing repo skills through repeated evals, prompt mutations, and score tracking. It runs test cycles across multiple prompts and configurations to identify superior variants and guide optimization.

Can I use autonomous evaluation loops with Codex-like environments?

Yes, the autonomous evaluation framework is designed for Codex and Codex-like environments. It requires the canonical autoresearch workflow and compatible evaluation tooling to track and compare performance across iterations within these platforms.

Do I need compatible evaluation tooling to track and compare skill performance?

Yes, compatible evaluation tooling is required to track and compare performance across iterations. The framework relies on these tools to collect metrics over iterations, enabling score tracking and comparison to guide prompt optimization.

What's the best way to iteratively enhance a skill by continuously evaluating variants?

The best way is to start an autoresearch cycle against an existing skill, then review the score history and mutate prompts for improvement. This allows teams to iteratively enhance skills by continuously evaluating variants and selecting the best-performing prompts.