What problem does it solve?
Skill Forge automates the iterative improvement of AI skills and generic codebases by running a closed-loop experiment that mutates instructions, evaluates changes, and preserves only beneficial improvements, enabling overnight optimization without manual intervention.
Core Features & Use Cases
- Autonomous Skill Improvement: Optimizes SKILL.md in Skill Mode by running evals, scoring with a composite metric, and keeping only high-impact mutations.
- Generic-Metric Optimization: Applies the same autoresearch paradigm to any file against a numeric shell metric (e.g., test coverage, bundle size, performance).
- Robust Workflow: Includes a setup wizard, dry-run validation, TSV experiment logs, coverage matrix, crash recovery, and an optional guided mode for hands-on control.
- Overnight and Scheduled Runs: Supports scheduled tasks that run unattended and deliver a morning report.
- Real-world scenarios: improve a LinkedIn-post skill, reduce Docker image size, or improve code quality.
Quick Start
Tell Skill Forge to auto-tune the target SKILL.md to maximize the designated evaluation metric.