What problem does it solve?
It helps reduce the recurring 30% of weak or inconsistent outputs by running an evidence-driven improvement loop over skills, memory, agents, and documentation, while also auditing and repairing “silent” memory drift.
Core Features & Use Cases
- Evidence-based self-improvement: Performs eval-loop style scoring and targeted mutation based on recent failures, file changes, corrections, and repeated friction.
- Memory audit and repair: Detects staleness, gaps, redundancy, and inconsistencies, then rewrites or consolidates memory entries and produces an audit report and changelog.
- Background conversation review (Hermes pattern): Every 10 turns (and when durable preferences appear), reviews the conversation and saves only genuinely durable insights without interrupting the main task.
- Eval-loop production workflow: Creates dashboards, results tracking (results.json/results.tsv), and a structured, binary eval suite to measure improvement over multiple runs.
Quick Start
Ask your agent to run auto-improve to audit recent evidence, score results with an eval loop, and update the most leverageful skill, documentation, or memory files.