What problem does it solve?
It helps you turn vague performance improvement goals into measurable, repeatable experimentation cycles with clear decisions to keep or discard changes.
Core Features & Use Cases
- Research-cycle execution: Runs a hypothesis → change → measure → evaluate loop aligned to a configured metric, surface, direction, and measurement window.
- Evidence-based evaluation: Measures outcomes using the cycle’s configured method (quantitative scripted/computed or qualitative scoring with required justification).
- Experiment governance: Supports approval-gated experiment creation and logs learnings for every run, including failures, to avoid repeating discarded hypotheses.
- Use case: Optimize a specific operational metric for an agent fleet (for example, improving “system_effectiveness”) by testing targeted modifications to a surface file and evaluating results against a baseline.
Quick Start
Assign and run an autoresearch cycle by telling the analyst to start a research cycle for the metric you want to improve, then wait for the experiments to be created and evaluated.