What problem does it solve?
It prevents researchers from over-claiming by using a structured evidence-gathering and expert judgment step to decide whether results actually support the intended research claim.
Core Features & Use Cases
- Evidence-to-judgment routing: Collects metrics, baselines, caveats, and intended claim context from sources like W&B runs and experiment logs, then sends them to Codex for an explicit support verdict.
- Verdict-driven next actions: Automatically routes outcomes into next steps (pivot, supplement, or confirm) and updates findings/wiki artifacts when available.
- Traceable decision workflow: Encourages forensic trace capture after Codex calls and records integrity/audit status when present, so judgments are reproducible.
Quick Start
After your experiments finish, run Codex via the result-to-claim gate by supplying your intended claim and key metrics (from W&B history or log tables) so you can get a supported/partial/not-supported verdict and the next experiments to run.