What problem does it solve?
Audit experiment integrity using cross-model reviewer backends to detect fake ground truth, score normalization fraud, phantom results, and insufficient scope across provided experiment artifacts.
It applies to post-experiment evaluation across ML experiments, enabling independent verification of datasets, metrics, and narrative claims.
It supports configurable reviewer backends (codex/manual) and artifact collection with traceable reporting.
Core Features & Use Cases
- Cross-model integrity verification via external reviewer backend; executor does not participate in integrity judgment.
- Guards against fake ground truth, score normalization fraud, phantom results, insufficient scope.
- Supports configurable reviewer backends and traceable auditing workflow across experiments.
Quick Start
Provide the path to the experiment directory and artifacts to audit, then run the audit workflow.