experiment-audit

Audit ML experiment integrity using configurable cross-model reviewer backends.

Updated Jun 10, 2026
One-click install
npx skills add https://github.com/xqinag/ARIS-new --skill experiment-audit-xqinag
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: experiment-audit
Source: https://github.com/xqinag/ARIS-new/tree/main/skills/experiment-audit
Command: npx skills add https://github.com/xqinag/ARIS-new --skill experiment-audit-xqinag

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Audit experiment integrity using cross-model reviewer backends to detect fake ground truth, score normalization fraud, phantom results, and insufficient scope across provided experiment artifacts. It applies to post-experiment evaluation across ML experiments, enabling independent verification of datasets, metrics, and narrative claims. It supports configurable reviewer backends (codex/manual) and artifact collection with traceable reporting.

Core Features & Use Cases

  • Cross-model integrity verification via external reviewer backend; executor does not participate in integrity judgment.
  • Guards against fake ground truth, score normalization fraud, phantom results, insufficient scope.
  • Supports configurable reviewer backends and traceable auditing workflow across experiments.

Quick Start

Provide the path to the experiment directory and artifacts to audit, then run the audit workflow.

Frequently Asked Questions about experiment-audit

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I verify ML experiment integrity and detect fake ground truth?

Verify ML experiment integrity by applying cross-model reviewer backends to independently audit datasets and metrics, detecting fake ground truth, score normalization fraud, and phantom results across provided experiment artifacts.

What is cross-model auditing for post-experiment evaluation?

Cross-model auditing is an independent review process where external reviewer backends evaluate experiment artifacts to verify narrative claims and detect fraud, ensuring the executor does not participate in integrity judgments.

How do I detect score normalization fraud in my ML experiment artifacts?

Detect score normalization fraud by providing your experiment directory path to an audit workflow that applies external reviewer backends to trace and verify metrics and datasets across experiment artifacts.

Can I use manual review backends for experiment integrity verification?

Yes, experiment integrity verification supports configurable reviewer backends including both codex and manual modes to perform independent cross-model audits and generate traceable reports.

What types of fraud does cross-model experiment auditing guard against?

Cross-model experiment auditing guards against fake ground truth, score normalization fraud, phantom results, and insufficient scope by applying external reviewer backends to evaluate provided experiment artifacts.

Does experiment auditing work without external dependencies?

Yes, experiment auditing operates without external dependencies, using configurable internal reviewer backends to collect artifacts and generate traceable reports for post-experiment evaluation.