eval-design-forensics

Audit research paper evaluation design and reporting validity for leakage and bias.

123|7|Updated Jun 26, 2026
One-click install
npx skills add https://github.com/wanshuiyin/Anti-Autoresearch --skill eval-design-forensics
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eval-design-forensics
Source: https://github.com/wanshuiyin/Anti-Autoresearch/tree/main/skills/eval-design-forensics
Command: npx skills add https://github.com/wanshuiyin/Anti-Autoresearch --skill eval-design-forensics

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill audits the evaluation design and reporting validity of research papers, ensuring that the reported results are reliable and accurately reflect the research methods.

Core Features & Use Cases

  • Evaluation Design Audit: Checks if the evaluation design in a paper measures what it claims and if the reporting is complete.
  • Leakage Detection: Identifies leakage between training and testing data.
  • Judge Validity Check: Verifies if the evaluation judge is unbiased and validated.
  • Selective Reporting Analysis: Inspects for selective reporting in the paper's evaluation.
  • Use Case: If you are reviewing a paper and want to ensure that the evaluation metrics are reliable, use this Skill to analyze the paper's evaluation design and reporting.

Quick Start

Use the eval-design-forensics skill to audit the evaluation design and reporting of the paper in the 'paper-dir'.

Frequently Asked Questions about eval-design-forensics

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I detect data leakage between training and testing sets in a research paper?

To detect data leakage in a research paper, you can audit its evaluation design to identify potential overlap or leakage between training and testing data, ensuring reported results are reliable.

What is evaluation validity and how do I check it during a paper review?

Evaluation validity ensures a paper's reported metrics accurately reflect its research methods. You can check it by auditing the evaluation design, verifying judge validity, and inspecting for selective reporting.

How can I verify if an evaluation judge is unbiased in machine learning research?

You can verify if an evaluation judge is unbiased by auditing the paper's evaluation design to check whether the judge is validated and free from systematic bias affecting the reported results.

How do I inspect a research paper for selective reporting of evaluation results?

To inspect a research paper for selective reporting, audit its evaluation design and reporting validity to identify if evaluation metrics are selectively chosen or incomplete, affecting result reliability.

Can I audit the evaluation design of a paper without writing custom scripts?

Yes, you can audit the evaluation design and reporting validity of papers without writing custom scripts by using an automated skill to analyze the paper's claims, methodology, and results directly.

What are the limitations of automated evaluation audits for research papers?

Automated evaluation audits rely on analyzing the paper's claims, methodology, and results, meaning they cannot detect issues omitted from the text or evaluate experimental execution beyond the reported design.