ref-hallucination-arena

Verify cited papers against Crossref, PubMed, arXiv, and DBLP to quantify hallucination rates.

775|63|Updated Jul 8, 2025
One-click install
npx skills add https://github.com/agentscope-ai/OpenJudge --skill ref-hallucination-arena
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ref-hallucination-arena
Source: https://github.com/agentscope-ai/OpenJudge/tree/main/skills/ref-hallucination-arena
Command: npx skills add https://github.com/agentscope-ai/OpenJudge --skill ref-hallucination-arena

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires matplotlib, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the critical issue of Large Language Models (LLMs) fabricating academic references, ensuring the reliability of AI-generated literature reviews and citations.

Core Features & Use Cases

  • Comprehensive Verification: Validates cited papers against major academic databases (Crossref, PubMed, arXiv, DBLP).
  • Detailed Metrics: Measures hallucination rate, per-field accuracy (title, author, year, DOI), and discipline-specific performance.
  • Use Case: When evaluating an LLM's ability to generate a literature review, use this Skill to automatically check if all the provided citations are real and accurate, quantifying its tendency to hallucinate.

Quick Start

Evaluate LLM reference recommendation capabilities using the provided configuration file.

Frequently Asked Questions about ref-hallucination-arena

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I check if an LLM is hallucinating academic references?

You detect LLM reference hallucinations by verifying generated citations against major academic databases like Crossref, PubMed, arXiv, and DBLP. This quantifies fabrication rates and measures per-field citation accuracy for AI-generated literature reviews.

What databases can I use to verify LLM citation accuracy?

You can verify LLM citation accuracy against Crossref, PubMed, arXiv, and DBLP. Validating cited papers across these databases measures discipline-specific compliance and detects fabricated academic references in generated text.

How do I benchmark LLM literature review generation for hallucinated papers?

Benchmark LLM literature reviews by validating recommended papers against Crossref, PubMed, arXiv, and DBLP. This quantifies hallucination rates and measures per-field citation accuracy, including title, author, year, and DOI match rates.

Does LLM hallucination detection support tool-augmented modes?

Yes, LLM hallucination detection supports tool-augmented modes for enhanced accuracy. This approach evaluates LLM reference recommendation capabilities by verifying cited papers against Crossref, PubMed, arXiv, and DBLP to ensure citation reliability.

What metrics are used to evaluate LLM reference recommendation capabilities?

Evaluating LLM reference recommendation capabilities uses metrics including hallucination rate, per-field accuracy for title, author, year, and DOI, plus discipline-specific compliance measured against verified academic databases like Crossref and PubMed.

Do I need matplotlib to benchmark LLM reference accuracy?

Yes, matplotlib is required to benchmark LLM reference accuracy. This dependency supports the visualization of hallucination rates and per-field citation accuracy metrics when evaluating AI-generated academic references.