OmniDocBench
Evaluate document parsing and OCR models against real-world PDFs
All Skills in This Repository (1)
Pure Emerald Level IndicatorsFrequently Asked Questions
FAQPage SchemaHow to install OmniDocBench?โผ
Run `npx skills add opendatalab/OmniDocBench --all -g -y` in your terminal to install the evaluation helper skill globally.
What is OmniDocBench used for?โผ
It is a benchmark for testing how well document parsing and OCR models handle real-world PDFs like papers, newspapers, and handwritten notes. It scores text, formulas, tables, layout, and reading order.
How do I evaluate my OCR model on OmniDocBench?โผ
Point the skill at your ground-truth JSON and prediction markdown folder, and it generates a config and runs the evaluation for you. It then reports Overall, Text, Formula, and Table scores with result file paths.
Does OmniDocBench work with Docker?โผ
Yes. Docker is the recommended way to run evaluations because it bundles the TeX, ImageMagick, and Ghostscript dependencies needed for CDM formula scoring.
Can I use OmniDocBench without coding experience?โผ
Yes. The AI agent validates your file paths, builds the config, runs the evaluation, and explains the scores based on your plain-English requests.
Related Repositories in Data & Analytics
View All in Data & AnalyticsโPaddleOCR
Extract text, tables, and formulas from PDFs and images
Scrapling
Scrape any website and bypass anti-bot protection with AI
last30days-skill
Research any topic across Reddit, X, YouTube, and the web