mineru

Convert PDFs, images, and web pages into Markdown, HTML, LaTeX, or DOCX.

1|Updated Mar 30, 2026
One-click install
npx skills add https://github.com/Arry8/openclaw-edge --skill mineru-arry8
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mineru
Source: https://github.com/Arry8/openclaw-edge/tree/main/skills/mineru
Command: npx skills add https://github.com/Arry8/openclaw-edge --skill mineru-arry8

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Manual document extraction and format conversion from PDFs, images, and web content is tedious and error-prone; MinerU CLI streamlines extraction and multi-format exports.

Core Features & Use Cases

  • Quick, token-free extraction for small tasks with flash-extract
  • Precision extraction with OCR, table and formula recognition, and multi-format outputs (md, html, latex, docx)
  • Batch processing and web crawling to convert pages to Markdown and other formats
  • Pipelines and piped workflows to integrate into data processing or research projects

Quick Start

Run mineru-open-api flash-extract on a local file to quickly convert it to Markdown.

Frequently Asked Questions about mineru

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert PDF to Markdown without losing table structures?

To convert PDF to Markdown with intact tables, precision extraction applies OCR and table recognition to preserve structures. It outputs clean Markdown alongside HTML, LaTeX, or DOCX formats for research and data analytics workflows.

What is the best way to batch extract text from multiple PDFs and images?

Batch extract text from PDFs and images using automated pipelines that apply OCR across multiple files simultaneously. This streamlines document extraction into Markdown or other formats for content management and analytics workflows.

Can I crawl web pages and convert the content directly to Markdown?

Crawl web pages and convert content directly to Markdown using integrated web crawling features. It extracts page content and exports to Markdown, HTML, LaTeX, or DOCX formats for data processing and research pipelines.

Does MinerU support token-free PDF extraction for quick conversions?

MinerU supports token-free PDF extraction via flash-extract. It quickly converts local files to Markdown without consuming API tokens, making it ideal for small tasks and rapid format conversions.

What formats can I export extracted PDF content into besides Markdown?

Besides Markdown, export extracted PDF content into HTML, LaTeX, and DOCX formats. These multi-format exports are supported during precision extraction, accommodating diverse content management and research documentation needs.

When do I need OCR for document extraction and format conversion?

You need OCR for document extraction when processing scanned PDFs or images with non-selectable text. It recognizes text, tables, and formulas, enabling accurate conversion into Markdown, HTML, LaTeX, or DOCX formats.