What problem does it solve?
This Skill eliminates the tedious, error-prone manual work of transcribing text from scanned documents, image-based PDFs, and complex file formats, saving hours of repetitive effort for researchers, analysts, and teams working with digital documents.
Core Features & Use Cases
- Dual Extraction Modes: Uses lightweight pymupdf for instant text extraction from standard PDFs, and high-accuracy marker-pdf for OCR, equation parsing, and complex layout analysis of scanned documents.
- Broad Format & Workflow Support: Works with PDFs, DOCX, PPTX, EPUB, and images, with built-in tools for splitting, merging, and searching across document collections, plus native support for Arxiv paper extraction.
- Real-World Use Case: A researcher can extract full text and structured table data from 100 scanned academic papers or quarterly financial reports in minutes, no manual transcription required.
Quick Start
Use the ocr-and-documents skill to extract all editable text and table data from the attached scanned quarterly report PDF and save it as a clean markdown file.