What problem does it solve?
This Skill removes the manual effort of reading scanned PDFs and images by converting them into searchable, editable text. It is built for documents that cannot be copied directly, such as photocopies, scans, screenshots, and image-based reports.
Core Features & Use Cases
- Multi-engine OCR: Choose between RapidOCR, RapidDoc, PaddleOCR, or SiliconFlow API depending on speed, structure, and accuracy needs.
- PDF and image support: Process scanned PDFs as well as common image formats such as JPG, PNG, BMP, GIF, TIFF, and WEBP.
- Structured extraction: Preserve reading order, recover layout where possible, and return Markdown when using the enhanced engine.
- Practical workflows: Use it for contracts, books, reports, forms, screenshots, and batch conversion of large document sets.
Quick Start
Ask the Skill to extract text from your scanned PDF or image using the OCR engine you prefer, and it will return the recognized text with page-by-page results.