ocr-and-documents

Extract text from PDFs, scanned documents, and images using OCR.

Updated Jun 22, 2026
One-click install
npx skills add https://github.com/ashiqcodeleaf/long-Run-Agents --skill ocr-and-documents-ashiqcodeleaf
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ocr-and-documents
Source: https://github.com/ashiqcodeleaf/long-Run-Agents/tree/main/skills/productivity/ocr-and-documents
Command: npx skills add https://github.com/ashiqcodeleaf/long-Run-Agents --skill ocr-and-documents-ashiqcodeleaf

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pymupdf, marker-pdf, python-docx, web-extract, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the process of extracting text from PDFs, scanned documents, images, and other file formats, saving time and reducing manual effort.

Core Features & Use Cases

  • Text Extraction: Automatically extract text from a wide range of file formats using OCR.
  • PDF Processing: Handle various PDF features such as tables, equations, code blocks, and more.
  • Use Case: Need to extract text from a PDF report but don't want to manually type it all out? Use this skill to quickly convert it into a more accessible format.

Quick Start

Use the ocr-and-documents skill to extract text from the attached PDF 'report.pdf'.

Frequently Asked Questions about ocr-and-documents

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from scanned PDFs and images?

To extract text from scanned PDFs and images, this Skill applies OCR to automatically recognize and digitize document content. It processes various file formats using libraries like pymupdf and marker-pdf, reducing manual typing by handling scanned content directly.

Does this PDF processing approach support extracting tables and equations?

PDF processing with this approach supports extracting complex document features including tables, equations, and code blocks. It leverages marker-pdf and pymupdf to accurately parse these structured elements alongside standard text during document digitization workflows.

What Python libraries do I need to process documents with OCR?

To process documents with OCR, you need the Python libraries pymupdf, marker-pdf, python-docx, and web-extract. These dependencies handle different file types and features, enabling comprehensive text extraction from PDFs, images, and other document formats.

Can I automate admin workflows by digitizing PDF reports?

You can automate admin workflows by digitizing PDF reports to quickly convert them into accessible text formats. This automation saves time and reduces manual effort by extracting text from reports, scanned documents, and images directly.

What is the best way to digitize documents for administrative tasks?

The best way to digitize documents for administrative tasks is using an automated OCR workflow that extracts text from PDFs and images. This method handles various document features like tables and equations, converting physical or scanned files into accessible digital text.