PDF OCR Extraction

Perform OCR on scanned PDFs to extract text and create searchable PDFs.

368|75|Updated Jan 29, 2026
One-click install
npx skills add https://github.com/claude-office-skills/skills --skill pdf-ocr-extraction
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: PDF OCR Extraction
Source: https://github.com/claude-office-skills/skills/tree/main/pdf-ocr
Command: npx skills add https://github.com/claude-office-skills/skills --skill pdf-ocr-extraction

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill solves the problem of extracting text from scanned or image-based PDF documents, making them searchable and editable.

Core Features & Use Cases

  • Optical Character Recognition (OCR): Converts images of text into machine-readable text.
  • Text Extraction: Extracts plain text or structured data from PDFs.
  • Searchable PDF Creation: Adds a text layer to image-based PDFs, allowing for text searching.
  • Use Case: Digitize a stack of old paper documents by scanning them into PDFs and then using this Skill to make all the text searchable and copyable.

Quick Start

Extract text from this scanned PDF.

Frequently Asked Questions about PDF OCR Extraction

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from a scanned PDF that isn't searchable?

To extract text from a scanned PDF, this Skill performs optical character recognition (OCR) to convert images of text into machine-readable plain text or structured data.

Can I add a text layer to an image-based PDF to make it searchable?

Yes, you can create a searchable PDF by adding a text layer to an image-based PDF, which allows you to search and copy text directly from the document.

Does OCR text extraction work with multiple languages?

Yes, OCR text extraction supports multiple languages, allowing you to accurately digitize and extract text from scanned PDF documents regardless of the language.

What image quality do I need for accurate OCR on scanned documents?

For accurate OCR on scanned documents, you need high image quality; this Skill provides specific guidance on image pre-processing to ensure optimal text extraction results.

What is the best way to digitize a stack of old paper documents?

The best way to digitize old paper documents is to scan them into PDFs and use OCR to make the text searchable, editable, and ready for structured data extraction.