pdf-text-extractor

Extract text and metadata from PDFs with OCR and multiple output formats.

Updated Apr 7, 2026
One-click install
npx skills add https://github.com/zhangyanbo2007/openclaw --skill pdf-text-extractor-zhangyanbo2007
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pdf-text-extractor
Source: https://github.com/zhangyanbo2007/openclaw/tree/main/workspace-fox-avatar/skills/pdf-text-extractor
Command: npx skills add https://github.com/zhangyanbo2007/openclaw --skill pdf-text-extractor-zhangyanbo2007

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

PDF content is often locked in PDFs, making it hard to search, reuse, or archive. PDF-Text-Extractor converts PDF content into editable text and structured output, with optional OCR for scanned documents, enabling fast digitization and analysis.

Core Features & Use Cases

  • Text extraction from PDFs (text-based) and OCR for scanned documents
  • Batch processing of multiple PDFs with metadata extraction
  • Output formats include plain text, JSON, Markdown, and HTML, plus page-aware structure and language detection

Quick Start

Quick start by running a single extraction on a PDF to obtain text and metadata.

Frequently Asked Questions about pdf-text-extractor

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from scanned PDF documents?

To extract text from scanned PDFs, this tool uses OCR to recognize and digitize text from images. It supports language detection and outputs page-aware results in formats like plain text, JSON, Markdown, or HTML for archiving.

Can I batch process multiple PDFs for text extraction at once?

Yes, you can batch process large sets of PDFs for text extraction simultaneously. The tool handles batch-processing while extracting metadata and digitizing content from multiple invoices or contracts into structured outputs.

Does this PDF text extractor require any external dependencies?

No, this PDF text extractor operates with zero dependencies. It runs text extraction and OCR directly on your documents without requiring additional library installations or external environment setups.

What output formats are available when extracting text from PDFs?

When extracting text from PDFs, available output formats include plain text, JSON, Markdown, and HTML. These formats provide page-aware structure and metadata, enabling easy digital archiving and content processing.

How does OCR language detection work during PDF digitization?

During PDF digitization, OCR language detection automatically identifies the language of the scanned text. This ensures accurate text extraction and proper rendering in the chosen output format for document archiving.

What is the best way to digitize invoices and contracts from PDF files?

The best way to digitize invoices and contracts is using PDF text extraction with OCR. It converts locked PDF content into editable text and structured JSON output, supporting batch-processing for efficient digital archiving.