pdf-reader

Extract text, tables, and forms from PDF documents with OCR.

4|1|Updated Mar 3, 2026
One-click install
npx skills add https://github.com/aegntic/clawreform --skill pdf-reader-aegntic
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pdf-reader
Source: https://github.com/aegntic/clawreform/tree/main/crates/clawreform-skills/bundled/pdf-reader
Command: npx skills add https://github.com/aegntic/clawreform --skill pdf-reader-aegntic

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the extraction, interpretation, and summarization of content from PDF documents, including text, tables, and forms, saving users significant time and effort in document analysis.

Core Features & Use Cases

  • Content Extraction: Extracts text, tables, and form data from various PDF types (text-based, scanned).
  • Document Analysis: Summarizes content, compares documents, and searches for specific information.
  • Use Case: A researcher needs to quickly understand the key findings from a lengthy academic paper. This Skill can extract the abstract, methodology, results, and conclusions, providing a concise summary.

Quick Start

Use the pdf-reader skill to summarize the content of the attached document 'research_paper.pdf'.

Frequently Asked Questions about pdf-reader

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text and tables from a scanned PDF document?

To extract text and tables from a scanned PDF, the pdf-reader utilizes OCR technology to recognize characters while preserving the original document structure. This allows accurate data extraction from image-based files.

Can I summarize and analyze academic papers or financial documents?

Yes, you can summarize and analyze academic papers, financial documents, and legal texts. The tool extracts abstracts, methodologies, and results to provide concise summaries and compare different documents.

How does PDF content extraction handle document structure and forms?

PDF content extraction preserves document structure during the process, accurately capturing text, tables, and form data. This ensures the extracted information maintains its original context and formatting.

What is the best way to search for specific information across multiple PDF files?

The best way to search for specific information across PDF files is using the built-in search functionality. It processes extracted text and tables to locate targeted data across various document types.

Does PDF extraction work without additional dependencies or libraries?

Yes, PDF extraction works without additional dependencies, as the Skill operates independently with its built-in scripts. It handles text-based and scanned documents natively without requiring external libraries.