document-parser

Extract text and structured fields from PDFs and images using OCR.

20|7|Updated Feb 2, 2026
One-click install
npx skills add https://github.com/Alexi5000/ClawKeeper --skill document-parser
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: document-parser
Source: https://github.com/Alexi5000/ClawKeeper/tree/main/skills/document-parser
Command: npx skills add https://github.com/Alexi5000/ClawKeeper --skill document-parser

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

OCR and parse documents including invoices, receipts, and bank statements, enabling automated processing of scanned and digital documents for faster workflows.

Core Features & Use Cases

  • OCR Text Extraction: Convert images to text and make content searchable.
  • PDF Parsing & Structured Data Extraction: Extract text, fields, and line items from PDFs and images.
  • Document Classification & Quality Assessment: Identify document types and assess OCR confidence to drive routing and review decisions.

Quick Start

Provide a PDF or image to the skill and request structured extraction of text and key fields.

Frequently Asked Questions about document-parser

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract structured data from PDF invoices and receipts?

You can extract structured data from PDFs by providing the document to the skill and requesting structured extraction of text and key fields. It leverages OCR technologies to parse invoices and receipts for automated processing.

Can I use Tesseract or Google Document AI for image processing?

Yes, you can use either Tesseract or Google Document AI for image processing and OCR text extraction. These options allow you to convert images to searchable text and extract structured data from scanned documents.

What is the best way to automate bank statement parsing?

The best way to automate bank statement parsing is using OCR technologies to extract fields and line items. This digitizes scanned and digital documents quickly, enabling automated validation within business workflows.

Does this approach support document classification and quality assessment?

Yes, this approach supports document classification and quality assessment by identifying document types and evaluating OCR confidence. This drives routing and review decisions for scanned images and PDFs.

How do I classify document types and assess OCR confidence for routing?

You can classify document types and assess OCR confidence by applying OCR technologies to your documents. This process identifies the document type and evaluates extraction quality to drive routing and review decisions.