pymupdf-pdf

Extract text and images from PDF files using PyMuPDF.

Updated Mar 20, 2026
One-click install
npx skills add https://github.com/s-k-y-h-i-g-h/OpenState --skill pymupdf-pdf-s-k-y-h-i-g-h
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pymupdf-pdf
Source: https://github.com/s-k-y-h-i-g-h/OpenState/tree/main/skills/pymupdf-pdf-parser-clawdbot-skill
Command: npx skills add https://github.com/s-k-y-h-i-g-h/OpenState --skill pymupdf-pdf-s-k-y-h-i-g-h

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pymupdf, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides rapid, lightweight extraction of text and images from PDF files to streamline data retrieval and document processing tasks.

Core Features & Use Cases

  • Fast text extraction: Converts PDF pages into Markdown or JSON formats quickly for immediate use.
  • Image extraction: Retrieves embedded images from PDF pages for visualization or analysis.
  • Use Case: Quickly extracting text and images from a batch of research papers or invoices for review or archiving.

Quick Start

Run the tool to parse a PDF file and generate Markdown, JSON, and images quickly without complex setup.

Frequently Asked Questions about pymupdf-pdf

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text and images from PDF files quickly?

You can extract text and images from PDF files quickly by running this Skill, which leverages PyMuPDF for fast local parsing and outputs structured Markdown, JSON, and image assets instantly.

Can I convert PDF pages to JSON format for data analysis?

Yes, you can convert PDF pages to JSON format for data analysis, as this Skill rapidly extracts text content and retrieves embedded images, outputting them into structured JSON files without complex setup.

Does this PDF text extraction tool require PyMuPDF?

Yes, this PDF text extraction tool requires PyMuPDF as its core dependency to enable fast local document parsing and ensure rapid retrieval of text and images with minimal setup overhead.

What is the best way to parse a batch of research papers or invoices for archiving?

The best way to parse a batch of research papers or invoices for archiving is using this lightweight Skill, which rapidly extracts text and images locally, prioritizing processing speed over detailed layout accuracy.

When should I not use this approach for PDF parsing?

You should not use this approach for PDF parsing when your task requires detailed layout accuracy, as this Skill is specifically designed for scenarios where fast text and image extraction speed outweighs precise structural fidelity.