pdf-context-refinery

Convert PDF files into structured Markdown documents with OCR and preserved tables and images.

Updated Apr 21, 2026
One-click install
npx skills add https://github.com/itda-skills/skills.pub --skill pdf-context-refinery
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pdf-context-refinery
Source: https://github.com/itda-skills/skills.pub/tree/main/itda-work/skills/pdf-context-refinery
Command: npx skills add https://github.com/itda-skills/skills.pub --skill pdf-context-refinery

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires poppler-utils, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the challenge of transforming complex PDFs into clean, structured Markdown text suitable for language models, eliminating manual markup and formatting effort.

Core Features & Use Cases

  • PDF Conversion: Automatically interpret and convert scanned or digital PDFs into structured, readable Markdown.
  • Knowledge Base Creation: Segment large PDF documents into sections for easy reference and retrieval.
  • Use Case: Convert a 200-page legal document with tables and images into Markdown, preserving structure and embedding images for review or AI training purposes.

Quick Start

Use the pdf context refinery skill to convert your PDF files into clean Markdown for knowledge base building.

Frequently Asked Questions about pdf-context-refinery

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert a PDF to Markdown while preserving tables and images?

This Skill converts PDF files into structured Markdown documents, preserving tables, images, and overall document structure during the conversion process. It automatically interprets both scanned and digital PDFs to retain formatting for easy editing and AI referencing.

Can I use OCR to extract text from scanned PDF documents for a knowledge base?

Yes, this Skill supports OCR to extract text from scanned PDFs and convert them into structured Markdown. It segments large documents into sections, making them suitable for easy reference and knowledge base retrieval.

Does poppler-utils support converting large legal PDF files into structured text?

Yes, this Skill utilizes poppler-utils to process large legal PDF files, accurately preserving complex formatting and embedding images. It converts extensive documents into structured Markdown suitable for review or AI training purposes.

What is the best way to transform complex PDFs into clean text for language models?

The best way to transform complex PDFs into clean text for language models is using this Skill to eliminate manual markup. It interprets document structures and outputs structured Markdown, making the content immediately suitable for AI referencing and training.

Do I need to manually format tables when converting PDF documents to Markdown?

No, you do not need to manually format tables when converting PDF documents to Markdown. This Skill automatically interprets and preserves tables, images, and document structure, eliminating manual formatting effort during the conversion process.