ref-pdf-converter

Extract text from PDFs and convert to Markdown with preserved structure and tables.

Updated Mar 19, 2026
One-click install
npx skills add https://github.com/sunLeee/optimization --skill ref-pdf-converter
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ref-pdf-converter
Source: https://github.com/sunLeee/optimization/tree/main/.claude/skills/data-fetch/ref-pdf-converter
Command: npx skills add https://github.com/sunLeee/optimization --skill ref-pdf-converter

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires markitdown, pypdf.

What problem does it solve?

PDF documents are often locked in binary formats that hinder content reuse. This skill converts PDFs into editable Markdown while preserving text layout, headings, and tables.

Core Features & Use Cases

  • Extract text from PDFs and convert to Markdown with preserved structure and tables.
  • Support multi-page documents and maintain metadata references for linking.
  • Use case: Convert technical reports or academic papers into a readable Markdown format for collaboration and archiving.

Quick Start

Convert a sample PDF to Markdown to verify text, headings, and table extraction.

Frequently Asked Questions about ref-pdf-converter

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert PDF to Markdown while preserving tables and headings?

To convert PDF to Markdown while preserving tables and headings, use this skill to extract text and maintain structural formatting. It processes academic papers and technical reports, outputting editable Markdown with accurate layout, headings, and table structures for collaboration.

What is the best way to extract text from multi-page PDF documents for archiving?

The best way to extract text from multi-page PDF documents for archiving is using a conversion tool that maintains metadata references and page structure. This skill handles multi-page PDFs, converting them into readable Markdown while safely preserving document metadata for linking.

Can markitdown handle encrypted PDF files during text extraction?

Yes, markitdown can handle encrypted PDF files during text extraction. The skill uses markitdown as the core tool with optional pypdf integration, safely processing encrypted documents and page metadata to convert them into editable Markdown text.

Does pypdf support table extraction when converting PDFs to Markdown?

Yes, pypdf supports table extraction when converting PDFs to Markdown. When integrated with markitdown, the skill extracts text from PDF documents and converts it into Markdown while preserving structural elements like tables, headings, and text layout.

Why does PDF conversion lose formatting and how can I keep the original structure?

PDF conversion often loses formatting because binary formats hinder content reuse, but this skill preserves the original structure. It converts PDFs into Markdown while specifically maintaining text layout, headings, and tables for accurate content extraction.

When do I need to extract text from technical reports into Markdown format?

You need to extract text from technical reports into Markdown format when you require accurate content extraction and readable formatting for collaboration and archiving. This skill converts academic papers, reports, and manuals from locked binary PDFs into editable Markdown.