document-extractor

Convert PDFs, Word, slides, spreadsheets, images, HTML, CSV, JSON, XML, ZIP, EPUB, Outlook, and YouTube content into Markdown.

4|Updated Mar 5, 2026
One-click install
npx skills add https://github.com/lirrensi/agent-cli-helpers --skill document-extractor
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: document-extractor
Source: https://github.com/lirrensi/agent-cli-helpers/tree/main/skills/document-extractor
Command: npx skills add https://github.com/lirrensi/agent-cli-helpers --skill document-extractor

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Converting hard-to-read documents into clean, structured Markdown reduces manual review time and makes content searchable and machine-friendly for downstream workflows.

Core Features & Use Cases

  • Converts PDFs, Word, PowerPoint, Excel, images, HTML, CSV, JSON, XML, ZIP, EPUB, Outlook, and YouTube inputs into readable Markdown.
  • Enables easy feeding of documents into LLMs, data pipelines, and automation scripts.
  • Provides installation guidance to ensure MarkItDown is ready for use in various environments.

Quick Start

Convert a sample file such as report.pdf to Markdown for quick inspection.

Frequently Asked Questions about document-extractor

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert PDFs and Word documents to Markdown for LLM workflows?

You can convert PDFs and Word documents to Markdown by using this tool to normalize content for AI workflows. It transforms various office documents into clean, structured Markdown, making them readable and machine-friendly for downstream pipelines.

Can I extract YouTube transcriptions and convert them to Markdown?

Yes, you can extract YouTube content and convert it to Markdown. The tool supports YouTube transcription extraction alongside other formats like HTML, CSV, and JSON, normalizing video content into readable text for automated processing.

What do I need to install before converting documents to Markdown?

You need to install MarkItDown via uv, pipx, or other supported package managers before converting documents. The Skill provides installation guidance to ensure the environment is ready for processing files like PDFs, Excel, and EPUB into Markdown.

Does this approach work with spreadsheets, images, and ZIP files?

Yes, this approach works with spreadsheets, images, and ZIP files. It supports a wide range of formats including Excel, PowerPoint, CSV, XML, and EPUB, extracting and normalizing their content into structured Markdown for inspection.

What is the best way to batch convert Outlook items and EPUB files into clean Markdown?

The best way to convert Outlook items and EPUB files into clean Markdown is using this extraction tool. It reduces manual review time by transforming hard-to-read formats into structured, searchable text suitable for data pipelines.

Why convert HTML and JSON files into Markdown instead of processing them directly?

Converting HTML and JSON files into Markdown eases inspection and makes content machine-friendly. Structured Markdown reduces manual review time and integrates smoothly with LLMs and automation scripts compared to processing raw formats directly.