kreuzberg

Extract text, tables, metadata, and images from 91+ document formats.

52|4|Updated Oct 23, 2025
One-click install
npx skills add https://github.com/vaayne/agent-kit --skill kreuzberg-vaayne
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: kreuzberg
Source: https://github.com/vaayne/agent-kit/tree/main/skills/kreuzberg
Command: npx skills add https://github.com/vaayne/agent-kit --skill kreuzberg-vaayne

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Kreuzberg automates the extraction of text, tables, metadata, and images from 91+ document formats, enabling rapid digitization and analysis of diverse files.

Core Features & Use Cases

  • Extract text, tables, metadata, and embedded images from PDFs, Office documents, HTML, emails, archives, and more.
  • Perform batch extractions, OCR-enabled data capture, and embeddings-ready outputs for semantic search and analytics.
  • Use cases include document digitization, data ingestion for analytics pipelines, and structured data retrieval via LLMs.

Quick Start

Run Kreuzberg on a sample document to extract text, tables, metadata, and images.

Frequently Asked Questions about kreuzberg

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text and tables from multiple document formats for batch processing?

To extract text and tables from multiple document formats for batch processing, this tool automates ingestion across 91+ formats including PDFs, Office documents, HTML, and emails. It handles batch extraction efficiently and outputs structured data ready for analytics pipelines.

What is the best way to prepare document extractions for semantic search and embeddings?

The best way to prepare document extractions for semantic search and embeddings is to use a tool that outputs structured, chunked text. This extraction Skill provides configurable chunking options and embeddings-ready outputs directly from 91+ document formats.

Can I use OCR to extract text from scanned PDFs and images?

Yes, you can use OCR to extract text from scanned PDFs and images. This extraction tool supports OCR backends to capture text from image-heavy files, ensuring accurate digitization and metadata retrieval for analytics.

Does this document extraction tool support archives and email formats?

Yes, this document extraction tool supports archives and email formats. It extracts text, tables, metadata, and embedded images from a wide range of 91+ formats including HTML, emails, and Office documents.

How do I configure chunking and output formats for extracted document data?

You can configure chunking and output formats for extracted document data using configurable CLI options. The tool allows you to specify output formats and chunking parameters to tailor the extraction results for server use or analytics ingestion.