kreuzberg

Extract text, tables, metadata, and images from 75+ file formats.

8.9k|539|Updated Jan 31, 2025
One-click install
npx skills add https://github.com/kreuzberg-dev/kreuzberg --skill kreuzberg
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: kreuzberg
Source: https://github.com/kreuzberg-dev/kreuzberg/tree/main/skills/kreuzberg
Command: npx skills add https://github.com/kreuzberg-dev/kreuzberg --skill kreuzberg

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Kreuzberg is a high-performance document intelligence library that enables programmatic extraction of text, tables, metadata, and images from 75+ formats across PDFs, Office documents, images (with OCR), HTML, email, archives, and academic formats.

Core Features & Use Cases

  • Efficient multi-format extraction: pull text, tables, and metadata from diverse document types for indexing, search, and downstream processing.
  • OCR-enabled data capture: automatically recognize text in scanned images and image-based PDFs via configurable backends.
  • Batch processing and extension points: process multiple files, with plugin support for post-processors, validators, and OCR backends.

Quick Start

Process a sample document to extract structured content and inspect results.

Frequently Asked Questions about kreuzberg

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text and tables from multiple document formats in one batch process?

Batch document extraction pulls text, tables, metadata, and images from 75+ file formats in a single run. Kreuzberg supports processing multiple files simultaneously with configurable options, enabling efficient indexing and downstream analysis across Python, Node.js, Rust, and CLI environments.

Does OCR work on scanned images and image-based PDFs for text extraction?

OCR enables text extraction from scanned images and image-based PDFs via configurable backends. Kreuzberg automatically recognizes text in image files, making previously inaccessible scanned content available for search indexing and downstream document processing workflows.

What's the best way to extract structured content from Office documents and HTML for search indexing?

Multi-format extraction from Office documents, HTML, email, and academic formats provides structured text and metadata ideal for search indexing. Kreuzberg handles diverse document types programmatically, enabling robust document intelligence pipelines without manual format-specific preprocessing.

Can I use language detection and embeddings for document chunking after extracting text?

Language detection and embeddings support document chunking after text extraction. Kreuzberg's extraction pipeline includes these capabilities to enable downstream analysis, ensuring extracted content from 75+ formats is properly segmented and processed for indexing workflows.

Are there plugin extensions for custom OCR backends and post-processing validators?

Plugin architecture supports custom OCR backends, post-processors, and validators for document extraction. Kreuzberg provides extension points that allow developers to customize the extraction pipeline, adding validation logic and specialized processing for document-heavy workflows.

What formats are supported for extracting metadata and images beyond standard PDFs?

Extraction supports 75+ formats including PDFs, Office documents, images, HTML, email, archives, and academic formats. Kreuzberg extracts metadata and images from these diverse file types, enabling comprehensive document intelligence across mixed-format collections.