kreuzberg

Extract text, tables, metadata, and images from over 75 document formats.

1|Updated Mar 5, 2026
One-click install
npx skills add https://github.com/novvoo/skill-router --skill kreuzberg-novvoo
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: kreuzberg
Source: https://github.com/novvoo/skill-router/tree/main/agent/skills/kreuzberg
Command: npx skills add https://github.com/novvoo/skill-router --skill kreuzberg-novvoo

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill automates the extraction of text, tables, metadata, and images from over 75 document formats, saving significant time and effort in data processing.

Core Features & Use Cases

  • Universal Document Parsing: Handles PDFs, Office docs, images (with OCR), HTML, emails, archives, and more.
  • Structured Data Output: Extracts text, tables, and metadata into easily usable formats.
  • Advanced Configuration: Supports OCR, chunking, embeddings, language detection, and more for tailored extraction.
  • Use Case: Automatically process a batch of scanned invoices to extract key information like vendor name, amount, and date, making them searchable and analyzable.

Quick Start

Use the kreuzberg skill to extract text from the document 'report.pdf'.

Frequently Asked Questions about kreuzberg

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from scanned PDFs and images using OCR?

OCR text extraction from scanned PDFs and images captures textual data from visual document formats. This Skill processes over 75 document formats, using OCR to extract text, tables, and metadata into structured, searchable output.

Can I parse tables and metadata from Office documents and emails?

Table and metadata parsing from Office documents and emails extracts structured data from diverse file types. It supports Office formats, HTML, emails, and archives, outputting text, tables, and metadata into easily usable formats.

Does this document extraction tool support chunking and language detection?

Document extraction supports chunking and language detection as advanced configuration features. It handles custom plugins for tailored data processing, extracting text while applying embeddings and language detection for optimized workflows.

What is the best way to automate data processing for batch invoices?

Automating data processing for batch invoices extracts key information like vendor names, amounts, and dates. This Skill processes scanned documents in batches, making extracted invoice data fully searchable and analyzable.

How do I extract content from HTML files and email archives?

Content extraction from HTML files and email archives pulls text and metadata from web and communication formats. It processes HTML, emails, and archive files, outputting structured data for downstream analysis and content intelligence.