Kreuzberg

Extracts text, metadata, and structured information from diverse file formats into Markdown with tables and code blocks.

8|1|Updated Mar 14, 2026
One-click install
npx skills add https://github.com/codata/croissant-toolkit --skill kreuzberg-codata
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Kreuzberg
Source: https://github.com/codata/croissant-toolkit/tree/main/.gemini/skills/kreuzberg
Command: npx skills add https://github.com/codata/croissant-toolkit --skill kreuzberg-codata

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires kreuzberg, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables rapid extraction of text, metadata, and structured data from a wide array of document formats, streamlining data collection and analysis tasks.

Core Features & Use Cases

  • Wide Format Support: Handles PDF, Word, Excel, PowerPoint, and over 300 programming languages.
  • Structured Output: Produces Markdown with tables and code blocks, preserving original formatting.
  • Use Case: For instance, extracting code snippets, tables, and summaries from academic research papers or legal documents efficiently.

Quick Start

Use this Skill to extract and convert your document or code file into organized Markdown with minimal effort.

Frequently Asked Questions about Kreuzberg

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text and tables from a PDF document into Markdown?

To extract text and tables from a PDF into Markdown, this Skill parses files and preserves original formatting like tables and code blocks. It converts multi-format documents into structured Markdown output suitable for review and analysis.

What is the best way to extract code snippets from academic research papers?

The best way to extract code snippets from academic research papers is using a multi-format extraction tool. This Skill handles over 300 programming languages, converting document contents into organized Markdown efficiently.

Does this document extraction tool support OCR for scanned files?

Yes, document extraction supports OCR for scanned files. It features a Rust core for high-speed processing and utilizes OCR capabilities to accurately extract text and metadata from complex, multi-format data sources.

Can I extract metadata from Word, Excel, and PowerPoint files simultaneously?

Yes, you can extract metadata from Word, Excel, and PowerPoint files. This Skill provides wide format support to extract text, metadata, and structured data from diverse file formats, streamlining your data collection tasks.

How do I convert source code files into structured Markdown for review?

To convert source code files into structured Markdown, the Skill extracts code blocks while preserving original formatting. It supports over 300 programming languages, allowing you to organize code snippets for analysis and review.

Do I need Python libraries to run multi-format document extraction?

Yes, you need Python libraries to run multi-format document extraction. This Skill depends on Python libraries while leveraging a Rust core to deliver fast, accurate, and high-performance processing for complex data sources.