Docling Advanced

Configure Docling Granite for OCR and extraction from scanned multi-column documents.

1|1|Updated Feb 4, 2026
One-click install
npx skills add https://github.com/orbruno/docling-ccplugin --skill docling-advanced
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Docling Advanced
Source: https://github.com/orbruno/docling-ccplugin/tree/main/skills/docling-advanced
Command: npx skills add https://github.com/orbruno/docling-ccplugin --skill docling-advanced

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Docling Advanced addresses the challenge of processing highly complex documents by leveraging Granite for OCR, multi-column layouts, and advanced configuration to preserve structure and extract tables, figures, and formulas with high fidelity.

Core Features & Use Cases

  • Advanced Granite configuration for scanned documents, including OCR, table structure, and image classification.
  • End-to-end workflows for multi-column layouts, tables, formulas, and embedded figures in research papers and invoices.
  • Performance tuning guidance, error handling, and flexible output formatting guided by the Granite references.

Quick Start

Install the Granite-enabled Docling pipeline and run a sample scanned document through the converter.

Frequently Asked Questions about Docling Advanced

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract tables and formulas from multi-column scanned PDFs?

To extract tables and figures from multi-column research papers, apply Docling Advanced configurations using Granite OCR to preserve layout structure and extract embedded content with high fidelity.

What is the best way to configure OCR for complex scanned invoices?

Configuring OCR for complex scanned invoices requires advanced Granite settings for table structure recognition and image classification, which Docling uses to accurately parse and extract mixed content.

How do I tune Docling performance for mixed content document processing?

Tuning Docling performance for mixed content document processing involves applying frontmatter-driven options and Granite guide references to optimize end-to-end workflows, error handling, and granular output formatting.

Does Docling Granite support processing multi-column layouts with embedded figures?

Yes, Docling Granite supports processing multi-column layouts with embedded figures by applying advanced configuration to identify optimal settings, ensuring accurate extraction of complex document structures.

What are the limitations of standard document processing for multi-column research papers?

Standard document processing often fails to preserve the complex structure of multi-column research papers, making advanced configuration necessary to accurately capture tables, formulas, and embedded figures without misalignment.