document-reader

Extract text and structured content from documents and images using GLM-OCR.

5|Updated Apr 15, 2026
One-click install
npx skills add https://github.com/47network/Sven --skill document-reader-47network
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: document-reader
Source: https://github.com/47network/Sven/tree/main/skills/ocr/document-reader
Command: npx skills add https://github.com/47network/Sven --skill document-reader-47network

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill extracts text from documents and images using GLM-OCR with multi-language, table, handwriting, math detection.

Core Features & Use Cases

  • Multi-language OCR and layout-aware text extraction for PDFs, images, and scans.
  • Table recognition, handwriting, math and code region detection for structured data capture.
  • Use Case: Digitize paper documents into searchable text and structured data for analytics.

Quick Start

Upload a document or image and ask the skill to extract the text in plain text or JSON.

Frequently Asked Questions about document-reader

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from scanned PDFs and images?

To extract text from scanned PDFs and images, this skill uses GLM-OCR to perform layout-aware text extraction. It processes documents and images across multiple languages, converting them into searchable text and structured data formats like text, markdown, json, or html.

Does OCR text extraction work with multiple languages?

Yes, the OCR text extraction supports multiple languages. It processes scanned documents, images, and handwritten notes across various languages, enabling you to digitize and archive international paper documents into searchable text or structured json data.

What is the best way to digitize paper documents for analytics?

The best way to digitize paper documents for analytics is using layout-aware OCR to extract structured content. This skill captures text, tables, and code regions from scans and photos, outputting searchable markdown or json data ready for downstream analysis.

Do I need GLM-OCR to extract structured content from forms?

Yes, GLM-OCR integration is required to extract structured content from forms. It provides the multi-language, handwriting, and table recognition capabilities needed to capture complex layout regions and output them as text, markdown, json, or html.

What output formats are available when extracting text from images?

When extracting text from images, available output formats include plain text, markdown, json, and html. This allows you to capture layout-aware text, tables, and handwritten notes, converting them into structured data suitable for searchable archives.