Aidenwu0209Aidenwu0209Communityยท2 Agent Skills Included

PaddleOCR-Skills

Extract text and structure from images, scans, and PDFs

Extracts text from screenshots, photos, scans, and PDFs with line-level accuracy and optional bounding boxes. Converts complex documents into clean Markdown and JSON, preserving tables, formulas, figures, and reading order. Handles Chinese and CJK text, multi-column layouts, and large PDFs without manual copy-pasting or retyping.
npx skills add Aidenwu0209/PaddleOCR-Skills --all -g -y

All Skills in This Repository (2)

Pure Emerald Level Indicators

Frequently Asked Questions

FAQPage Schema
How to install PaddleOCR-Skills?โ–ผ

Run `npx skills add Aidenwu0209/PaddleOCR-Skills --all -g -y` in your terminal to install both skills globally for your agent.

How to extract text from images and PDFs?โ–ผ

The text recognition skill pulls exact text from screenshots, photos, scans, and PDFs, including Chinese and handwritten content. Just give your agent the file path or URL and it returns the full text.

Can it convert PDFs to Markdown with tables?โ–ผ

Yes. The document parsing skill rebuilds tables, formulas as LaTeX, figures, and multi-column layouts into structured Markdown or JSON with correct reading order.

Does PaddleOCR-Skills work with Claude Code and Cursor?โ–ผ

Yes. Both skills follow the universal SKILL.md standard and run in Claude Code, Codex, Cursor, GitHub Copilot, OpenCode, and OpenClaw.

Do I need an API key to use PaddleOCR-Skills?โ–ผ

Yes. You need a free access token and endpoint URL from paddleocr.com, set as environment variables. The skills guide you through configuration on first run.

Related Repositories in Content & Communication

View All in Content & Communicationโ†’