pdf

Extract text and structured data from PDF documents into report-ready outputs.

12|7|Updated Jan 28, 2026
One-click install
npx skills add https://github.com/georgemarmelstein/sistema-marmelstein --skill pdf-georgemarmelstein
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pdf
Source: https://github.com/georgemarmelstein/sistema-marmelstein/tree/main/.claude/spec/templates
Command: npx skills add https://github.com/georgemarmelstein/sistema-marmelstein --skill pdf-georgemarmelstein

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It eliminates the tedious manual work of extracting information from PDFs and turning document content into usable, structured outputs.

Core Features & Use Cases

  • PDF-to-text extraction: Converts PDF content into editable text for downstream analysis and drafting.
  • Structured data extraction: Helps derive report-ready structure from document content for easier review.
  • Repeatable workflow: Suitable for recurring document intake where accuracy and consistency matter.

Quick Start

Use the skill when you need to extract information from a PDF named file 'contrato.pdf' and want the result as a structured report for review.

Frequently Asked Questions about pdf

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract structured data from a PDF contract?

To extract structured data from a PDF contract, you need a tool-based workflow that reads the input PDF from the workspace, parses the content, and writes the resulting structured artifacts back to disk without inventing missing details.

What is the best way to automate PDF data extraction for administrative records?

Automating PDF data extraction for administrative records involves using a repeatable workflow that converts document content into editable text and report-ready structure for consistent parsing and review.

Can I convert PDF content to text for downstream document processing workflows?

Yes, you can convert PDF content to text for downstream document processing workflows by applying an extraction process that transforms the document into usable text and structured outputs.

How do I handle repeated PDF intake tasks with consistent parsing?

Handling repeated PDF intake tasks requires a repeatable workflow that reads each input PDF from the workspace, performs consistent parsing, and writes the resulting structured artifacts back to disk.

Does PDF extraction work with workspace files directly?

Yes, PDF extraction works directly with workspace files by reading the input PDF from the disk and writing the resulting text and structured artifacts back to the workspace.

What are the limitations of extracting data from contract PDFs?

A key limitation of extracting data from contract PDFs is that the workflow requires a safe, tool-based approach and will not invent missing details, meaning incomplete documents will yield incomplete structured outputs.