What problem does it solve?
Manually transcribing text from PDFs, scanned documents, images, and other digital file formats is time-consuming and prone to human error, especially for security analysts, researchers, and professionals who regularly work with large volumes of unstructured document content.
Core Features & Use Cases
- Lightweight Local Extraction: Uses pymupdf for fast, low-overhead text, table, and image extraction from text-based PDFs, EPUBs, and other common document formats, with no large model downloads required.
- High-Accuracy OCR & Layout Analysis: Uses marker-pdf for scanned documents, OCR across 90+ languages, equation extraction, and complex layout parsing when working with non-text-based or poorly formatted files.
- Use Case: A security investigator can quickly extract full text from a scanned incident report PDF, or a researcher can pull the complete content of an Arxiv paper to reference during threat analysis, without manual transcription.
Quick Start
Use the ocr-and-documents skill to extract all editable text and tables from the attached local scanned incident report PDF.