What problem does it solve? Media content like videos, voice notes, PDFs, and screenshots cannot enter a knowledge base until it is transcribed, OCR-extracted, or visually described. This Skill converts media into structured, entity-linked brain pages so nothing gets lost. ## Core Features & Use Cases - Multi-format extraction pipelines: Transcribe YouTube videos and audio via Whisper, OCR PDFs with pymupdf or marker-pdf, and describe screenshots with vision analysis. - Entity extraction and cross-linking: Identify people, companies, projects, and dates, then link new pages to existing brain entries and add timeline entries. - Verbatim preservation: Voice notes and direct quotes are stored with exact phrasing, while long-form content is summarized. - Use Case: A user sends a voice note saying a client wants to renew in Q3. The Skill transcribes it verbatim, links it to the client's company page, and adds a timeline entry for the renewal signal. ## Quick Start Ingest this voice note into gbrain, transcribe it exactly, extract the entities, and link it to the matching company page.