What problem does it solve?
Manually creating notes from videos that sync screenshots with transcript text is extremely time-consuming, especially for long talks or presentations with dense slide content. This Skill automates the entire end-to-end workflow to produce a ready-to-use markdown document in minutes.
Core Features & Use Cases
- End-to-end video processing: Automatically downloads videos from URLs or uses local files, extracts slide frames and key visual moments, transcribes audio, and merges everything into a single timestamped markdown file.
- Three quality modes: Choose between fast swift mode for quick first-pass notes, context mode for OCR-enriched slide text integration, or polish mode for LLM-refined, high-quality deliverables.
- Use cases: Perfect for generating study notes from recorded lectures, creating shareable recaps of conference talks, or building searchable archives of video content with synced visuals and text.
Quick Start
Use the multimodal-extraction skill to turn the attached video file or URL into a markdown timeline with synced screenshots and full transcript text.