What problem does it solve? Watching long videos to extract key information is time-consuming, and different videos need different levels of analysis. This Skill automates transcription and analysis of videos from YouTube, Loom, Vimeo, Riverside, Zoom recordings, social platforms, and local files, letting you choose between a fast free transcript, a visual frame-by-frame analysis, or full multimodal AI ingestion. ## Core Features & Use Cases - Three depth modes: transcript (fast, free text extraction), visual (transcript plus ffmpeg frame extraction and Claude vision pass on key moments), and multimodal (Gemini native video ingestion or dense Claude vision). - Multi-source support: Downloads via yt-dlp from YouTube, Loom, Vimeo, X/IG/TikTok, or accepts local MP4/MOV/WebM files, preferring platform-provided transcripts before falling back to local MLX-Whisper transcription. - Structured outputs: Saves transcript, metadata, key moments, and a summary with flagged action items and decisions to an organized workdir, with optional capture to a second-brain notes vault. - Use Case: A teammate shares a 45-minute Loom walkthrough of a new feature. Run the visual mode to get a transcript, timestamped key moments of UI changes, and a summary with action items — without watching the full recording. ## Quick Start Ask the AI to watch and summarize a video by providing its URL, for example: transcribe and summarize this Loom recording at https://www.loom.com/share/abc123 using visual mode.