What problem does it solve?
Audio transcripts lack visual context, making it hard to understand the full story of a video. This Skill enriches transcripts with detailed visual descriptions, providing a comprehensive overview for intelligent editing.
Core Features & Use Cases
- Frame Extraction: Uses FFmpeg to intelligently extract key frames from video files.
- AI Visual Analysis: Analyzes extracted frames to generate descriptive text about subjects, settings, and actions.
- Visual Transcript Creation: Integrates visual descriptions directly into the audio transcript, creating a "visual transcript" for rough cut generation.
- Use Case: After transcribing a product review video, use this skill to add descriptions like "Close-up of product packaging" or "User demonstrating feature" at relevant timestamps, making the transcript much more useful for editing.
Quick Start
Analyze the video at '/path/to/my/product_review.mov' and add visual descriptions to its transcript for the 'product-launch' library.