What problem does it solve? Raw talking-head, interview, or podcast footage lacks visual structure; this Skill packages an existing clip with timed, designed graphic overlay cards (kinetic titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture) synced to the spoken transcript, while the original video plays untouched underneath. ## Core Features & Use Cases - Transcript-driven card design: Extracts audio with ffmpeg, transcribes locally via Whisper (hyperframes transcribe), and derives card timing and content from word-level timestamps. - Visual design library: Ships 10 card styles, 4 layouts, and 3 video frames as self-contained HTML references that can be freely mixed across 16:9, 9:16, and 4:5 canvases. - HTML-to-MP4 rendering: The agent writes each card's HTML fragment, assembles a GSAP-driven composition, and renders the final MP4 via the hyperframes CLI. - Use Case: Take a 2-minute founder interview clip, auto-infer roughly 17 cards from its information density, pick a portrait 9:16 canvas with a stack layout, and render a social-ready video with animated takeaway cards. ## Quick Start Package my talking-head video interview.mp4 with designed graphic overlay cards synced to what the speaker says, rendered for TikTok in 9:16.