What problem does it solve? Raw talking-head, interview, or podcast footage lacks visual structure; this Skill layers designed, timed graphic overlays (titles, lower-thirds, data callouts, quotes, side panels, picture-in-picture) onto the untouched clip, synced to what is actually being said. ## Core Features & Use Cases - Transcript-driven card design: Extracts audio with ffmpeg, transcribes locally via Whisper (hyperframes transcribe), and derives card timing from word-level timestamps. - Visual design library: Ships 10 card styles, 4 layouts, and 3 video frames as self-contained HTML references, mixable across 16:9, 9:16, and 4:5 canvases. - GSAP composition rendering: Assembles per-card HTML fragments into one composition and renders to MP4 via the hyperframes CLI, with no third-party API keys. - Use Case: Take a 2-minute interview clip, auto-infer roughly 17 cards from its information density, pick a portrait 9:16 canvas with a stack layout, and produce a social-ready video with kinetic titles and data callouts. ## Quick Start Package my interview clip videos/interview.mp4 with graphic overlay cards and render it as a 9:16 portrait MP4.