What problem does it solve? Turning raw talking-head footage and b-roll into a finished vertical UGC post requires many error-prone editing steps — cutting retakes and filler, choosing a layout, syncing b-roll, and placing captions — where small mistakes (wrong track index, unscoped captions, wrong aspect ratio) destroy the timeline. ## Core Features & Use Cases - Transcript-driven cutting: Reads the full word list and removes retakes, filler words, false starts, and dead air in a single pass. - Layout selection and application: Chooses between straight intercut, stacked split (b-roll top or bottom), or floating overlay, and applies it via apply_layout with fallback transforms. - Scoped captioning: Generates styled captions (highlightPop/wordPop) scoped to A-roll only, placed at the seam or lower third depending on format. - Use Case: A creator has a 69-second phone recording plus product demo clips and wants a 30-second TikTok ad — the skill sets 9:16, cuts bad takes, tiles muted b-roll on a new track, applies a stacked split, and adds gold-highlight captions. ## Quick Start Edit my raw talking-head footage and b-roll clips into a captioned 9:16 UGC video with a stacked split layout.