What problem does it solve? Adding styled captions to talking-head videos normally requires manual editing, keyframing, and rotoscoping. This Skill automates the full pipeline locally: it transcribes speech, segments the subject from the background, and composites captions into the scene without altering the original footage. ## Core Features & Use Cases - 35-style identity catalog: Pick one visual identity (e.g. cream, ink, anchor, terminal, vhs) from a single catalog; the engine, compiler, and authoring file are derived automatically. - Rail + embed caption model: A verbatim lower-third rail carries most text while scarce peak words are composited behind the subject using matte occlusion for a cinematic depth effect. - Local end-to-end pipeline: One prepare script runs subject matting, WhisperX transcription, and safe-zone analysis in parallel, followed by JSON authoring, fast preview-frame QA, and a gated render to final.mp4. - Use Case: Given a 30-second founder update clip, probe the footage, pick the keynote identity, author a small JSON of caption blocks, preview composite frames, and render a captioned video with the climax word embedded behind the speaker. ## Quick Start Ask the agent to add captions to your talking-head video file and let it probe the clip, recommend an identity from the catalog, and render the captioned result.