What problem does it solve?
This Skill converts approved first frames and tagged A-roll script segments into scroll-stopping talking-head video clips with synced lip movement and dialogue audio, reducing the anti-polish effort needed to make AI video feel human.
Core Features & Use Cases
- A-roll clip generation (image-to-video with dialogue): Produces one talking-head MP4 clip per
[A-ROLL] script segment using Sora 2 as the primary provider.
- Provider fallback chain for reliability: Automatically falls back from Sora 2 (fal.ai) to Kling 3, and then to Replicate Kling 3 if needed; you can force Kling by using
--provider kling.
- Structured motion prompting: Uses a labeled prompt schema (Camera/Subject/Dialogue/Audio/Environment & light/Style & mood) to keep the result realistic (iPhone/selfie feel) and reduce content-filter failures.
- Generation logging and brand memory integration: Reads the canonical
frame1.png, the [A-ROLL] script, and creator references; writes MP4 outputs and append-only output-log.md entries.
Quick Start
Run the animate workflow for A-roll by generating a portrait clip from your approved frame1.png using motion-prompt.txt and an A-roll script that includes [A-ROLL] segments.