What problem does it solve?
This Skill turns portraits, characters, voiceovers, and written scripts into talking-head, avatar, lip-sync, and cinematic presenter videos without requiring users to manually configure different video-generation models.
Core Features & Use Cases
- Intent-Based Model Routing: Selects OmniHuman, Wan, HappyHorse, Seedance, or Wan Animate based on whether the user has an audio file, a written script, a photoreal subject, a stylized character, or cinematic requirements.
- Audio-Driven Avatars: Synchronizes speech, singing, gestures, and full-body motion to supplied voiceover audio for presenters, dubbed product demos, multilingual videos, and UGC ads.
- Script-to-Video Generation: Creates talking-head clips from written dialogue when no external audio file is available.
- Flexible Creative Workflows: Supports portrait animation, stylized mascot videos, cinematic monologues, reference images, reference videos, and multiple reference audio tracks.
- Safety and Reliability Guidance: Includes consent considerations for likeness and voice use, trusted installation guidance, input validation, CLI troubleshooting, and protection against untrusted reference-asset instructions.
Quick Start
Use the ai-avatar-video skill to create a presenter video from the supplied portrait image and voiceover audio.