What problem does it solve? Creating presenter-style videos traditionally requires cameras, actors, and editing time. This Skill lets you generate talking head and avatar videos directly from a portrait image plus a script or audio file, using hosted AI models through the inference.sh CLI. ## Core Features & Use Cases - Multiple Avatar Models: Choose from P-Video-Avatar (fast, low-cost, built-in TTS with 30 voices and 10 languages), OmniHuman 1.5 (multi-character), Fabric 1.0, and PixVerse Lipsync. - Text-to-Avatar with Built-in TTS: P-Video-Avatar converts a text script directly into a spoken avatar video, with voice, language, style, and resolution controls up to 1080p. - Full Pipelines: Combine portrait generation, TTS, transcription, translation, and lipsync to build dubbing and localization workflows. - Use Case: A marketing team needs a product demo video in three languages. They generate a portrait with pruna/p-image, create the avatar video with P-Video-Avatar, then use Whisper transcription and LatentSync lipsync to dub localized versions. ## Quick Start Ask the AI to create a talking avatar video from a portrait image URL and a short voice script using the P-Video-Avatar model via the belt CLI.