What problem does it solve?
Creating presenter-style videos traditionally requires cameras, actors, and editing time. This Skill turns a single portrait image into a speaking avatar video with built-in text-to-speech, removing the need for filming or separate voiceover production.
Core Features & Use Cases
- Text-to-Avatar Generation: Provide a portrait image and a voice script to produce a lip-synced talking head video at 720p or 1080p.
- Audio-Driven Avatars: Supply your own audio file to drive the avatar instead of using the built-in TTS with 30 voices across 10 languages.
- Style and Behavior Control: Use voice_prompt and video_prompt parameters to control tone, pacing, emotion, backgrounds, and body language.
- Use Case: A marketing team generates a product demo video by creating a portrait with pruna/p-image, then animating it with a scripted voiceover in Spanish and English versions for localized campaigns.
Quick Start
Ask the AI to generate a talking avatar video from a portrait image URL with a short welcome script using the p-video-avatar skill.