What problem does it solve?
Creating a multi-person talking head podcast video normally requires cameras, actors, voice talent, and editing software. This Skill orchestrates the full AI pipeline — character image generation, text-to-speech, avatar animation, and video stitching — so you can produce a complete multi-speaker video from a script alone.
Core Features & Use Cases
- Character Creation: Generate real humans (with identity-consistent Phota profiles), brand mascots, or illustrated characters, including alternate angles and logo placement on clothing.
- Voice & Script Production: Write natural conversational scripts with duration targets, then generate TTS audio with per-character voices and controlled speaking rates.
- Avatar Video & Merge: Animate each character frame with its audio into talking head clips, then stitch all clips into a final MP4 video.
- Use Case: A marketing team wants a 60-second two-host podcast-style promo video. The Skill generates both host portraits, produces approved TTS voices, animates each turn sequentially, and merges the clips into one finished video.
Quick Start
Ask the agent to create a 60-second two-person podcast video about your product, starting by generating the two host characters and proposing voice samples for approval.