What problem does it solve?
Creating digital human videos (talking photos, lip-synced dubs, face swaps) normally requires stitching together multiple AI models, TTS engines, and billing logic by hand. This Skill orchestrates the full DashScope pipeline so an agent can produce a finished MP4 from a photo and a sentence of text.
Core Features & Use Cases
- Five generation modes: photo_speak (talking photo), video_relip (lip-sync replacement), video_reface (person swap), avatar_compose (multi-image character fusion), and pose_drive (motion transfer), with cost preview before any paid submission.
- Multi-backend and dual TTS: runs on Alibaba DashScope, RunningHub, or local ComfyUI, with CosyVoice or free Edge-TTS for speech synthesis, plus voice cloning and figure libraries.
- Use Case: A user says "make this headshot say a 5-second welcome message in Chinese" — the Skill runs face detection, synthesizes speech, submits the wan2.2-s2v job, estimates the cost (about ¥0.31), and returns the task id with the finished MP4.
Quick Start
Ask the agent to make the attached portrait photo speak the sentence "Hello, welcome to my channel" using the default voice at 480P resolution.