ai-avatar-video

Create avatar videos from audio or scripts via RunComfy model routing.

31|9|Updated Apr 30, 2026
One-click install
npx skills add https://github.com/agentspace-so/runcomfy-agent-skills --skill ai-avatar-video-agentspace-so
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-avatar-video
Source: https://github.com/agentspace-so/runcomfy-agent-skills/tree/main/ai-avatar-video
Command: npx skills add https://github.com/agentspace-so/runcomfy-agent-skills --skill ai-avatar-video-agentspace-so

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It solves the problem of quickly creating realistic avatar and talking-head videos from a voiceover or a script, without manually juggling multiple tools and model-specific inputs.

Core Features & Use Cases

  • Model routing by intent: chooses the best RunComfy avatar route (OmniHuman, Wan 2-7 with audio_url, HappyHorse, Seedance v2 Pro, or Wan 2-2 Animate) based on whether you have an audio file or only text and what style you want.
  • Audio-driven lip-sync options: supports portrait+MP3 lip-sync (OmniHuman), prompt+audio_url scene generation (Wan 2-7), and script-driven in-pass speech (HappyHorse).
  • Cinematic, multi-modal generation: composes reference subject visuals and reference audio for more cinematic outcomes (Seedance v2 Pro).

Quick Start

Ask your AI to generate an audio-driven avatar video by providing a portrait image URL and a voiceover audio URL, then run the skill with those inputs so it selects the correct RunComfy route automatically.

Frequently Asked Questions about ai-avatar-video

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create a talking-head video from a portrait and an audio file?

To create a talking-head video, provide a portrait image URL and a voiceover audio URL, and the skill routes your inputs to the appropriate RunComfy avatar model for portrait lip-sync. It generates a video matching the audio track to the portrait.

Can I generate an avatar video from just a text script without a pre-recorded MP3?

Yes, you can generate an avatar video from a text script using the script-to-video speech-insertion route. This synthesizes speech from your script and drives the character animation without requiring a separate MP3 file.

What's the best way to achieve cinematic avatar generation with reference visuals?

For cinematic avatar generation, use the Seedance v2 Pro route to compose reference subject visuals and reference audio. This produces more cinematic outcomes by combining multi-modal reference inputs.

Do I need the RunComfy CLI to use audio-driven lip sync models?

Yes, you need to use the RunComfy CLI with the correct model endpoints and JSON input schema for image_url, audio_url, and prompt fields. The CLI is required to route your request to the correct avatar model.

How does the model routing decide which RunComfy route to use for my video?

Model routing selects the RunComfy avatar route by checking whether you have an audio file or only text, and your desired style. It automatically chooses between OmniHuman, Wan 2-7, HappyHorse, Seedance v2 Pro, or Wan 2-2 Animate.

Can I generate stylized character animation from an image using prompt and audio_url?

Yes, you can generate stylized character animation by providing a prompt and an audio_url to the Wan 2-7 route. This creates a scene driven by both your text prompt and the provided audio track.