ai-avatar-video

Generate avatar talking-head videos from portrait images and voice input via RunComfy models.

12|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/runcomfy-com/skills --skill ai-avatar-video-runcomfy-com
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-avatar-video
Source: https://github.com/runcomfy-com/skills/tree/main/ai-avatar-video
Command: npx skills add https://github.com/runcomfy-com/skills --skill ai-avatar-video-runcomfy-com

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Turns a user’s intent into an AI avatar or talking-head video so they can speak, present, and lip-sync without manually producing or editing video.

Core Features & Use Cases

  • Routes to the best model by intent: selects between OmniHuman, Wan 2-7 (audio_url), Wan 2-2 Animate, HappyHorse (script-generated audio), and Seedance v2 Pro for cinematic setups.
  • Supports audio-driven avatar creation: generates speaking motion synchronized to user-provided voiceover audio or to script-generated speech.
  • Covers practical media scenarios: UGC voiceover, virtual presenter, dubbed product demos, lip-synced characters, “make this portrait talk,” and cinematic monologues with reference subject/audio.

Quick Start

Use ai-avatar-video to generate an avatar talking-head by giving a portrait URL and a voiceover URL, with the agent routing to OmniHuman automatically.

Frequently Asked Questions about ai-avatar-video

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an AI avatar video from a portrait image and voiceover?

To create an AI avatar video, provide a portrait image URL and a voiceover audio URL. The system routes your request to a model like OmniHuman to generate a lip-synced talking-head video.

Can I generate a virtual presenter video using a script instead of pre-recorded audio?

Yes, you can generate a virtual presenter using a script. The system routes your request to models like HappyHorse, which generates speech from your script and synchronizes it with a portrait image.

What is the best way to produce a cinematic monologue with an AI talking head?

The best way to produce a cinematic monologue is to use a reference subject and audio. The system routes your request to the Seedance v2 Pro model for cinematic avatar setups.

How does the system choose the right model for my lip sync or voiceover request?

The system chooses the right model for your lip sync request by evaluating your input intent. It automatically selects between OmniHuman, Wan 2-7, Wan 2-2 Animate, HappyHorse, and Seedance v2 Pro based on your inputs.

Do I need the RunComfy CLI to generate a talking head video?

Yes, you need the RunComfy CLI to generate a talking head video. It requires a CLI invocation with the correct model endpoint and JSON input fields such as image_url, audio_url, and prompt.