lipsync

Synchronize facial mouth movements with audio to produce lip-synced video.

31|9|Updated Apr 30, 2026
One-click install
npx skills add https://github.com/prime-skills/runcomfy-agent-skills --skill lipsync-prime-skills
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: lipsync
Source: https://github.com/prime-skills/runcomfy-agent-skills/tree/main/lipsync
Command: npx skills add https://github.com/prime-skills/runcomfy-agent-skills --skill lipsync-prime-skills

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill eliminates the complexity of matching a face's mouth movements to spoken audio, helping you create convincing dubbed videos, talking avatars, and synchronized speech content through RunComfy.

Core Features & Use Cases

  • Video Lip-Sync: Apply a voiceover to existing footage while preserving the subject's appearance, motion, lighting, and background.
  • Avatar Generation: Turn a portrait and audio track into a talking-head or full-body avatar video.
  • Script-to-Video: Generate speech and synchronized video from a written script when no audio file is available.
  • Model Routing: Select among Sync Labs, ByteDance OmniHuman, Kling, Creatify, Wan, and HappyHorse based on the input format, quality needs, and budget.
  • Use Case: Dub a product video into multiple languages by pairing the original footage with translated voiceovers and selecting a premium lip-sync route.

Quick Start

Use the lipsync skill to synchronize the provided video or portrait with the specified audio track and save the generated result locally.

Frequently Asked Questions about lipsync

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I synchronize facial mouth movements with audio for video dubbing?

To synchronize facial mouth movements with audio for video dubbing, you provide an existing video and an audio track. The Skill routes the input through RunComfy models to generate a natural lip-synced video while preserving the subject's appearance and lighting.

Can I generate a talking avatar from a portrait and a voiceover script?

Yes, you can generate a talking avatar from a portrait and a voiceover script. The Skill supports script-to-video generation, creating synchronized speech and matching facial animation even when no pre-recorded audio file is provided.

What's the best way to create multilingual voiceovers for existing footage?

The best way to create multilingual voiceovers for existing footage is to pair the original video with translated audio tracks. The Skill applies audio-driven lip-sync to match the new speech with the subject's mouth movements, preserving the original background and motion.

Does this lip-sync tool work with stylized character animation?

Yes, this lip-sync tool works with stylized character animation. It synchronizes audio-driven facial animations for various formats, including portrait-based talking avatars and stylized characters, by routing inputs to compatible models like ByteDance OmniHuman or Kling.

What do I need to provide for audio-driven video generation?

For audio-driven video generation, you need accessible media URLs for your video or portrait, valid CLI authentication, and compatible audio and video durations. You must also obtain consent from the people whose faces or voices are used.

Why is my lip-synced video not generating correctly?

Your lip-synced video may not generate correctly if the audio and video durations are incompatible, CLI authentication is invalid, or media URLs are inaccessible. Ensure your RunComfy model routing matches your input format and quality requirements to avoid processing errors.