comfyui-voice-pipeline

Automate character voice creation and lip-sync for video production.

86|24|Updated Feb 6, 2026
One-click install
npx skills add https://github.com/MCKRUZ/ComfyUI-Expert --skill comfyui-voice-pipeline
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: comfyui-voice-pipeline
Source: https://github.com/MCKRUZ/ComfyUI-Expert/tree/main/skills/comfyui-voice-pipeline
Command: npx skills add https://github.com/MCKRUZ/ComfyUI-Expert --skill comfyui-voice-pipeline

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Generates and synchronizes character voices for video projects using TTS, voice cloning, and lip-sync tools, reducing manual voice production effort and enabling consistent dialogue across media assets.

Core Features & Use Cases

  • Voice synthesis & cloning: Generate natural-sounding speech with multiple engines (Chatterbox, F5-TTS, TTS Audio Suite, RVC, ElevenLabs) and clone target voices from reference samples.
  • Lip-sync integration: Align generated speech with video footage using Wav2Lip, SadTalker, LivePortrait, and related post-processing pipelines for realistic mouth movements.
  • Character workflow: Create voice profiles for characters, manage multi-voice scenes, and reuse profiles across projects.

Quick Start

Specify a character voice and target video to generate audio and automatically lip-sync it to the footage.

Frequently Asked Questions about comfyui-voice-pipeline

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate character voice creation and lip-sync for video production?

Character voice creation and lip-sync for video production are automated by generating speech with TTS engines and aligning the audio to footage using Wav2Lip or SadTalker for realistic mouth movements.

Can I clone a specific voice and apply it to a video avatar?

Voice cloning allows you to replicate target voices from reference samples using engines like RVC or ElevenLabs, which can then be applied to lip-synced avatars for consistent character dialogue.

How does lip-sync integration work with generated TTS audio?

Lip-sync integration aligns generated TTS audio with video footage using post-processing pipelines like Wav2Lip, SadTalker, and LivePortrait to create realistic mouth movements for characters.

What's the best way to manage multi-voice scenes across different languages?

Multi-voice scenes across languages are managed by creating distinct voice profiles for characters, allowing you to reuse these profiles across projects and maintain consistent dialogue.

Does this TTS and voice cloning workflow support multiple speech engines?

The workflow supports multi-engine integration including Chatterbox, F5-TTS, TTS Audio Suite, RVC, and ElevenLabs for flexible voice synthesis, cloning, and post-processing refinement.

When do I need voice profiling workflows for video post-processing?

Voice profiling workflows are needed for production-grade voice design and post-processing refinement when your project requires managing multiple character voices and ensuring consistent dialogue across media assets.