wan-2-7

Generate text-to-video clips with optional audio-driven lip-sync via RunComfy.

12|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/runcomfy-com/skills --skill wan-2-7
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: wan-2-7
Source: https://github.com/runcomfy-com/skills/tree/main/wan-2-7
Command: npx skills add https://github.com/runcomfy-com/skills --skill wan-2-7

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Generate high-quality text-to-video clips with motion that matches the prompt, including audio-driven lip-sync when you provide a voice track.

Core Features & Use Cases

  • Audio-driven lip-sync: Supply an audio_url to drive speaking motions and timing in the generated video.
  • Prompt expansion and negative prompting: Improve output consistency with automatic prompt rewriting (toggleable) and targeted exclusions via negative_prompt.
  • Model routing for best results: Use Wan 2.7 when you explicitly ask for Wan/Wan 2.7, but route to alternatives (HappyHorse 1.0, Seedance 2.0 Pro, Kling Video O1, LTX 2) based on the user’s intent.

Quick Start

Ask for a Wan 2.7 talking-head video and include your audio track URL so the result is lip-synced to your voice.

Frequently Asked Questions about wan-2-7

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a lip-sync video from text and audio?

To generate a lip-sync video, provide a text prompt and an audio_url containing a 3–30 second WAV or MP3 file. The Wan 2.7 model processes these inputs to create a video where the subject's speaking motions match your supplied voice track.

What audio formats and size limits are supported for lip-sync video generation?

Lip-sync video generation supports WAV and MP3 audio files. The audio track must be between 3 and 30 seconds long and cannot exceed 15MB in size to be processed by the model.

Can I control the aspect ratio and resolution of text-to-video clips?

Yes, you can control the aspect ratio, resolution, and duration of your text-to-video clips. These configurable parameters allow you to tailor the output dimensions to fit specific cinematic or spokesperson video requirements.

How does prompt expansion improve text-to-video consistency?

Prompt expansion improves text-to-video consistency by automatically rewriting your input prompt. Toggling this feature on helps the model better understand the intent, resulting in generated motion that more accurately matches your requested scenario.

What is the best way to exclude unwanted elements from generated videos?

The best way to exclude unwanted elements is by using the negative_prompt parameter. Supplying targeted exclusions prevents the model from generating specific objects, styles, or motions you do not want in the final video.

When should I use Wan 2.7 instead of other text-to-video models?

Use Wan 2.7 when you explicitly need Wan 2.7 features like audio-driven lip-sync or prompt-controlled cinematic motion. The system routes requests to alternative models like HappyHorse 1.0, Seedance 2.0 Pro, Kling Video O1, or LTX 2 based on your specific video generation intent.