wan-2-7

Generate Wan 2.7 videos from text prompts via RunComfy CLI.

5|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/doany-ai/skills --skill wan-2-7-doany-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: wan-2-7
Source: https://github.com/doany-ai/skills/tree/main/wan-2-7
Command: npx skills add https://github.com/doany-ai/skills --skill wan-2-7-doany-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Wan-2-7 simplifies the creation of high-quality text-to-video content by providing a flexible, lip-sync capable model with multi-reference conditioning, enabling creators to produce cinematic video outputs directly within RunComfy.

Core Features & Use Cases

  • Multi-reference conditioning for richer motion control and cinematic quality.
  • Audio-driven lip-sync via audio_url to synchronize speech with visuals.
  • Prompt expansion enabled by default to improve output quality; easily disable for literal prompts.
  • Clear routing to alternative models (HappyHorse 1.0, Seedance 2.0 Pro, Kling, LTX 2) when a different capability is required.
  • Support for standard aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4) and resolutions (720p, 1080p); durations from 2–15s.

Quick Start

Provide a prompt and an optional audio_url to generate a Wan 2.7 video through RunComfy.

Frequently Asked Questions about wan-2-7

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a lip-synced video from text and audio?

To generate a lip-synced video, provide a text prompt and an optional audio_url to synchronize speech with visuals. The system processes these inputs via the RunComfy CLI to produce a cinematic video file.

What resolutions and aspect ratios are supported for text-to-video generation?

Text-to-video generation supports 720p and 1080p resolutions. It includes a 16:9 family of aspect ratios, specifically supporting 16:9, 9:16, 1:1, 4:3, and 3:4 formats.

Does prompt expansion improve cinematic video quality automatically?

Prompt expansion is enabled by default to improve cinematic video output quality. You can easily disable this feature if you require the model to strictly follow your literal text prompts.

What is multi-reference conditioning used for in video generation?

Multi-reference conditioning provides richer motion control and cinematic quality in video generation. It allows the model to process multiple visual references simultaneously to guide the output.

How long can a generated video clip be?

A generated video clip can be between 2 and 15 seconds long. You can specify the exact duration within this range when configuring your text-to-video generation parameters.

When should I use alternative video generation models instead of Wan 2.7?

You should use alternative models like HappyHorse 1.0, Seedance 2.0 Pro, Kling, or LTX 2 when a different capability is required. Wan 2.7 is optimized for lip-syncing and multi-reference cinematic outputs.