wan-t2v-video

Generate videos from text prompts using dual-model UNETs and CLIP encoders.

530|85|Updated Feb 15, 2026
One-click install
npx skills add https://github.com/artokun/comfyui-mcp --skill wan-t2v-video
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: wan-t2v-video
Source: https://github.com/artokun/comfyui-mcp/tree/main/plugin/skills/wan-t2v-video
Command: npx skills add https://github.com/artokun/comfyui-mcp --skill wan-t2v-video

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill enables the creation of original video content directly from textual descriptions, bypassing the need for manual video editing or complex animation software.

Core Features & Use Cases

  • Text-to-Video Generation: Creates videos based on detailed text prompts, specifying scene, motion, and style.
  • Dual-Model Architecture: Utilizes specialized HighNoise and LowNoise models for structured motion and refined details.
  • Advanced Control: Supports VACE modules, Lightning LoRAs, and custom sampler settings for fine-tuning output quality and speed.
  • Use Case: Generate a short, cinematic clip of a futuristic cityscape at sunset, with flying vehicles and dynamic lighting, all from a descriptive prompt.

Quick Start

Use the wan-t2v-video skill to generate a video from the prompt 'A majestic dragon soaring over a medieval castle at dawn'.

Frequently Asked Questions about wan-t2v-video

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate AI video from text prompts using ComfyUI?

To generate AI video from text prompts, this Skill uses a dual-model architecture with specialized UNETs and CLIP encoders to synthesize motion. You provide a descriptive prompt, and it outputs a generated video clip without manual animation software.

How does dual-model architecture improve text-to-video generation?

Dual-model architecture improves text-to-video generation by using HighNoise models for structured motion and LowNoise models for refined details. This separation ensures the final AI video maintains both coherent movement and high visual fidelity.

What resolution and frame count limits apply to AI video synthesis here?

AI video synthesis here handles resolutions up to 720p and frame counts up to 121 frames. These limits ensure manageable processing times while maintaining detailed motion synthesis and output quality.

Can I use VACE modules and Lightning LoRAs for AI video generation?

Yes, you can use VACE modules and Lightning LoRAs for AI video generation to fine-tune output quality and speed. These advanced control features allow you to customize sampler settings alongside the specialized UNETs.

Do I need complex video editing software to create cinematic clips from text?

No, you do not need complex video editing software to create cinematic clips from text. This Skill bypasses manual editing by directly translating detailed text descriptions specifying scene, motion, and style into video content.