seedance-v2

Generate cinematic short-form videos with native lip-synced audio via RunComfy.

12|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/runcomfy-com/skills --skill seedance-v2
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: seedance-v2
Source: https://github.com/runcomfy-com/skills/tree/main/seedance-v2
Command: npx skills add https://github.com/runcomfy-com/skills --skill seedance-v2

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Users need a fast way to generate cinematic short-form videos with native lip-synced audio using Seedance 2.0 Pro, without manually stitching references, timing, and playback.

Core Features & Use Cases

  • Cinematic video generation with native lip-sync: Produces 4–15s outputs with synchronized in-pass audio from prompts and audio references.
  • Multimodal reference support (identity, scene, voice tone): Combines up to 9 images, up to 3 short video clips, and up to 3 short audio references to keep the result coherent.
  • Routing guidance vs sibling models: Uses Seedance for multimodal cinematic lip-synced shorts, while advising when to switch to HappyHorse, Wan, Kling, or LTX.

Quick Start

Tell the skill to create a 9:16 spokesperson-style video using Seedance 2.0 Pro with a prompt describing the shot and dialogue tone, optionally including an image reference for identity.

Frequently Asked Questions about seedance-v2

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate cinematic short videos with native lip-synced audio?

To generate cinematic short videos with native lip-synced audio, use Seedance 2.0 Pro via RunComfy CLI. Submit a JSON schema with your prompt, reference arrays, aspect ratio, and duration to produce synchronized 4–15s video outputs.

Can I use multiple image and audio references to keep a spokesperson video coherent?

You can combine up to 9 images, 3 video clips, and 3 audio references to maintain multimodal continuity. This ensures identity, scene, and voice tone remain coherent across your generated short-form video.

What is the best way to create a 9:16 spokesperson ad with synchronized dialogue?

The best way to create a 9:16 spokesperson ad with synchronized dialogue is running the Seedance 2.0 Pro endpoint. Provide a prompt describing the shot and dialogue tone, alongside an optional image reference for identity continuity.

Does Seedance 2.0 Pro support multimodal inputs like video clips and audio URLs?

Seedance 2.0 Pro supports multimodal continuity by accepting prompt, image_url, video_url, and audio_url inputs. This allows cohesive short sequences with native lip-synced in-pass audio generation.

When should I choose Seedance over other video generation models like Wan or Kling?

Choose Seedance for multimodal cinematic lip-synced shorts requiring native audio. The skill provides routing guidance to switch to HappyHorse, Wan, Kling, or LTX when your specific video generation needs differ from this use case.

What are the limitations of generating short-form video with Seedance 2.0 Pro?

Limitations of generating short-form video with Seedance 2.0 Pro include output durations constrained to 4–15 seconds. It also requires RunComfy CLI execution and a specific JSON input schema to properly process references and download produced URLs.