seedance-v2

Generate lip-synced cinematic videos from image, video, and audio references.

5|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/doany-ai/skills --skill seedance-v2-doany-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: seedance-v2
Source: https://github.com/doany-ai/skills/tree/main/seedance-v2
Command: npx skills add https://github.com/doany-ai/skills --skill seedance-v2-doany-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Create cinematic short-form videos that preserve identity across multiple media references and deliver native lip-synced audio without manual editing.

Core Features & Use Cases

  • Multi-modal references: combine up to 9 images, 3 videos, and 3 audio clips to drive a single scene.
  • In-pass lip-sync: generate natural-sounding speech synchronized to references.
  • Routing flexibility: switch to other RunComfy models as needed (e.g., HappyHorse, Wan, Kling) for different outputs.
  • Simple 4–15s schema: concise, deterministic prompts that produce short, repeatable results.

Quick Start

Provide a prompt and optional media URLs to generate a lip-synced cinematic video using Seedance 2.0 Pro.

Frequently Asked Questions about seedance-v2

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a lip-synced cinematic video from multiple images and audio clips?

To generate a lip-synced cinematic video, provide a prompt along with up to 9 image references, 3 video references, and 3 audio clips. The tool outputs a 4–15 second 720p video with synchronized native speech.

Can I use audio references to drive lip-sync in multi-modal video generation?

Yes, you can use up to 3 audio references to drive natural-sounding speech synchronization. The generation process applies in-pass lip-sync to align the visual output with the provided audio inputs.

Does multi-modal video generation work for multi-language marketing narratives?

Multi-modal video generation is built for marketing, branding, and multi-language narratives. It preserves identity across image, video, and audio references while delivering native lip-synced audio without manual editing.

What is the maximum duration and resolution for short-form cinematic video generated from references?

Short-form cinematic videos generated from multi-modal references are output at 720p resolution with a duration of 4 to 15 seconds. This concise schema ensures repeatable, deterministic results for short scenes.

Do I need manual video editing to achieve lip-sync with multi-modal references?

No manual video editing is needed to achieve lip-sync with multi-modal references. The tool applies in-pass lip-sync during generation, automatically aligning natural-sounding speech to the visual output.

When should I switch to other models instead of Seedance for multi-modal video generation?

You should switch to other RunComfy models like HappyHorse, Wan, or Kling when you need routing flexibility for different video outputs. Seedance 2.0 Pro handles 4–15s lip-synced scenes, but alternative models may suit other specific generation requirements.