music-video-subtitle-generator

Generates beat-synced multi-shot MV prompts and lyric typography plans from music and lyrics.

38|4|Updated Jan 22, 2026
One-click install
npx skills add https://github.com/Qo-qiao/ComfyUI-omni-llm --skill music-video-subtitle-generator-qo-qiao
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: music-video-subtitle-generator
Source: https://github.com/Qo-qiao/ComfyUI-omni-llm/tree/main/skills/music-video-subtitle-generator
Command: npx skills add https://github.com/Qo-qiao/ComfyUI-omni-llm --skill music-video-subtitle-generator-qo-qiao

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve? Creating AI music videos with on-screen lyric typography requires coordinating beat timing, lyric mapping, shot breakdown, and visual consistency across clips, which is difficult to plan manually. This Skill turns music, lyrics, and reference images into a locked multi-shot prompt script with stitching rules so generated clips assemble into a coherent MV. ## Core Features & Use Cases - Beat-Synced Shot Planning: Analyzes vocal timing and drum hits to split videos longer than 15 seconds into 2-5 second shots mapped to lyric timestamps and the beat grid. - Spatial Lyric Typography: Designs on-screen text as a dynamic graphic layer that reacts to snare hits, 808 drops, and hi-hat rolls, with word-for-word matching to performed lyrics. - Reference Role Separation: Assigns character, scene, and typography reference cards to isolated jobs so visual elements do not cross-contaminate. - Use Case: A musician uploads a 30-second trap track and lyrics, confirms a 9:16 vertical format, and receives a complete multi-shot prompt script with transition logic that video and editing agents use to generate and stitch a finished MV. ## Quick Start Create a 30-second vertical MV prompt script with lyric typography from my uploaded trap beat and these lyrics, split into beat-synced shots.

Frequently Asked Questions about music-video-subtitle-generator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an AI music video with lyric typography?

Provide your music track, lyrics, and optional reference images, then confirm the target duration and aspect ratio. The Skill locks the lyrics, splits the video into beat-mapped shots, and writes a complete multi-shot prompt script to a canvas node for generation agents.

How to make videos longer than 15 seconds with AI video models?

Use a multi-shot stitching workflow that splits the video into 2-5 second shots mapped to the beat grid. Each shot uses the previous shot's tail frame as its head frame for continuity, and all clips are aligned to one master audio track during editing.

Can I use this Skill outside the MiniMax Hub environment?

No, the Skill requires the MiniMax Hub agent with its canvas workspace and Hub generation and routing tools. It is not portable to generic agent harnesses because locked prompts must be written to Hub canvas text nodes.

What reference images do I need for MV generation?

You can provide up to three reference cards with isolated roles: a character card for identity and wardrobe, a scene card for environment and lighting, and a typography card for text style and motion. Each card controls only its assigned aspect to prevent cross-contamination.

When should I not use this MV prompt generator?

Avoid it for ordinary subtitle burn-in, generic video editing, non-music product ads, or simple single-image and single-clip requests without MV structure. It is designed for stylized music visuals where lyrics, rhythm, and typography must be designed together.

Why does the Skill require locking lyrics before generating prompts?

Locked lyrics are the single source text for both vocal performance and visible typography, ensuring on-screen words match the performed lines word-for-word. Without locking, generated shots could show text that diverges from the actual sung or rapped lyrics.