What problem does it solve? Writing prompts for MiniMax H3 video generation requires strict field structures, reference labels, keyframe alignment, and audio sections that are hard to compose correctly by hand. This Skill rewrites multimodal requests into the exact H3 prompt format so generated videos match the intended timeline, shots, and sound. ## Core Features & Use Cases - Five Input Modes: Supports T2VA (text-to-video), I2VA (first-frame), FL2VA (first-and-last-frame), L2VA (last-frame), and full-reference Ref2VA rewrites. - Structured Output Fields: Produces integrated_multimodal_description, overall_soundscape, and non_diegetic_music for base modes, plus subject_definitions, summary, retention_analysis, and detailed_description for Ref2VA. - Reference Label Management: Defines and tracks <Subject N>, <Picture N>, <Video N>, and <Audio N> labels consistently across all sections, including speaker IDs and dialogue formatting. - Use Case: Given a first-frame image and a request for an 8-second clip, the Skill emits an I2VA prompt with the frame-alignment instruction, shot-by-shot camera motion, dialogue in <d> tags, and matching soundscape sections. ## Quick Start Ask the agent to rewrite your video idea into an H3 prompt, specifying the mode (T2VA, I2VA, FL2VA, L2VA, or Ref2VA), the target duration, and any reference images, videos, or audio.