What problem does it solve?
Setting up MiniMax H3 (Hailuo) video generation locally in ComfyUI involves confusing choices between local-weight nodes and paid API nodes, multiple model files, turbo LoRAs, and nonstandard frame math, and mistakes produce broken graphs or wasted API credits.
Core Features & Use Cases
- Local vs API path selection: Distinguishes the free local-weight nodes (MiniMaxH3ImageToVideo, MiniMaxH3ReferenceToVideo) from the paid partner API nodes so the right cost model is used.
- Model and template setup: Lists the Comfy-Org INT8 diffusion models, Qwen3-VL text encoder, dual VAEs, and 4/8-step turbo LoRAs, plus how to load the official Template Library graphs.
- Output and chaining guidance: Specifies 24 fps, 17k+5 frame-length math, 768px sizing, VRAM tiers down to 8 GB, and last-frame chaining for clips longer than 15 seconds.
- Use Case: A user on a 12 GB GPU asks for a 10-second stereo-audio clip from a text prompt; the skill loads the T2V template, enables the turbo LoRA and Sage attention, sets the correct length, and queues the render.
Quick Start
Ask the agent to build a local MiniMax H3 text-to-video workflow in ComfyUI using the Comfy-Org INT8 template and turbo LoRA for a short clip with audio.