What problem does it solve?
This Skill removes the complexity of integrating multiple multimodal media APIs by providing a single, consistent interface to create voice, music, images, and video through MiniMax MCP and REST fallbacks, while managing local outputs and async tasks.
Core Features & Use Cases
- Unified MCP toolset: text-to-audio, voice cloning, AI voice design, music generation, text/image-to-video, asynchronous video generation, and image generation accessible from one skill.
- Local file management & async workflows: configurable output directory, support for returning task IDs for long-running video generation, and tools to query task status.
- REST fallback and configuration: automatic MCP installation guidance, environment-driven configuration of MINIMAX_API_KEY and host, and REST API endpoints for cases where MCP is unavailable.
- Use cases: produce narrated voiceovers, clone a speaker's timbre from samples, generate background music from lyrics, create short promotional videos, or batch-generate images for creative assets.
Quick Start
Ask the agent to "Generate a 6-second seaside sunset video with TTS narration and save outputs to my MiniMax-Output directory."