minimax-speech

Generate speech and narration through MiniMax APIs with production controls.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill minimax-speech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: minimax-speech
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/minimax-speech
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill minimax-speech

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps media teams produce reliable MiniMax speech deliverables without guessing at API capabilities, model choices, voice controls, consent requirements, temporary artifact windows, or production QA.

Core Features & Use Cases

  • Speech Generation: Plan and produce narration, advertising voiceovers, audiobook segments, localized dialogue, product walkthroughs, and interactive voice responses using MiniMax HTTP, WebSocket, and asynchronous T2A workflows.
  • Voice Production: Select and audition system voices, design custom voices, or clone approved voices with documented source-quality, consent, retention, and disclosure safeguards.
  • Production Readiness: Handle model selection, pronunciation, language, emotion, subtitle timestamps, pricing, rate limits, artifact custody, localization review, and final audio quality assurance.

Quick Start

Use the minimax-speech skill to create a production plan and MiniMax generation request for a premium 30-second SaaS voiceover with an audition workflow, word-level captions, rights notes, cost estimate, and final QA checklist.

Frequently Asked Questions about minimax-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate localized voiceovers using MiniMax text to speech APIs?

MiniMax text to speech generates localized voiceovers through HTTP, WebSocket, and asynchronous T2A workflows. You must select endpoint-aware models, audition system voices, configure pronunciation and language controls, and manage temporary artifact retrieval for final audio assets.

Can I clone a custom voice for audiobook narration with MiniMax?

Yes, MiniMax supports approved voice cloning for audiobook narration. You must document source-quality, verify consent, manage retention policies, apply disclosure safeguards, and complete rights checks before generating localized dialogue or narration assets.

What is the best way to plan production costs for advertising voiceovers?

Planning advertising voiceover costs requires MiniMax pricing and rate-limit analysis. You evaluate endpoint-aware model selection, estimate generation requests, audition voices, and map temporary artifact retrieval windows to avoid unmanaged production risks.

Does MiniMax speech generation support word-level subtitle timestamps?

Yes, MiniMax speech generation supports subtitle timing controls. You can produce word-level captions alongside audio assets, which requires configuring pronunciation, language, and emotion parameters during the T2A workflow generation request.

What are the limitations of using MiniMax for low-latency interactive assistants?

MiniMax low-latency interactive assistants require strict endpoint-aware model selection and rate-limit planning. Limitations include avoiding unsupported capabilities, managing temporary artifact retrieval windows, and mitigating unmanaged production risks during real-time voice responses.