XiaomiMiMoXiaomiMiMoOfficial·1 Agent Skills Included

MiMo-Skills

Text-to-speech synthesis, voice cloning, and voice design

Generates natural speech from text using preset voices, custom voice design, or voice cloning from audio samples. Controls emotion, tone, dialects, and singing through plain-language instructions and inline audio tags. Eliminates manual audio editing and sends finished voice messages directly to Feishu chats.
npx skills add XiaomiMiMo/MiMo-Skills --all -g -y

All Skills in This Repository (1)

Pure Emerald Level Indicators

Frequently Asked Questions

FAQPage Schema
How to install MiMo-Skills?

Run `npx skills add XiaomiMiMo/MiMo-Skills --all -g -y` in your terminal to install the skill globally. You will also need a MiMo API key set as the MIMO_API_KEY environment variable.

How to convert text to speech with MiMo?

Ask your agent to read text aloud and it runs the mimo_tts.py script with your chosen preset voice. You can control emotion, speed, and style using plain-language instructions or inline tags.

Can MiMo clone a voice from an audio sample?

Yes. Provide an mp3 or wav sample under 10 MB and the voice cloning script reproduces that voice for any new text you supply.

Does MiMo TTS support singing and dialects?

Yes. Preset voices support singing via a simple tag, and you can apply dialects like Cantonese or Sichuanese plus dozens of emotion and tone styles.

Can I send generated voice messages to Feishu?

Yes. The included script converts the audio to opus, uploads it, and sends it as a native voice message to Feishu private or group chats.

Related Repositories in Content & Communication

View All in Content & Communication