vllm-omni-audio-tts

Generate speech and audio with vLLM-Omni using Qwen3-TTS, MiMo-Audio, and Stable-Audio.

84|27|Updated Mar 3, 2026
One-click install
npx skills add https://github.com/hsliuustc0106/vllm-omni-skills --skill vllm-omni-audio-tts
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: vllm-omni-audio-tts
Source: https://github.com/hsliuustc0106/vllm-omni-skills/tree/main/skills/vllm-omni-audio-tts
Command: npx skills add https://github.com/hsliuustc0106/vllm-omni-skills --skill vllm-omni-audio-tts

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Generate multi-model audio content and TTS outputs by orchestrating Qwen3-TTS, MiMo-Audio, and Stable-Audio within vLLM-Omni, simplifying cross-model workflows for speech and sound generation.

Core Features & Use Cases

  • TTS & Voice Cloning: Convert text to realistic speech including cloning voices from reference samples.
  • Voice Design & Audio Generation: Design unique voice personas and compose audio with music and sound effects.
  • Use Case: Create a narrated video with a cloned presenter voice and synchronized background music using a single API.

Quick Start

Provide a text prompt and select a model to generate speech or audio using Omni.

Frequently Asked Questions about vllm-omni-audio-tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate speech and audio using multiple models like Qwen3-TTS and Stable-Audio?

Generate speech and audio by orchestrating Qwen3-TTS, MiMo-Audio, and Stable-Audio within vLLM-Omni. This approach simplifies cross-model workflows for text-to-speech, voice cloning, and sound effect generation through a unified API.

Can I clone a voice from a reference sample for text-to-speech?

Voice cloning from reference samples is supported for text-to-speech generation. You provide a text prompt and reference audio to generate realistic speech that mimics the cloned voice persona using the configured models.

What is the best way to create narrated audio with synchronized background music?

Create narrated audio with synchronized background music by using multi-model audio generation. You can combine cloned presenter speech with audio composed of music and sound effects generated by Stable-Audio and MiMo-Audio.

Do I need to configure an Omni client and server to run text-to-audio tasks?

You need Omni client and server setup to run text-to-audio tasks. The workflow requires configuring the listed models and environment details described in the documentation for both offline and server modes.

Does vLLM-Omni support voice design and unique voice persona creation?

vLLM-Omni supports voice design and unique voice persona creation. You can design custom voices and generate audio content tailored to specific character profiles or application requirements using Qwen3-TTS.