What problem does it solve? Creating voiceovers, audiobooks, podcasts, or narration traditionally requires recording equipment, voice talent, and editing time. This Skill lets you generate natural-sounding speech from plain text using multiple AI voice models through a single CLI. ## Core Features & Use Cases - Multiple TTS Models: Choose from ElevenLabs (premium, 22+ voices, 32 languages), DIA TTS (conversational), Kokoro (fast), Chatterbox, Higgs Audio (emotional control), and VibeVoice (long-form podcasts). - Expressive & Multi-Speaker Speech: Generate conversational dialogue, emotional delivery, and podcast-length audio. - Video Pipeline Integration: Combine generated speech with talking-head avatar tools like OmniHuman for narrated videos. - Use Case: Write a product demo script, run it through Kokoro TTS to produce a voiceover, then feed the audio into an avatar video generator for a complete narrated presentation. ## Quick Start Ask the assistant to generate speech from your text using the Kokoro TTS model via the belt CLI, for example: convert 'Welcome to our tutorial' into an audio file.