What problem does it solve? Creating voiceovers, audiobook narration, or podcast audio traditionally requires recording equipment, voice talent, and editing time. This Skill generates natural-sounding speech from text via the inference.sh CLI, covering multiple TTS models and voice styles. ## Core Features & Use Cases - Multiple TTS Models: Access ElevenLabs (22+ voices, 32 languages), Kokoro TTS, DIA, Chatterbox, Higgs, and VibeVoice through a unified CLI interface. - Voice Library & Control: Choose from American and British English voices with adjustable speed (0.8-1.2) and punctuation-based pacing control. - Long-Form & Multi-Voice Workflows: Chunk long texts, merge audio segments, create multi-speaker conversations, and combine voiceovers with video or talking-head avatars. - Use Case: A content creator needs narration for a 10-minute tutorial video. They generate the voiceover with a professional Kokoro voice, then merge it with their video using the media-merger app. ## Quick Start Use the ai-voice-cloning skill to generate a warm female voiceover reading my intro paragraph with the Kokoro af_sarah voice.