text-to-speech

Convert text into speech using multiple TTS models like Kokoro TTS.

688|95|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/inference-sh/skills --skill text-to-speech-inference-sh
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: text-to-speech
Source: https://github.com/inference-sh/skills/tree/main/tools/audio/text-to-speech
Command: npx skills add https://github.com/inference-sh/skills --skill text-to-speech-inference-sh

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the need to convert written text into spoken audio, enabling content creators and developers to generate natural-sounding speech for various applications without manual recording.

Core Features & Use Cases

  • Multi-Model Support: Utilizes diverse TTS engines like DIA TTS, Kokoro TTS, Chatterbox, Higgs Audio, and VibeVoice.
  • Voice Cloning & Expressiveness: Offers capabilities for voice cloning and fine-grained emotional control in speech.
  • Use Case: Generate a podcast episode script into an audio file using VibeVoice, or create a voiceover for a video with expressive narration using Higgs Audio.

Quick Start

Use the text-to-speech skill to convert the text 'Hello, welcome to our product demo.' into speech using the Kokoro TTS model.

Frequently Asked Questions about text-to-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech for a podcast episode?

You can convert text to speech for a podcast by using the VibeVoice model, which generates natural-sounding audio directly from your episode script. It supports multi-speaker dialogue generation and expressive speech synthesis for realistic audio output.

Can I clone a voice for text to speech generation?

Yes, voice cloning is supported during text to speech generation, allowing you to replicate specific vocal characteristics. This works alongside fine-grained emotional control to produce expressive speech synthesis tailored to your voiceover or audiobook needs.

What is the best way to generate expressive voiceovers from text?

The best way to generate expressive voiceovers from text is using models like Higgs Audio, which provide fine-grained emotional control during speech synthesis. This enables expressive narration tailored for video voiceovers and accessibility applications.

Does this text to speech tool support multi-speaker dialogue generation?

Yes, this text to speech tool supports multi-speaker dialogue generation natively. You can assign different voices within a single text script to automatically produce conversational audio, making it ideal for podcast generation and interactive voice assistant responses.

Which AI voice models work with this speech synthesis tool?

This speech synthesis tool works with multiple AI voice models including DIA TTS, Kokoro TTS, Chatterbox, Higgs Audio, and VibeVoice. Each model offers distinct capabilities for voice cloning, multi-speaker dialogue, and expressive speech generation.

Can I use text to speech for IVR systems and video narration?

Yes, you can use text to speech generation for IVR systems and video narration. The supported AI models convert written scripts into natural-sounding speech, providing automated voice responses and expressive narration without requiring manual audio recording.