tts

Generate speech audio from text using Hume Octave, Inworld TTS, and Google Gemini TTS.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/mshuffett/dotfiles --skill tts-mshuffett
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tts
Source: https://github.com/mshuffett/dotfiles/tree/main/agents/knowledge/atoms/claude-skill-archive/text-to-speech
Command: npx skills add https://github.com/mshuffett/dotfiles --skill tts-mshuffett

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the creation of speech audio from text, enabling voiceovers, podcasts, and audio content generation.

Core Features & Use Cases

  • Multiple TTS Providers: Integrates with Hume Octave, Inworld TTS (MAX), and Google Gemini TTS.
  • Voice Selection & Emotion Control: Offers a wide range of voices and allows for emotional expression and style customization.
  • Use Case: Generate a podcast intro with a professional voice, create an audiobook narration, or produce voiceovers for explainer videos.

Quick Start

Use the tts skill to generate speech audio from the text "Hello, world!" using the Inworld TTS service with the Noah voice.

Frequently Asked Questions about tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate speech audio from text for a podcast?

To generate speech audio from text for a podcast, you can use this skill to synthesize voiceovers and multi-speaker audio content through providers like Hume Octave, Inworld TTS, and Google Gemini TTS, complete with emotional control and pre-defined voices.

Can I create multi-speaker podcast audio with emotional control using TTS?

Yes, you can create multi-speaker podcast audio with emotional control using TTS through this skill's integration with Hume Octave and Inworld TTS, which support dynamic voice generation, pre-defined voices, and style customization for expressive narration.

Do I need API keys to use Hume Octave, Inworld, and Google Gemini TTS services?

Yes, you need API keys to use Hume Octave, Inworld, and Google Gemini TTS services, as this skill requires separate authentication credentials for each provider to authenticate requests and generate speech audio from text.

What is the best way to produce an audiobook narration with different voices?

The best way to produce an audiobook narration with different voices is using this skill's voice selection features, allowing you to choose from a wide range of pre-defined voices across Inworld TTS, Hume Octave, and Google Gemini TTS to match your content's tone.

Does this speech synthesis skill support dynamic voice generation?

Yes, this speech synthesis skill supports dynamic voice generation, enabling you to create custom voice profiles and adjust emotional expression on the fly when generating audio content from text via the supported TTS providers.

What are the limitations when using multiple TTS providers for voiceover generation?

A key limitation for voiceover generation using multiple TTS providers is the hard dependency on external services; you must obtain and manage valid API keys for Hume, Inworld, and Google Gemini separately to ensure successful speech synthesis processing.