tts

Convert input text into speech audio files with configurable voices and formats.

Updated May 30, 2026
One-click install
npx skills add https://github.com/zeroix07/mcp-skill-agent --skill tts-zeroix07
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tts
Source: https://github.com/zeroix07/mcp-skill-agent/tree/main/tts
Command: npx skills add https://github.com/zeroix07/mcp-skill-agent --skill tts-zeroix07

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires z-ai-web-dev-sdk, and includes scripts (resource) components.

What problem does it solve?

This Skill eliminates the time and cost of manual voice recording, letting you generate natural-sounding speech audio from any text in seconds without hiring voice actors or setting up recording equipment.

Core Features & Use Cases

  • Multi-Voice Support: Choose from 7 distinct voice options to match your content tone, from warm and friendly to professional and clear.
  • Customizable Audio Output: Adjust playback speed (0.5x to 2x), volume, and output format (WAV, MP3, PCM) to fit your needs.
  • Scalable Content Generation: Process single text snippets or batch multiple scripts for audiobooks, e-learning modules, accessibility content, and automated announcements.
  • Use Case: If you need to create audio versions of 20 blog posts, use this Skill to generate consistent, high-quality narration for all posts in a single batch run.

Quick Start

Use the tts skill to convert the provided product announcement text into a 1.3x speed MP3 audio file saved as announcement.mp3.

Frequently Asked Questions about tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech for e-learning content creation?

To convert text to speech for e-learning content, this Skill synthesizes input text into natural-sounding audio files. It eliminates manual voice recording by generating speech with configurable voice options, playback speed, and output formats. You can process text snippets or batch multiple scripts to produce consistent narration for modules.

Can I generate audio in different formats and speeds for accessibility?

Yes, you can generate accessibility audio in WAV, MP3, or PCM formats with adjustable playback speeds ranging from 0.5x to 2x. This text-to-speech Skill allows you to configure volume and select from 7 distinct voice options to match your content's tone and meet diverse audio generation requirements.

Does this text-to-speech tool support long text chunking for audiobook production?

Yes, this text-to-speech tool supports long text chunking for audiobook production. It processes extensive scripts by dividing them into manageable segments, ensuring consistent and high-quality speech synthesis across large-scale content generation tasks without manual intervention.

What's the best way to batch process multiple scripts into voice audio?

The best way to batch process multiple scripts into voice audio is using this Skill's scalable content generation feature. It handles multiple text inputs in a single run, applying your chosen voice and speed settings uniformly to produce consistent speech files for automated announcements or blog post narrations.

Do I need the z-ai-web-dev-sdk to use this voice synthesis Skill?

Yes, you need the z-ai-web-dev-sdk dependency installed to use this voice synthesis Skill. The SDK provides the underlying framework required to execute the text-to-speech scripts, convert input text into audio, and manage the configurable voice selection and output settings.