TTS

Convert written text into spoken audio files with selectable voices and formats.

1|Updated Jan 3, 2026
One-click install
npx skills add https://github.com/MO196931/documentosZai --skill tts-mo196931
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: TTS
Source: https://github.com/MO196931/documentosZai/tree/main/skills/TTS
Command: npx skills add https://github.com/MO196931/documentosZai --skill tts-mo196931

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires z-ai-web-dev-sdk, and includes scripts (resource) components.

What problem does it solve?

This Skill eliminates the tedious manual work of generating spoken audio from text, which is a major bottleneck for content creators, educators, and developers building voice-enabled applications.

Core Features & Use Cases

  • Multi-Voice Speech Generation: Choose from 7 natural-sounding voices suited for different tones and scenarios.
  • Customizable Audio Output: Adjust speech speed (0.5x to 2x) and volume, and export to WAV, MP3, or PCM formats.
  • Use Case: An e-learning platform can use this Skill to automatically generate narration for course materials, or a content creator can turn long-form blog posts into podcast-ready audio files.

Quick Start

Use the TTS skill to convert the provided text into a natural-sounding speech audio file and save it as output.wav.

Frequently Asked Questions about TTS

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech audio for e-learning content?

To convert text to speech audio for e-learning content, this Skill processes written text input to generate natural-sounding spoken audio files. It provides 7 natural-sounding voices suited for different tones and scenarios.

Can I use z-ai-web-dev-sdk for batch audio generation?

Yes, you can use z-ai-web-dev-sdk for batch audio generation. This Skill supports batch audio processing to automatically generate multiple spoken audio files from text inputs for diverse application use cases.

How do I adjust speech speed and volume for voice synthesis?

You adjust speech speed and volume for voice synthesis by configuring the customizable audio output settings. This Skill allows you to set speech speed ranging from 0.5x to 2x and adjust volume levels before exporting.

What audio formats can I export from text-to-speech processing?

You can export WAV, MP3, or PCM formats from text-to-speech processing. This Skill provides customizable audio output options to save your generated natural-sounding speech in these standard formats.

Does this voice synthesis tool work for backend integration?

Yes, this voice synthesis tool works for backend integration. It integrates with the z-ai-web-dev-sdk to automate public announcement systems, voice assistant responses, and e-learning narration within your applications.