TTS

Generate spoken audio from written text with configurable voices and formats.

1|Updated May 1, 2026
One-click install
npx skills add https://github.com/e2662020/QuickMovie --skill tts-e2662020
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: TTS
Source: https://github.com/e2662020/QuickMovie/tree/main/skills/TTS
Command: npx skills add https://github.com/e2662020/QuickMovie --skill tts-e2662020

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires z-ai-web-dev-sdk, and includes scripts (resource) components.

What problem does it solve?

This Skill eliminates the need for manual voice recording to generate spoken audio from written text, saving time and resources for content creators, developers, and teams building voice-enabled applications or accessibility features.

Core Features & Use Cases

  • Multi-Voice Support: Choose from 7 distinct natural-sounding voices for different tones and use cases.
  • Configurable Audio Parameters: Adjust speech speed (0.5x to 2x) and volume to match your content needs.
  • Flexible Output & Processing: Generate audio in WAV, MP3, or PCM formats, with support for streaming and batch processing of long text content.
  • Use Case Example: A podcast producer can use this Skill to convert full episode scripts into narrated audio files in minutes, without hiring voice actors.

Quick Start

Use the TTS skill to convert the text Welcome to QuickMovie into a 1.2x speed audio file named welcome.mp3 using the tongtong voice.

Frequently Asked Questions about TTS

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate natural speech audio from text for backend integration?

Generate natural speech audio by passing written text input to this Skill, which leverages the z-ai-web-dev-sdk for backend integration. It supports configurable voice selection, adjustable speech speed, and multiple output formats for automated notification systems.

What audio formats can I output when converting text to speech?

Text-to-speech conversion supports WAV, MP3, and PCM audio output formats. The Skill also enables streaming playback and batch processing to handle long-form written text input efficiently.

Can I adjust voice speed and volume for text-to-speech generation?

Yes, text-to-speech generation supports configurable audio parameters. You can choose from 7 distinct natural-sounding voices and adjust speech speed from 0.5x to 2x to match your content needs.

Does the z-ai-web-dev-sdk support batch audio processing for long text?

Yes, this Skill uses the z-ai-web-dev-sdk to support batch audio processing for long text. This enables efficient generation of extensive spoken audio content like audiobooks and e-learning narration.

What is the best way to automate audiobook production from written scripts?

Automate audiobook production by using this Skill to convert full written scripts into narrated audio files. It eliminates manual voice recording by offering 7 natural voices, batch processing, and adjustable speech speed.