qwen3-tts

Generate and validate Qwen3-TTS speech with model, voice, and route selection.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill qwen3-tts-calesthio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: qwen3-tts
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/qwen3-tts
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill qwen3-tts-calesthio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps agents plan and produce reliable Qwen3-TTS speech without confusing hosted and open-weight deployment, selecting incompatible models or voices, mishandling temporary audio URLs, or overlooking consent, privacy, and delivery QA.

Core Features & Use Cases

  • Route and Model Selection: Choose between non-real-time HTTP or SSE synthesis, realtime WebSocket speech, instruction-controlled models, Voice Design, Voice Cloning, and open-weight checkpoints.
  • Production Voice Workflows: Create narration, advertisements, audiobooks, assistants, characters, multilingual content, and localized media with voice, language, dialect, prosody, and pronunciation guidance.
  • Artifact and Compliance Control: Download expiring audio immediately, preserve request and generation metadata, manage consent and provenance for custom voices, and validate pronunciation, performance, timing, loudness, and file integrity.
  • Use Case: For a SaaS launch video, select and audition a compatible Qwen voice, generate segmented English narration with confident and playful instruction control, download each result, record its manifest, and perform final audio and subtitle QA.

Quick Start

Use the qwen3-tts skill to create a polished 60-second English product-video narration with an appropriate Qwen3-TTS model, expressive direction, artifact manifest, rights notes, and final QA checklist.

Frequently Asked Questions about qwen3-tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate multilingual speech for audiobooks using text-to-speech?

To generate multilingual speech for audiobooks, select a compatible Qwen3-TTS model and route, specify the target language and dialect, apply prosody and pronunciation guidance, download the expiring audio URL immediately, and validate file integrity.

What's the best way to clone a voice for production-ready audio?

The best way to clone a voice for production-ready audio is using the Qwen3-TTS Voice Cloning route, which requires explicit consent management, provenance tracking, and technical audio QA to validate pronunciation, loudness, and performance.

Does Qwen3-TTS support realtime audio generation for voice assistants?

Yes, Qwen3-TTS supports realtime audio generation for voice assistants via its WebSocket speech route, offering an alternative to non-real-time HTTP or SSE synthesis for interactive applications.

How do I avoid mishandling temporary audio URLs when synthesizing speech?

To avoid mishandling temporary audio URLs during speech synthesis, you must download the generated audio immediately upon completion, preserve request metadata in an artifact manifest, and verify technical audio QA before the link expires.

Can I use instruction-controlled models to design character speech?

Yes, you can use Qwen3-TTS instruction-controlled models and Voice Design features to create character speech, applying expressive direction like confident or playful tones for narration and localized media.

What are the limitations of using open-weight checkpoints versus hosted speech synthesis?

Open-weight checkpoints require route-specific model and region verification, unlike hosted speech synthesis, and both deployment boundaries necessitate explicit language controls and rights safeguards to ensure production-ready compliance.