What problem does it solve?
This Skill helps agents plan and produce reliable Qwen3-TTS speech without confusing hosted and open-weight deployment, selecting incompatible models or voices, mishandling temporary audio URLs, or overlooking consent, privacy, and delivery QA.
Core Features & Use Cases
- Route and Model Selection: Choose between non-real-time HTTP or SSE synthesis, realtime WebSocket speech, instruction-controlled models, Voice Design, Voice Cloning, and open-weight checkpoints.
- Production Voice Workflows: Create narration, advertisements, audiobooks, assistants, characters, multilingual content, and localized media with voice, language, dialect, prosody, and pronunciation guidance.
- Artifact and Compliance Control: Download expiring audio immediately, preserve request and generation metadata, manage consent and provenance for custom voices, and validate pronunciation, performance, timing, loudness, and file integrity.
- Use Case: For a SaaS launch video, select and audition a compatible Qwen voice, generate segmented English narration with confident and playful instruction control, download each result, record its manifest, and perform final audio and subtitle QA.
Quick Start
Use the qwen3-tts skill to create a polished 60-second English product-video narration with an appropriate Qwen3-TTS model, expressive direction, artifact manifest, rights notes, and final QA checklist.