What problem does it solve?
This Skill streamlines the creation of high-quality AI voiceover audio from plain text so teams can produce narration, dialogue, and multilingual voice tracks without manual recording or complex audio tooling.
Core Features & Use Cases
- Generate speech: Produce 24 kHz, 16-bit mono WAV audio from text using Gemini TTS.
- Multi-speaker dialogue: Create two-speaker conversations with per-speaker voice selection and style control.
- Asset tooling: Measure exact durations, inspect audio quality and pronunciation, and iterate by chunking to meet timeline targets.
- Use Case: Create podcast snippets, video narration, character dialogue for games, or timed voiceover tracks for editor timelines.
Quick Start
Call the generate_speech function with your script text and optional voice_name or speakers to save a WAV audio file (for example, call generate_speech with text set to Welcome to our product tour).