What problem does it solve? Producing expressive voiceover narration for videos requires a TTS provider that supports emotion control, many languages, and reusable cloned voices, and this Skill provides the exact API usage, model selection, and cost guidance for fish.audio. ## Core Features & Use Cases - Model Selection: Choose between s2.1-pro, s2.1-pro-free, s2-pro, and s1 backends, with guidance on quality, emotion-tag support, and per-byte billing. - Voice Cloning Reuse: Pass a playground-created voice model as reference_id to reuse cloned voices across narration jobs. - Emotion & Prosody Control: Use inline emotion tags like [laugh] or [whispers] plus latency, temperature, speed, and bitrate tuning for delivery control. - Use Case: A video producer drafts narration on the free promotional model, gets voice approval from a 10-15 second sample, then renders the final hero narration on s2.1-pro. ## Quick Start Generate a short sample narration with the fish.audio provider using model s1 and my playground voice reference_id, saving the audio to my project's assets folder.