elevenlabs-tts

Produce and quality-control speech from scripts using ElevenLabs text-to-speech models.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill elevenlabs-tts-calesthio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: elevenlabs-tts
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/elevenlabs-tts
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill elevenlabs-tts-calesthio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you turn scripts into credible, consistent, and delivery-ready speech while avoiding pronunciation errors, poor model choices, audible seams, unsupported controls, and voice-rights failures.

Core Features & Use Cases

  • Model and Voice Selection: Choose and audition ElevenLabs models and voices for narration, character performance, multilingual speech, streaming, and long-form production.
  • Directed Speech Production: Prepare spoken text, normalize numbers and brands, control pronunciation, pacing, emotion, code-switching, chunking, and API rendering behavior.
  • Production QA and Delivery: Manage captions, timing, manifests, retries, local repairs, loudness, true peak, artifact checks, listening evaluation, rights, consent, privacy, and disclosure.
  • Use Case: Create a 60-second product narration with exact brand and pricing pronunciation, aligned captions, a documented render manifest, and a localization plan for future French delivery.

Quick Start

Use the elevenlabs-tts skill to create and quality-control a polished narration from the provided script, including model and voice selection, pronunciation handling, render settings, captions, rights checks, and final QA.

Frequently Asked Questions about elevenlabs-tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create production-ready AI voiceovers without pronunciation errors?

Production-ready AI voiceovers require model and voice preflight, deterministic text normalization for numbers and brands, and listening-based QA to catch pronunciation errors. This Skill manages script preparation, render manifests, and delivery mastering to ensure credible, consistent speech output.

What is the best way to handle pronunciation repair and text normalization for text-to-speech?

Pronunciation repair and text normalization for text-to-speech require controlling pacing, emotion, and chunking while preparing spoken text. This Skill applies deterministic text normalization and supported pronunciation controls to ensure exact brand and pricing pronunciation in your voiceover.

Can I use ElevenLabs voice cloning and voice design for long-form narration?

Yes, ElevenLabs voice cloning and voice design support long-form narration and character performance when paired with model preflight. This Skill manages consent, privacy, and disclosure checks alongside render manifests and batch rendering for extensive media production.

Does this text-to-speech workflow support multilingual and code-switched media production?

Multilingual and code-switched media production is supported through model and voice selection tailored for diverse speech. This Skill handles code-switching controls, caption validation, and localization planning to ensure accurate delivery across languages.

How do I quality-control AI voiceover delivery for loudness, true peak, and artifacts?

Quality-controlling AI voiceover delivery for loudness, true peak, and artifacts requires listening-based QA and delivery mastering. This Skill manages artifact checks, local repairs, retry procedures, and render manifests to produce polished, delivery-ready speech.

What are the limitations of streaming text-to-speech rendering for real-time audio?

Streaming text-to-speech rendering requires handling API rendering behavior, chunking, and timing validation to avoid audible seams. This Skill applies retry and recovery procedures alongside caption alignment to manage constraints in streamed or batch rendering workflows.