livekit-tts

Configure multi-provider TTS models for LiveKit agents with voice customization.

1|Updated Mar 11, 2026
One-click install
npx skills add https://github.com/FutureAtoms/claude-skills-backup --skill livekit-tts
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: livekit-tts
Source: https://github.com/FutureAtoms/claude-skills-backup/tree/main/livekit-tts
Command: npx skills add https://github.com/FutureAtoms/claude-skills-backup --skill livekit-tts

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires livekit.plugins.cartesia, livekit.plugins.elevenlabs, livekit.plugins.openai, livekit.plugins.google, livekit.plugins.azure, dotenv, numpy, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill simplifies the process of configuring and integrating various Text-to-Speech (TTS) models for LiveKit agents, enabling dynamic and high-quality voice synthesis.

Core Features & Use Cases

  • Multi-Provider Support: Integrates with popular TTS services like Cartesia, ElevenLabs, OpenAI, Google Cloud, and Azure.
  • Voice Customization: Allows detailed control over voice selection, speed, stability, and emotion.
  • Advanced Features: Supports voice cloning, emotion control, word-level timing, and pronunciation adjustments.
  • Use Case: An agent needs to respond to user queries with a specific emotional tone and a custom voice. This skill allows configuring the agent to use ElevenLabs with a cloned voice and a "positive" emotion setting.

Quick Start

Configure the LiveKit TTS agent to use the Cartesia Sonic provider with the 'sweet_lady' voice.

Frequently Asked Questions about livekit-tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I configure text-to-speech for LiveKit agents?

Configuring text-to-speech for LiveKit agents involves integrating provider plugins like Cartesia or ElevenLabs. This enables fine-grained control over voice selection, speed, and emotion within dynamic audio agent sessions.

Which TTS providers can I use with LiveKit for voice synthesis?

LiveKit supports multiple text-to-speech providers for voice synthesis, including Cartesia, ElevenLabs, OpenAI, Google Cloud, and Azure. These plugins enable detailed customization of voice output within real-time agent sessions.

Can I use voice cloning and emotion control in LiveKit TTS?

Yes, LiveKit TTS supports advanced features like voice cloning and emotion control. You can configure agents to use specific providers, such as ElevenLabs, with custom cloned voices and targeted emotional tone settings.

Does LiveKit TTS support word-level timing and pronunciation adjustments?

Yes, LiveKit TTS supports advanced features including word-level timing and pronunciation adjustments. These capabilities enable developers to fine-tune real-time audio output for highly dynamic agent responses.

What is the best way to integrate ElevenLabs voice synthesis into a LiveKit agent?

The best way to integrate ElevenLabs voice synthesis into a LiveKit agent is by using the ElevenLabs plugin. This enables detailed configuration of voice stability, speed, and custom cloned voices for real-time agent sessions.