elevenlabs-mcp

Generate speech, transcribe audio, and manage voices via the ElevenLabs MCP Server.

2|Updated Feb 13, 2026
One-click install
npx skills add https://github.com/pmarashian/cursor-agent-skills --skill elevenlabs-mcp
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: elevenlabs-mcp
Source: https://github.com/pmarashian/cursor-agent-skills/tree/main/elevenlabs-mcp
Command: npx skills add https://github.com/pmarashian/cursor-agent-skills --skill elevenlabs-mcp

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill streamlines complex audio tasks, from generating natural-sounding speech and music to transcribing audio and managing custom voices, all through a unified interface.

Core Features & Use Cases

  • Text-to-Speech (TTS): Convert text into realistic speech with various voices and models.
  • Speech-to-Text (STT): Transcribe audio files with speaker diarization.
  • Voice Management: Clone existing voices or create new ones from descriptions.
  • AI Agents: Build conversational AI agents with voice capabilities and knowledge bases.
  • Music Composition: Generate music from text prompts or structured plans.
  • Audio Processing: Isolate audio, transform voices, and play audio files.
  • Phone Integration: Make outbound calls using AI agents.

Quick Start

Use the elevenlabs-mcp skill to convert the text "Hello, world!" into speech using the Adam voice.

Frequently Asked Questions about elevenlabs-mcp

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech using AI voices?

Text-to-speech generation converts written text into realistic speech using various AI voices and models. You can generate natural-sounding audio by providing text prompts and selecting specific voice profiles.

Can I clone a custom voice for audio generation?

Voice cloning allows you to replicate existing voices or create new ones from text descriptions. This enables custom voice generation for text-to-speech workflows without requiring original voice actors.

How does speech-to-text transcription handle multiple speakers?

Speech-to-text transcription processes audio files with speaker diarization, which identifies and separates different speakers in the recording. This provides structured text output distinguishing who spoke when.

What is the best way to generate music from text prompts?

Music composition from text prompts generates original audio tracks based on your descriptive input or structured plans. This creates custom music without requiring manual composition or production tools.

Can I build conversational AI agents with voice capabilities?

AI agent creation builds conversational interfaces with voice capabilities and integrated knowledge bases. These agents can handle spoken interactions and be deployed for outbound phone call integration.

Are there API cost warnings for specific audio processing operations?

Yes, API cost warnings apply to specific audio processing operations like voice cloning and phone call integration. You must configure the MCP server and monitor usage to manage operational costs effectively.