elevenlabs

Generate voiceovers, sound effects, and music from text prompts and audio samples.

1|Updated Apr 11, 2026
One-click install
npx skills add https://github.com/shige1014-dev/backup-OpenMontage --skill elevenlabs-shige1014-dev
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: elevenlabs
Source: https://github.com/shige1014-dev/backup-OpenMontage/tree/main/.agents/skills/elevenlabs
Command: npx skills add https://github.com/shige1014-dev/backup-OpenMontage --skill elevenlabs-shige1014-dev

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill centralizes audio generation for video and multimedia projects, eliminating the need for manual recording, complex audio assembly, or sourcing royalty-free assets by producing TTS, sound effects, music, and voice clones from descriptions and samples.

Core Features & Use Cases

  • Text-to-Speech: produce scene-level voiceovers with selectable models and voice settings for stability, similarity, style, and speed.
  • Voice Cloning: create instant or professional voice clones from sample audio for consistent narration across projects.
  • Sound Effects & Music: generate short SFX and longer musical tracks with prompt influence, instrument control, and duration settings.
  • Integration & Timing: export per-scene MP3s and a manifest.json with durations to sync audio to Remotion or other composition systems, plus examples for error handling and timing-aware playback.

Quick Start

Generate per-scene voiceover MP3s and a timing manifest from VOICEOVER-SCRIPT.md using your ElevenLabs API key so the files can be consumed by your Remotion composition.

Frequently Asked Questions about elevenlabs

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate per-scene voiceover MP3s with a timing manifest for video narration?

To generate per-scene voiceover MP3s with a timing manifest, provide your script and ElevenLabs API key. The Skill outputs individual MP3 audio files alongside a manifest.json containing durations for direct integration into your composition renderer.

Can I clone a voice from an audio sample for consistent podcast narration?

Yes, you can clone a voice from an audio sample for consistent podcast narration. The Skill creates instant or professional voice clones from provided samples, enabling consistent voice application across multiple audio generation projects.

Does this Skill support SSML for flash and turbo text-to-speech models?

Yes, this Skill supports SSML for flash and turbo text-to-speech models. You can select specific models and adjust voice settings including stability, similarity, style, and speed during audio generation.

What's the best way to generate short sound effects and longer music tracks from text prompts?

The best way to generate short sound effects and longer music tracks from text prompts is using this Skill's prompt influence and instrument control features. You can specify duration settings to produce tailored SFX or musical tracks for game audio.

Do I need an ElevenLabs API key to produce AI voiceovers and soundtracks?

Yes, you need an ElevenLabs API key to produce AI voiceovers and soundtracks. The API key is required to authenticate requests for text-to-speech, voice cloning, sound effects, and music generation processing.

Can I export streaming text-to-speech outputs for real-time audio generation applications?

Yes, you can export streaming text-to-speech outputs for real-time audio generation applications. The Skill supports streaming outputs and saves files, allowing flexible integration with various playback systems and composition renderers.