What problem does it solve?
This Skill centralizes audio generation for video and multimedia projects, eliminating the need for manual recording, complex audio assembly, or sourcing royalty-free assets by producing TTS, sound effects, music, and voice clones from descriptions and samples.
Core Features & Use Cases
- Text-to-Speech: produce scene-level voiceovers with selectable models and voice settings for stability, similarity, style, and speed.
- Voice Cloning: create instant or professional voice clones from sample audio for consistent narration across projects.
- Sound Effects & Music: generate short SFX and longer musical tracks with prompt influence, instrument control, and duration settings.
- Integration & Timing: export per-scene MP3s and a manifest.json with durations to sync audio to Remotion or other composition systems, plus examples for error handling and timing-aware playback.
Quick Start
Generate per-scene voiceover MP3s and a timing manifest from VOICEOVER-SCRIPT.md using your ElevenLabs API key so the files can be consumed by your Remotion composition.