audio-gen

Generate multilingual speech, voice clones, and sound effects via ElevenLabs TTS.

267|63|Updated Feb 15, 2026
One-click install
npx skills add https://github.com/modu-ai/cowork-plugins --skill audio-gen
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audio-gen
Source: https://github.com/modu-ai/cowork-plugins/tree/main/moai-media/skills/audio-gen
Command: npx skills add https://github.com/modu-ai/cowork-plugins --skill audio-gen

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires elevenlabs, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill enables seamless AI-driven audio production by generating realistic voices, dubbing, and sound effects to streamline multimedia content creation.

Core Features & Use Cases

  • Text-to-Speech and Voice Cloning: Produce multilingual narration and replicate voices from short samples for branding or personalization.
  • Multilingual Dubbing: Automatically translate and synchronize audio tracks for videos in multiple languages, saving time on manual dubbing.
  • Sound Effect Generation: Create custom sound effects for movies, games, or podcasts based on descriptive prompts.
  • Use Case: You can generate a professional narration in Korean, clone a voice from a 1-minute sample, and dub a video into English or Japanese with LipSync adjustments.

Quick Start

Input your text prompt requesting an audio narration in your preferred language, select a voice preset, and then generate the audio file for immediate use.

Frequently Asked Questions about audio-gen

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate multilingual speech for video dubbing?

Voice cloning replicates a specific voice from a short one-minute audio sample for branding or personalization. You provide the sample, and the system generates new narration matching the original voice characteristics.

Can I create custom sound effects from text prompts?

You can create custom sound effects by inputting descriptive text prompts. The system generates tailored audio assets for movies, games, or podcasts based directly on your provided descriptions.

Do I need Python libraries to use ElevenLabs text-to-speech?

You need Python libraries for audio processing and voice synthesis to use this ElevenLabs text-to-speech integration. These dependencies handle the underlying audio generation and file output operations.

What is the best way to clone a voice from a short audio sample?

The best way to clone a voice is providing a clear one-minute audio sample to the system. It replicates the voice characteristics to produce new narration for branding or personalization.