audiocraft-audio-generation

Generate custom music, sound effects, and compressed audio from text descriptions using Meta's AudioCraft.

1|Updated Jun 25, 2026
One-click install
npx skills add https://github.com/Signmanal/VIGIL --skill audiocraft-audio-generation-signmanal
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/Signmanal/VIGIL/tree/main/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/Signmanal/VIGIL --skill audiocraft-audio-generation-signmanal

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill eliminates the need for manual audio production or expensive stock audio subscriptions by enabling on-demand generation of custom music, sound effects, and compressed audio from text descriptions.

Core Features & Use Cases

  • Text-to-Audio Generation: Create custom music with MusicGen (including melody-conditioned and stereo output) and sound effects with AudioGen from simple text prompts.
  • Audio Compression: Use EnCodec for high-fidelity neural audio compression to reduce file size without significant quality loss.
  • Use Case: A game developer can use this Skill to generate unique background music and environmental sound effects for a game level directly from text descriptions, avoiding the time and cost of sourcing or recording custom audio.

Quick Start

Use the audiocraft-audio-generation skill to generate a 10-second upbeat electronic dance music track from the text prompt "upbeat electronic dance music with synthesizer leads and punchy drums at 128 bpm" and save it as output.wav.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate custom music and sound effects from text descriptions?

To generate custom audio from text, you can use Meta's AudioCraft library to create music with MusicGen and sound effects with AudioGen directly from simple text prompts. This eliminates manual audio production by enabling on-demand creation of game audio assets and content production sound design.

Can I generate stereo music with melody conditioning using AudioCraft?

Yes, AudioCraft supports generating custom music with MusicGen, including specific features for melody-conditioned and stereo output. You can create unique background music tailored to your exact specifications, such as an upbeat electronic dance music track with synthesizer leads and punchy drums.

What's the best way to compress audio files without significant quality loss?

The best way to compress audio without losing significant quality is using EnCodec for high-fidelity neural audio compression. This AudioCast feature reduces file size efficiently, making it ideal for audio storage or transmission while maintaining high fidelity.

Do I need a CUDA-enabled GPU and PyTorch 2.0 to run AudioCraft models?

Yes, running AudioCraft models requires PyTorch 2.0+, the audiocraft Python package, and compatible CUDA-enabled hardware. This specific hardware and software setup is necessary to achieve accelerated inference for MusicGen, AudioGen, and EnCodec.

Does AudioCraft work for generating game audio assets and environmental sound effects?

Yes, AudioCraft works perfectly for game audio asset creation by generating unique background music and environmental sound effects. A game developer can create custom audio for a game level directly from text descriptions, avoiding the time and cost of sourcing or recording custom audio.

What are the limitations of using text-to-music generation for content production?

Limitations of text-to-music generation include the requirement for specific hardware like CUDA-enabled GPUs and PyTorch 2.0+ for accelerated inference. While it enables music prototyping and eliminates stock audio subscriptions, users must manage these technical prerequisites for effective content production sound design.