What problem does it solve?
AudioCraft enables you to turn text prompts into realistic audio—covering both text-to-music (MusicGen) and text-to-sound-effects (AudioGen)—so you don’t have to start from scratch when creating audio for demos, prototypes, or creative projects.
Core Features & Use Cases
- Text-to-music with melody control (MusicGen): Generate short to medium-length musical audio directly from prompts, including melody-conditioned variants for tighter musical structure.
- Text-to-sound-effects (AudioGen): Produce environmental and effect sounds from descriptive prompts for game/audio production workflows.
- Neural audio codec (EnCodec): Support higher-fidelity audio tokenization and reconstruction as part of the model pipeline.
- Stereo generation and style transfer: Use stereo model variants and MusicGen-Style for reference-based style conditioning.
Example use case: You need a 15-second ambient electronic bed for a research demo; generate multiple variations from different prompts, pick the best one, and refine by adjusting guidance and sampling settings.
Quick Start
Use the audiocraft-audio-generation skill to generate a short track from the prompt "happy upbeat electronic dance music with synths" and save it as a WAV file.