What problem does it solve?
AudioCraft eliminates the need for manual composition by turning natural-language descriptions into generated audio, including full text-to-music and text-to-sound effects.
Core Features & Use Cases
- Text-to-music with MusicGen: Generate melodies and complete songs from text prompts, optionally with melody conditioning and style transfer for reference-guided output.
- Text-to-audio effects with AudioGen: Create short sound effects (e.g., ambience, environment noises, or character/scene SFX) from descriptive prompts.
- Neural audio codec with EnCodec: Enable high-fidelity audio tokenization and reconstruction workflows for improved generation quality and controllability.
- Use Case: Describe a scene and generate a matching soundtrack and sound effects (e.g., “sunset cinematic strings with soft percussion” plus “distant thunder and rain ambience”) for a prototype, demo, or creative pipeline.
Quick Start
Use the audiocraft-audio-generation skill to generate an 8-second track from the prompt “happy upbeat electronic dance music with synths” and save the resulting WAV output.