What problem does it solve?
AudioCraft enables users to turn natural-language prompts into synthesized audio (music and sound effects), avoiding the need for manual composition or expensive studio production for early drafts and prototypes.
Core Features & Use Cases
- Text-to-Music (MusicGen): Generate complete music tracks from textual descriptions, including melody-conditioned and stereo variants.
- Text-to-Sound Effects (AudioGen): Produce short, prompt-driven sound effects such as environments, actions, and ambience.
- Audio Codec Power (EnCodec): Compress and reconstruct audio representations to support higher-fidelity audio workflows, including advanced generation and processing pipelines.
- Common Use Case: Prototype a soundtrack concept by generating several 10–30 second variations from prompts like “epic orchestral soundtrack with strings and brass,” then audition and iterate on style before committing to final production.
- When to Use: You need fast iteration, flexible generation length, and practical controls such as duration, sampling diversity, and text adherence.
Quick Start
Request: “Generate 20 seconds of upbeat electronic dance music from this description: happy upbeat electronic dance music with synths, and save it as output.wav using AudioCraft MusicGen on my machine.”