What problem does it solve?
AudioCraft automates the creation of audio (music and sound effects) from natural-language prompts, so you can quickly prototype or produce ideas without manual composition or sound design.
Core Features & Use Cases
- Text-to-music generation (MusicGen): Create short or longer music tracks from text, including melody-conditioned variants and stereo models.
- Text-to-sound effects generation (AudioGen): Generate realistic sound effects such as nature ambience, UI sounds, and environmental noise from descriptions.
- Neural audio codec & enhancement (EnCodec, MBD): Decode high-fidelity neural audio tokens back into waveforms and optionally improve perceived quality with MultiBand Diffusion.
- Use cases: Music brainstorming for games and videos, rapid SFX prototyping, dataset preparation and fine-tuning workflows, and building an audio generation API or UI demo.
Quick Start
Use the audiocraft-audio-generation skill to generate a 10-second text-to-music track for the prompt "upbeat electronic dance music with punchy drums" and save the resulting audio as a WAV file.