audiocraft-audio-generation

Generate music and sound effects from text prompts with AudioCraft.

Updated Apr 1, 2026
One-click install
npx skills add https://github.com/founderphantom/zola-agent --skill audiocraft-audio-generation-founderphantom
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/founderphantom/zola-agent/tree/main/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/founderphantom/zola-agent --skill audiocraft-audio-generation-founderphantom

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft enables generating music and sound effects from natural language prompts, reducing manual composition time and enabling rapid iteration in multimedia projects.

Core Features & Use Cases

  • MusicGen: text-to-music generation for soundtrack ideas and background tracks.
  • AudioGen: text-to-sound effects generation for UI, games, and podcasts.
  • Melody conditioning & stereo output: supports melody-based guidance and stereo playback, with EnCodec for high-quality audio.
  • End-to-end workflows: integration with PyTorch and Transformers for deployment and experimentation.

Quick Start

Describe the audio you want and let AudioCraft generate music or sound effects.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music and sound effects from text prompts?

Text-to-audio generation converts natural language prompts into music and sound effects using AudioCraft's MusicGen and AudioGen models, enabling rapid creation of soundtrack ideas and game audio without manual composition.

Can I condition text-to-music generation using an existing melody?

Melody conditioning supports melody-based guidance for text-to-music generation, allowing you to steer the musical output by providing an input melody while generating stereo playback tracks with MusicGen.

Do I need PyTorch and Transformers to run AudioCraft for audio generation?

Yes, generating music and sound effects from text requires the AudioCraft library alongside dependencies like PyTorch and Transformers to execute end-to-end audio rendering workflows and EnCodec encoding.

What's the difference between MusicGen and AudioGen for sound design?

MusicGen handles text-to-music generation for background tracks and film scoring, whereas AudioGen produces text-to-sound effects tailored for UI elements, game audio, and podcasts.

Does AudioCraft support stereo output and configurable duration?

Yes, AudioCraft supports configurable duration and stereo output for generated audio, utilizing EnCodec encoding for high-quality sound rendering from your natural language text prompts.

When should I use text-to-audio generation for game audio and film scoring?

Use text-to-audio generation for rapid prototyping in multimedia projects to reduce manual composition time, enabling fast iteration on game audio, film scoring, and soundtrack ideas.