audiocraft-audio-generation

Generate music and sound effects from text prompts using MusicGen and AudioGen.

1|Updated Apr 21, 2026
One-click install
npx skills add https://github.com/ChangZhou-xj/zxj_skill --skill audiocraft-audio-generation-changzhou-xj
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/ChangZhou-xj/zxj_skill/tree/main/mlops/models/audiocraft
Command: npx skills add https://github.com/ChangZhou-xj/zxj_skill --skill audiocraft-audio-generation-changzhou-xj

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft automates the creation of audio content from text prompts, reducing manual music and sound design effort.

Core Features & Use Cases

  • MusicGen: text-to-music generation with melody conditioning
  • AudioGen: text-to-sound effects generation
  • EnCodec: high-fidelity neural audio codec
  • Multiple model sizes: small to large
  • Stereo support and style transfer
  • Melody conditioning and reference-based style transfer
  • Use cases include generating music for videos, game audio, and sound design tasks

Quick Start

Install the audiocraft package and load a pretrained MusicGen or AudioGen model to generate audio from a text prompt.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text prompts for video production?

Text-to-music generation uses MusicGen to create audio content directly from text prompts, supporting melody conditioning and style transfer for video production. Multiple model sizes are available to match your project scale.

Can I create sound effects from text for game audio?

AudioGen generates sound effects from text prompts for game audio and interactive experiences. It works alongside MusicGen and EnCodec to deliver high-fidelity neural audio codec outputs.

What is melody conditioning in AI audio generation?

Melody conditioning is a text-to-music generation feature that allows reference-based style transfer. It enables you to guide the audio creation process by providing a melody reference alongside your text prompt.

Does AudioCraft support stereo audio generation?

Stereo support is included in the audio generation capabilities. The system uses EnCodec for high-fidelity neural audio encoding across its small to large model sizes.

How do I get started with text-to-audio generation using AudioCraft?

Install the audiocraft package and load a pretrained MusicGen or AudioGen model to generate audio from a text prompt. The workflow requires a SKILL.md frontmatter with a name and description.

What are the limitations of neural audio codec for sound design?

Neural audio codec generation requires manual effort reduction but depends on model size selection from small to large. Sound design tasks should consider the appropriate model scale for desired fidelity.