audiocraft-audio-generation

Generate music and sound from text using AudioCraft's MusicGen and AudioGen models.

Updated Jun 28, 2026
One-click install
npx skills add https://github.com/jleechanorg/hermes-agent --skill audiocraft-audio-generation-jleechanorg
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/jleechanorg/hermes-agent/tree/main/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/jleechanorg/hermes-agent --skill audiocraft-audio-generation-jleechanorg

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Audio production often involves manual, repetitive tasks to create music and sound effects from text prompts. This Skill provides an end-to-end setup to generate audio using AudioCraft's MusicGen, AudioGen, and EnCodec workflows, enabling rapid prototyping and production-ready results.

Core Features & Use Cases

  • Text-to-music generation with MusicGen for melodies and mood-based prompts.
  • Text-to-sound generation with AudioGen for sound effects and ambient audio.
  • Melodic conditioning and high-fidelity encoding with EnCodec for scalable audio pipelines.
  • Use cases include game audio, film sound design, and rapid audio prototyping from prompts.

Quick Start

Install the AudioCraft package and generate sample audio from a short text description.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music and sound effects from text prompts?

Text-to-audio generation uses AudioCraft models like MusicGen and AudioGen to convert text descriptions into music and sound effects. This automates manual production tasks, enabling rapid prototyping for game audio and film sound design.

What is the difference between MusicGen and AudioGen for audio generation?

MusicGen handles text-to-music generation for melodies and mood-based prompts, while AudioGen focuses on text-to-sound generation for ambient audio and sound effects. Both use EnCodec for high-fidelity encoding within scalable audio pipelines.

Do I need PyTorch 2.x to run AudioCraft models?

Yes, running AudioCraft models requires the audiocraft package, PyTorch 2.x or newer, and compatible hardware. These prerequisites are necessary to load pre-trained models like MusicGen and AudioGen for generation workflows.

Can I use a melody to condition text-to-music generation?

Yes, AudioCraft supports melody conditioning to guide text-to-music generation. This allows you to input an existing melodic structure alongside text prompts to generate music that aligns with your desired tune.

What are the best use cases for AudioGen and EnCodec workflows?

AudioGen and EnCodec workflows suit game audio, film sound design, and rapid audio prototyping from text prompts. They provide an end-to-end setup for generating production-ready sound effects and scalable audio pipelines.