audiocraft-audio-generation

Generate audio from text prompts using AudioCraft's MusicGen, AudioGen, and EnCodec.

2|2|Updated Apr 16, 2026
One-click install
npx skills add https://github.com/huidge/hermes-skills --skill audiocraft-audio-generation-huidge
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/huidge/hermes-skills/tree/main/mlops/models/audiocraft
Command: npx skills add https://github.com/huidge/hermes-skills --skill audiocraft-audio-generation-huidge

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft enables rapid creation of audio content by turning text into music and sound effects using MusicGen, AudioGen, and EnCodec.

Core Features & Use Cases

  • Text-to-music generation for composing tracks from descriptions.
  • Text-to-sound effects generation for Foley and ambience.
  • Melody conditioning and style transfer across multiple model variants.
  • Quick experimentation and prototyping for audio workflows across genres and formats.

Quick Start

Describe your desired audio prompt and let AudioCraft generate the track.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text prompts using MusicGen?

Text-to-sound generation creates sound effects and ambience from text descriptions using AudioGen. It applies to sound design workflows like Foley generation, allowing rapid prototyping of environmental audio from descriptive prompts.

Can I use melody conditioning for style transfer across different audio tracks?

Melody conditioning enables style transfer by guiding audio generation with an existing melody across multiple model variants. This allows you to condition MusicGen outputs to match specific melodic structures while altering the surrounding musical style.

How do I set up the environment for reproducible audio generation results?

Environment setup for reproducible audio generation involves specifying dependencies and configuring generation parameters. The Skill defines recommended workflows including model variant selection and specific configurations to ensure consistent, reproducible audio outputs.

Does AudioCraft support both music composition and sound effect generation?

AudioCraft supports both music composition and sound effect generation through its MusicGen and AudioGen models. It leverages EnCodec to process audio, enabling text-to-music creation alongside text-to-sound effects for comprehensive audio prototyping.