audiocraft-audio-generation

Generate music and sound effects from text prompts with AudioCraft models.

1|Updated May 18, 2026
One-click install
npx skills add https://github.com/rickyananda/hermes-skills --skill audiocraft-audio-generation-rickyananda
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/rickyananda/hermes-skills/tree/main/mlops/models/audiocraft
Command: npx skills add https://github.com/rickyananda/hermes-skills --skill audiocraft-audio-generation-rickyananda

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill removes the friction of turning written ideas into usable audio by helping you generate music, sound effects, and audio continuations with Meta AudioCraft models.

Core Features & Use Cases

  • Text-to-Music: Create original tracks with MusicGen for background music, demos, prototypes, and creative production.
  • Text-to-Sound: Generate environmental audio and sound effects with AudioGen for games, media, and product experiences.
  • Advanced Audio Workflows: Support melody conditioning, stereo output, style transfer, audio compression, deployment, and troubleshooting for more production-ready use cases.
  • Use Case: A developer building a game prototype can use this Skill to generate ambient music, footsteps, rain, and other scene-specific sounds from short descriptions.

Quick Start

Ask the skill to generate a music or sound effect prompt for your target scene, then use the recommended AudioCraft model and settings to produce and save the audio.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music and sound effects from text prompts?

You generate music and sound effects from text prompts by applying AudioCraft models like MusicGen and AudioGen to your descriptions. This allows you to create original tracks, environmental audio, and scene-specific sounds for games or media production.

Can I use an existing melody to condition text-to-music generation?

Yes, you can use an existing melody to condition text-to-music generation. AudioCraft supports melody-conditioned generation, allowing you to influence the musical output by providing a reference melody alongside your text prompt.

What's the best way to create stereo output and perform style transfer with AudioCraft?

To create stereo output and perform style transfer with AudioCraft, you apply the advanced audio workflows supported by the models. This involves configuring generation parameters for stereo channels and using style transfer techniques for production-ready audio results.

Does AudioCraft support audio continuation and EnCodec compression workflows?

Yes, AudioCraft supports audio continuation and EnCodec compression workflows. You can extend existing audio segments and apply EnCodec compression to manage audio file sizes while maintaining compatibility with generation parameters and sample-rate handling.

How do I handle sample-rate and generation parameters for AudioGen sound effects?

You handle sample-rate and generation parameters for AudioGen by configuring the model settings according to your target scene requirements. Proper sample-rate handling ensures the generated environmental audio and sound effects meet production quality standards.

Why do I need reference material for AudioCraft fine-tuning and troubleshooting?

You need reference material for AudioCraft fine-tuning and troubleshooting to address advanced production issues. Reference material provides the necessary baseline for adjusting model behavior, resolving generation errors, and optimizing output quality.