audiocraft-audio-generation

Generate music and sound effects from text prompts using AudioCraft models.

Updated May 12, 2026
One-click install
npx skills add https://github.com/hungthinh04/Hermes_AI_Agent --skill audiocraft-audio-generation-hungthinh04
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/hungthinh04/Hermes_AI_Agent/tree/main/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/hungthinh04/Hermes_AI_Agent --skill audiocraft-audio-generation-hungthinh04

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft removes the manual work of turning prompts into music, sound effects, and conditioned audio variations for prototypes, demos, and creative production.

Core Features & Use Cases

  • Music generation: Create songs, background scores, and genre-specific clips with MusicGen.
  • Sound design: Generate environmental audio and effects with AudioGen.
  • Advanced workflows: Use melody conditioning, stereo output, style transfer, and audio continuation when you need more control.
  • Use case: Build a music-generation app, produce sound effects for a game, or prototype an audio demo from a simple text description.

Quick Start

Ask the skill to generate a 10-second upbeat electronic music clip from my prompt and save the result as a WAV file.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from a text prompt?

To generate music from a text prompt, you describe the genre and mood, and the skill uses MusicGen to create songs, background scores, and genre-specific audio clips. It removes manual composition work by directly translating text descriptions into audio.

Can I create sound effects from text for a game prototype?

Yes, you can create sound effects from text for game prototypes using AudioGen. It generates environmental audio and specific effects directly from text descriptions, providing quick assets for creative production.

Does this skill support melody conditioning and stereo output?

Yes, the skill supports melody conditioning, stereo output, style transfer, and audio continuation. These advanced workflows provide greater control over text-to-music generation and conditioned audio variations.

What do I need to run AudioCraft models for text-to-audio generation?

You need AudioCraft models, PyTorch, torchaudio, and Transformers-compatible generation parameters to run text-to-audio generation. These dependencies handle sampling, conditioning, and waveform export.

How do I save generated audio clips as a WAV file?

You save generated audio clips as a WAV file by requesting the skill to generate audio from a prompt and export the result. It processes the text-to-music or text-to-sound workflow and outputs the final waveform.

Is there a way to transfer musical styles using text descriptions?

Yes, style transfer is an advanced workflow supported by the skill. You can use text descriptions to guide the generation process and transfer specific musical styles into your output audio clips.