audiocraft-audio-generation

Generate music and sound effects from text prompts using AudioCraft models.

31|3|Updated May 7, 2026
One-click install
npx skills add https://github.com/markwang2658/hermes-windows-native --skill audiocraft-audio-generation-markwang2658
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/markwang2658/hermes-windows-native/tree/main/hermes-agent/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/markwang2658/hermes-windows-native --skill audiocraft-audio-generation-markwang2658

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Audio production often requires multiple tools and manual tuning; this Skill enables end-to-end audio generation from natural language prompts using AudioCraft's MusicGen, AudioGen, and EnCodec, consolidating workflows for music and sound design.

Core Features & Use Cases

  • Generate text-to-music with MusicGen, including melody-conditioned variants.
  • Generate text-to-sound effects with AudioGen and EnCodec-based workflows.
  • Rapid prototyping for game audio, film SFX, and content creation, with reference to advanced usage.

Quick Start

Use AudioCraft to generate audio content from a descriptive prompt.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text prompts using MusicGen?

Generate music from text prompts using MusicGen by providing a descriptive natural language prompt to the AudioCraft model, which handles end-to-end text-to-music synthesis for rapid audio prototyping.

Can I create sound effects from text for game audio and film SFX?

Create sound effects from text for game audio and film SFX using AudioGen and EnCodec-based workflows, which consolidate audio production into a single natural language prompt-driven process.

What is melody conditioning in text-to-music generation?

Melody conditioning in text-to-music generation is a feature of MusicGen that allows you to guide the audio synthesis process by providing an existing melody alongside your text prompt.

Does AudioCraft support text-to-sound and text-to-music in one workflow?

AudioCraft supports text-to-sound and text-to-music in one workflow by utilizing MusicGen and AudioGen models, eliminating the need for multiple manual tuning tools during audio production.

Why does my AudioCraft audio generation output need troubleshooting references?

Audio generation output needs troubleshooting references when rapid prototyping encounters issues, and this skill provides basic troubleshooting references to resolve common MusicGen and AudioGen generation parameter errors.