audiocraft-audio-generation

Generate music and sound effects from text using Meta's AudioCraft models.

1|Updated Jul 31, 2026
One-click install
npx skills add https://github.com/icyzh/hermes-web --skill audiocraft-audio-generation-icyzh
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/icyzh/hermes-web/tree/main/optional-skills/creative/audiocraft-audio-generation
Command: npx skills add https://github.com/icyzh/hermes-web --skill audiocraft-audio-generation-icyzh

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires audiocraft, torch, transformers, torchaudio, and includes references (resource) components.

What problem does it solve?

This Skill solves the challenge of creating custom, high-quality audio assets like background music or sound effects without requiring professional studio equipment or manual recording.

Core Features & Use Cases

  • MusicGen: Generate text-to-music with melody conditioning and style transfer.
  • AudioGen: Create realistic environmental sound effects and audio textures.
  • EnCodec: Perform high-fidelity neural audio compression and reconstruction.
  • Use Case: Quickly generate a 30-second upbeat electronic track for a video project or create specific sound effects like city traffic or nature sounds for game development.

Quick Start

Use the audiocraft-audio-generation skill to generate a 10-second upbeat electronic dance music track with synths.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate high-fidelity music and sound effects from text descriptions?

You can generate high-fidelity music and sound effects from text descriptions using Meta's AudioCraft models. The Skill supports text-to-music generation with MusicGen and realistic environmental sound creation with AudioGen.

Can I create realistic environmental sound effects like city traffic or nature sounds?

Yes, you can create realistic environmental sound effects using AudioGen. It generates audio textures and specific sounds like city traffic or nature sounds for game development and production workflows directly from text prompts.

Do I need PyTorch and transformers to run AudioCraft model inference?

Yes, you need PyTorch, torchaudio, and transformers to execute AudioCraft model inference and audio synthesis. These dependencies are required to run the text-to-music and text-to-sound generation processes locally.

What is the best way to generate a short upbeat electronic track for a video project?

The best way to generate a short upbeat electronic track is using MusicGen for text-to-music synthesis. You can quickly produce a 30-second track with synths by providing a descriptive text prompt for your video project.

Does AudioCraft support neural audio compression and reconstruction?

Yes, AudioCraft supports neural audio compression and reconstruction through EnCodec. It performs high-fidelity neural audio compression to efficiently encode and decode audio assets for creative workflows.

Can I apply melody conditioning and style transfer when generating AI music?

Yes, you can apply melody conditioning and style transfer when generating AI music. MusicGen allows you to condition text-to-music generation on existing melodies to transfer specific musical styles.