audiocraft-audio-generation

Generate music and sound effects from text descriptions using AudioCraft.

3|Updated Feb 21, 2026
One-click install
npx skills add https://github.com/ihatesea69/HieuNghi-AI-Skills --skill audiocraft-audio-generation-ihatesea69
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/ihatesea69/HieuNghi-AI-Skills/tree/main/airesearch_skills/18-multimodal/audiocraft
Command: npx skills add https://github.com/ihatesea69/HieuNghi-AI-Skills --skill audiocraft-audio-generation-ihatesea69

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires audiocraft, torch, transformers, scipy, torchaudio, and includes references (resource) components.

What problem does it solve?

This Skill enables the creation of original music and sound effects directly from textual descriptions, eliminating the need for complex audio editing software or extensive musical knowledge.

Core Features & Use Cases

  • Text-to-Music Generation: Create music in various genres and styles based on descriptive prompts (e.g., "epic orchestral soundtrack").
  • Text-to-Sound Effects: Generate realistic sound effects for games, films, or other media (e.g., "dog barking in a park").
  • Melody Conditioning: Generate music that follows a specific melodic input.
  • Use Case: A game developer needs a unique sound effect for a magical spell. They can describe the sound ("sparkling magical chime with a short echo") and generate multiple options instantly.

Quick Start

Use the audiocraft skill to generate a 10-second clip of upbeat electronic dance music.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music and sound effects from text descriptions?

You can generate music and sound effects from text descriptions using the AudioCraft library, which supports text-to-music and text-to-sound creation. Descriptive prompts like "epic orchestral soundtrack" or "dog barking in a park" produce original audio clips instantly.

Can I generate sound effects for games and films using text-to-sound?

Text-to-sound generation creates realistic sound effects for games, films, and other media from descriptive prompts. This process eliminates the need for complex audio editing software or extensive musical knowledge to produce unique audio assets.

Does AudioCraft support melody-conditioned music generation?

AudioCraft supports melody-conditioned music generation, allowing you to generate music that follows a specific melodic input. This feature works alongside standard text-to-music generation to provide more control over the musical output.

Do I need PyTorch and torchaudio to run text-to-music generation?

Running text-to-music generation requires PyTorch, torchaudio, and transformers for advanced features. These dependencies support the underlying processing logic of the AudioCraft library to generate audio from textual descriptions.

What is the best way to create an upbeat electronic dance music clip from a prompt?

The best way to create an upbeat electronic dance music clip is to input a descriptive prompt into the AudioCraft text-to-music generator. This enables you to generate short audio clips, such as a 10-second segment, directly from your text description.

What are the limitations of using AudioGen for text-to-sound generation?

AudioGen is limited to generating sound effects and audio assets based on textual input rather than full musical compositions. It requires PyTorch, torchaudio, and transformers to function, relying entirely on the accuracy of your descriptive prompts to achieve the desired audio output.