audiocraft-audio-generation

Generate music and sound effects from text prompts using AudioCraft.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/choice5346/BiSHE --skill audiocraft-audio-generation-choice5346
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/choice5346/BiSHE/tree/main/.github/skills/audiocraft
Command: npx skills add https://github.com/choice5346/BiSHE --skill audiocraft-audio-generation-choice5346

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires audiocraft, torch, transformers, torchaudio, scipy, and includes references (resource) components.

What problem does it solve?

This Skill enables the creation of original music and sound effects directly from textual descriptions, removing the need for specialized audio production skills or extensive libraries.

Core Features & Use Cases

  • Text-to-Music Generation: Create music in various genres and styles based on descriptive prompts.
  • Text-to-Sound Effects: Generate realistic sound effects for games, videos, or other media.
  • Melody Conditioning: Generate music that follows a specific melodic input.
  • Use Case: A game developer needs a unique sound effect for a magical spell. They can simply describe "a shimmering, ethereal magical sound with a rising pitch" and generate multiple options instantly.

Quick Start

Use the audiocraft skill to generate a 10-second clip of "happy upbeat electronic dance music with synths".

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text prompts using AudioCraft?

This Skill uses the AudioCraft library to generate music from text prompts via its MusicGen functionality. You provide a natural language description of the desired audio, and the tool outputs an original music clip matching that description without requiring production skills.

Can I generate realistic sound effects for games from a text description?

Yes, you can generate sound effects from text using the AudioGen functionality within AudioCraft. By providing a descriptive prompt for your game or video, the tool synthesizes realistic sound effects like a shimmering magical sound without needing manual audio production.

Does this text-to-music Skill support melody conditioning and stereo output?

Yes, this AudioCraft implementation supports both melody conditioning and stereo output. It leverages PyTorch and torchaudio to execute models that generate music following a specific melodic input while producing stereo audio tracks.

Do I need PyTorch and Transformers installed to use AudioCraft for audio generation?

Yes, you need PyTorch, Transformers, torchaudio, and scipy installed to run this AudioCraft Skill. These dependencies handle the model execution and audio manipulation required to synthesize music and sound effects from text prompts.

What is the best way to create original audio without specialized production skills?

Using text-to-music and text-to-sound generation is the best way to create original audio without specialized production skills. This Skill translates natural language text prompts directly into audio content, removing the need for extensive audio libraries or manual production.