audiocraft-audio-generation

Generate music and sound effects from text descriptions using MusicGen and AudioGen models.

1|Updated Apr 29, 2026
One-click install
npx skills add https://github.com/bailynlove/STARK-TOWER --skill audiocraft-audio-generation-bailynlove
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/bailynlove/STARK-TOWER/tree/main/opencrew/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/bailynlove/STARK-TOWER --skill audiocraft-audio-generation-bailynlove

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires audiocraft, torch, transformers, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill solves the problem of generating music and sound effects from text descriptions, eliminating the need for manual composition and providing a quick and efficient way to create audio content.

Core Features & Use Cases

  • Text-to-Music: Convert textual descriptions into music using the MusicGen model.
  • Text-to-Sound: Create sound effects from text descriptions using the AudioGen model.
  • EnCodec: Compress and decompress audio files for efficient storage and transfer.
  • Use Case: If you need to create a background track for a video or a sound effect for a game, you can use this Skill to generate the desired audio directly from a text description.

Quick Start

Use the audiocraft-audio-generation skill to generate a happy upbeat electronic dance music track with synths from the text description 'upbeat electronic dance music with synths'.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music and sound effects from text descriptions?

To generate music and sound effects from text descriptions, you can use this Skill to process your textual prompts through the MusicGen and AudioGen models, outputting audio content for creative applications.

Can I create background tracks for videos and games using text-to-sound generation?

Yes, text-to-sound generation supports creating background tracks for videos and sound effects for games by converting your descriptive text prompts directly into the desired audio content.

Do I need torch and transformers to perform text-to-music generation?

Yes, you need the torch and transformers dependencies installed in your environment to run the audiocraft models required for text-to-music and text-to-sound generation tasks.

What is the best way to convert text descriptions into upbeat electronic dance music?

The best way to convert text descriptions into upbeat electronic dance music is using the MusicGen model, which processes prompts like 'upbeat electronic dance music with synths' to generate the audio track.

Does audiocraft support compressing and decompressing audio files for storage?

Yes, audiocraft includes the EnCodec component to compress and decompress audio files, providing efficient storage and transfer of the generated music and sound effects.

What are the limitations of using AI for text-based audio content generation?

Text-based audio content generation eliminates manual composition but is primarily suited for creative and entertainment applications, meaning it may lack the precise control required for professional audio engineering.