audiocraft-audio-generation

Generate music and sounds from text prompts using MusicGen and AudioGen.

Updated Mar 31, 2026
One-click install
npx skills add https://github.com/quiznat/Hermes_Sapho --skill audiocraft-audio-generation-quiznat
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/quiznat/Hermes_Sapho/tree/main/.hermes/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/quiznat/Hermes_Sapho --skill audiocraft-audio-generation-quiznat

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Generate music and sounds from text descriptions, enabling rapid audio content creation for creative projects.

Core Features & Use Cases

  • MusicGen: text-to-music generation including melody conditioning
  • AudioGen: text-to-sound generation for effects and atmospheres
  • EnCodec integration: high-fidelity/compact audio encoding and decoding
  • Model variety: supports multiple MusicGen/AudioGen sizes and stereo outputs
  • Use cases: game audio, film scoring, and quick prototyping from prompts

Quick Start

Install audiocraft and run a simple prompt to generate audio.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text descriptions for game audio?

Generate music from text prompts using MusicGen to create audio assets for game design and film scoring. It processes descriptive prompts to produce stereo outputs across multiple model sizes.

What is the difference between text-to-music and text-to-sound generation?

Text-to-music generation creates musical tracks with melody conditioning, while text-to-sound generation produces effects and atmospheres. Both use EnCodec workflows for high-fidelity audio encoding and decoding.

Do I need PyTorch and Transformers to run AudiCraft for audio generation?

Yes, running AudiCraft requires PyTorch 2.x and Transformers to execute text-to-audio generation. This environment supports multiple MusicGen and AudioGen model sizes for high-fidelity outputs.

Can I condition music generation on an existing melody?

Yes, melody conditioning allows you to guide text-to-music generation using an existing melody. This feature is part of the MusicGen workflow for rapid audio prototyping and film scoring.

What are the limitations of using text-to-sound generation for rapid prototyping?

Text-to-sound generation for rapid prototyping depends on available model sizes and EnCodec processing limits. It is designed for quick asset generation rather than replacing professional sound design workflows.