audiocraft-audio-generation

Generate audio from text prompts and style references using AudioCraft models.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/matlee0409/cronus --skill audiocraft-audio-generation-matlee0409
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/matlee0409/cronus/tree/main/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/matlee0409/cronus --skill audiocraft-audio-generation-matlee0409

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft simplifies generating high-quality music and sound effects from text prompts, enabling rapid exploration and production of audio content with state-of-the-art multimodal models.

Core Features & Use Cases

  • Text-to-music generation with MusicGen, including melody-conditioned and stereo variants.
  • Text-to-sound generation with AudioGen, plus EnCodec-based high-fidelity compression workflows.
  • Quick deployment in research, prototyping, game audio, and content creation pipelines.

Quick Start

Install audiocraft and run a minimal example to generate audio from a text prompt.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from a text prompt using AudioCraft models?

Text-to-music generation with AudioCraft uses MusicGen models to synthesize audio directly from text prompts. You can generate audio by configuring duration and tempo, with optional style conditioning for research or game audio pipelines.

Can I generate sound effects and high-fidelity audio with AudioGen and EnCodec?

AudioGen handles text-to-sound generation for sound effects, while EnCodec provides high-fidelity audio compression workflows. Both are supported within AudioCraft to enable rapid prototyping and content creation pipelines.

Does AudioCraft support melody-conditioned and stereo music generation?

AudioCraft supports melody-conditioned outputs and stereo variants through its MusicGen workflows. This allows you to apply style references and conditioning when generating music for game audio or research applications.

What is the best way to create style-conditioned audio for game pipelines?

The best way to create style-conditioned audio for game pipelines is using AudioCraft's configurable MusicGen and AudioGen workflows. You can apply optional style conditioning to generate music and sound effects with specific duration and tempo settings.

Do I need style references to generate audio with AudioCraft, or can I use text prompts only?

You do not need style references to generate audio with AudioCraft; text prompts alone are sufficient. Style references are optional and can be applied for melody-conditioned outputs and specific style conditioning across MusicGen and AudioGen workflows.

Are there limitations when using AudioCraft for text-to-sound generation in prototyping?

Limitations of AudioCraft for text-to-sound generation in prototyping include the need to configure duration and tempo manually. It supports AudioGen and EnCodec workflows, but complex style conditioning may require additional references for high-fidelity outputs.