audiocraft-audio-generation

Generate music and sound effects from text prompts with MusicGen and AudioGen.

Updated May 14, 2026
One-click install
npx skills add https://github.com/SethyPagna/Secretary-Jarvis --skill audiocraft-audio-generation-sethypagna
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audiocraft-audio-generation
Source: https://github.com/SethyPagna/Secretary-Jarvis/tree/main/src/capabilities/skills/mlops/models/audiocraft
Command: npx skills add https://github.com/SethyPagna/Secretary-Jarvis --skill audiocraft-audio-generation-sethypagna

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

AudioCraft removes the friction of creating custom music, sound effects, and audio prototypes by turning natural-language ideas into usable audio outputs with controllable quality, style, and duration.

Core Features & Use Cases

  • Text-to-Music: Generate original music with MusicGen for background tracks, demos, game audio, and concepting.
  • Text-to-Sound: Create sound effects and environmental audio with AudioGen for product mockups, media production, and rapid iteration.
  • Advanced Control: Use melody conditioning, stereo output, EnCodec compression, fine-tuning, batch generation, and deployment patterns for production workflows.
  • Use Case: A developer can describe “upbeat cinematic orchestral music with strings and brass” and produce a ready-to-review audio clip without manual composition.

Quick Start

Use the audiocraft-audio-generation skill to create a short, upbeat electronic music clip from a text prompt with MusicGen.

Frequently Asked Questions about audiocraft-audio-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text prompts using MusicGen?

Generate music from text prompts using MusicGen by providing a natural-language description like “upbeat cinematic orchestral music with strings and brass” to produce a ready-to-review audio clip without manual composition.

What is the best way to create sound effects from text for game audio?

Create sound effects from text for game audio and media production by using AudioGen to turn natural-language ideas into environmental audio and custom sound effect prototypes with controllable duration.

Can I use melody conditioning and stereo output for text-to-music generation?

Yes, melody conditioning and stereo output are supported for text-to-music generation, alongside advanced controls like EnCodec compression, batch generation, temperature, and guidance scale handling.

Do I need PyTorch and torchaudio to run AudioGen and MusicGen?

Yes, you need PyTorch-compatible generation settings and torchaudio to run MusicGen and AudioGen, as the audio generation workflows require these frameworks to manage sample rates and audio processing.

How do I fine-tune MusicGen models for custom audio generation workflows?

Fine-tune MusicGen models for custom audio generation workflows by applying the deployment patterns and generation settings included in the Skill to adapt audio output style, quality, and duration to specific production needs.

What are the limitations of using EnCodec for audio compression?

EnCodec provides audio compression for production workflows within the AudioCraft ecosystem, but its application is limited to processing outputs generated through the MusicGen and AudioGen text-to-audio pipelines.