stable-audio

Generate and validate Stable Audio music and sound assets with rights tracking.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill stable-audio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: stable-audio
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill stable-audio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps media teams create production-ready music, sound effects, ambience, foley, loops, and sonic-branding assets with Stable Audio while avoiding unsuitable speech workflows, rights risks, and unreliable delivery decisions.

Core Features & Use Cases

  • Model and API Routing: Choose between hosted Stable Audio 2, 2.5, and 3 endpoints or open-weight models based on duration, latency, quality, and local-data requirements.
  • Audio Generation and Editing: Plan text-to-audio, audio-to-audio, inpainting, and continuation workflows with appropriate prompts, durations, steps, seeds, formats, and source-audio strength.
  • Production Guardrails: Enforce rights-cleared uploads, document provenance, handle asynchronous Stable Audio 3 polling, and review audio for artifacts, editability, loudness, continuity, and delivery readiness.
  • Use Case: Create a 40-second optimistic instrumental product-launch bed under narration, or generate multiple variants of a rights-cleared game chime while preserving its core motif.

Quick Start

Use the stable-audio skill to create a rights-safe 30-second SaaS launch music bed with a narration-friendly arrangement, generation parameters, provenance notes, and a final audio QA checklist.

Frequently Asked Questions about stable-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate rights-safe music and sound effects for media production?

Generate rights-safe music and sound effects by using text-to-audio workflows with documented provenance, rights-cleared source uploads, and final audio quality checks to ensure production readiness. This prevents unsuitable speech workflows and rights risks.

How does audio inpainting work for editing existing sound assets?

Audio inpainting works by applying audio-to-audio transformations with appropriate source-audio strength and prompts to regenerate or fill missing sections of existing loops, foley, or ambience while preserving the original motif and continuity.

Can I use Stable Audio 3 endpoints for generating long duration sound beds?

Stable Audio 3 endpoints support long duration sound beds through asynchronous job polling, allowing you to generate extended instrumental beds, ambience, and sonic branding assets while managing latency and parameter limits.

What is the best way to choose between hosted and local open-weight audio generation models?

Choose between hosted and local open-weight audio generation models by evaluating duration, latency, quality, and local-data requirements to route your text-to-audio, inpainting, and continuation workflows to the appropriate endpoint.

What are the limitations when uploading source audio for audio-to-audio transformations?

Limitations for uploading source audio include strict rights-cleared content requirements and provenance documentation, preventing rights-uncleared uploads and unsuitable speech workflows from entering the audio transformation pipeline.

Why do my generated music beds have artifacts and fail audio quality checks?

Generated music beds have artifacts and fail audio quality checks due to improper generation steps, seeds, or parameter limits, requiring a final review for loudness, continuity, editability, and delivery readiness before media project integration.