ace-step

Generate, repair, and extend music via RunComfy CLI with structured JSON inputs.

31|9|Updated Apr 30, 2026
One-click install
npx skills add https://github.com/prime-skills/runcomfy-agent-skills --skill ace-step-prime-skills
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ace-step
Source: https://github.com/prime-skills/runcomfy-agent-skills/tree/main/ace-step
Command: npx skills add https://github.com/prime-skills/runcomfy-agent-skills --skill ace-step-prime-skills

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill lets you generate, repair, and extend music without the high costs or limited editing workflows of premium music-generation services.

Core Features & Use Cases

  • Text-to-Audio Generation: Create instrumental or vocal tracks from genre, mood, instrument, tempo, and lyric tags using ACE Step or ACE Step 1.5.
  • Audio Inpainting: Regenerate a specific time range in an existing track to fix a chorus, replace a bridge, or revise another section.
  • Audio Outpainting: Extend tracks before or after the original recording to add intros, outros, fades, or longer arrangements.
  • Use Case: Produce inexpensive background-music drafts, multilingual vocal songs, game loops, or a polished longer cut from a short musical hook through the RunComfy CLI.

Quick Start

Use the ace-step skill to generate a 90-second mellow lo-fi instrumental with Rhodes piano, soft drums, vinyl crackle, and a 75 BPM tempo.

Frequently Asked Questions about ace-step

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate background music from text tags for a specific genre and tempo?

Text-to-audio generation creates instrumental or vocal tracks from genre, mood, instrument, tempo, and lyric tags. You can produce background-music drafts, multilingual vocal songs, or game loops by providing structured JSON inputs to the model.

What is audio inpainting and how does it repair an existing music track?

Audio inpainting regenerates a specific time range in an existing track to fix a chorus, replace a bridge, or revise another section. It requires HTTPS source audio and endpoint selection based on your inpainting intent.

Can I extend a short musical hook by adding intros or outros to the recording?

Audio outpainting extends tracks before or after the original recording to add intros, outros, fades, or longer arrangements. This allows you to create a polished longer cut from a short musical hook.

Do I need the RunComfy CLI and authentication to use the ACE Step models?

Yes, using the ACE Step models requires the RunComfy CLI and authentication. You must provide structured JSON inputs and select the appropriate endpoint based on whether your intent is generation, inpainting, or outpainting.

Does this approach support generating vocal songs with multilingual lyrics?

Yes, text-to-audio generation supports multilingual lyrics for vocal composition. You can specify lyric tags alongside genre, mood, instrument, and tempo tags to create inexpensive vocal tracks.

What are the limitations when editing existing tracks with audio inpainting?

Audio inpainting requires HTTPS source audio for editing workflows and is limited to regenerating specific time ranges. You must select the correct endpoint and provide structured JSON inputs to execute the repair successfully.