ace-step

Generate, inpaint, and outpaint music with ACE Step via RunComfy CLI.

5|2|Updated May 18, 2026
One-click install
npx skills add https://github.com/doany-ai/skills --skill ace-step-doany-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ace-step
Source: https://github.com/doany-ai/skills/tree/main/ace-step
Command: npx skills add https://github.com/doany-ai/skills --skill ace-step-doany-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

ACE Step enables rapid music creation, inpainting, and outpainting using open-weights models via RunComfy, dramatically reducing time and cost to prototype audio content.

Core Features & Use Cases

  • Tag-driven music generation with lyrics and structure markers.
  • Time-bound inpainting and bidirectional outpainting for edits and expansion.
  • Multilingual lyrics support (1.5 variant) and low-cost draft iterations for prototyping.

Quick Start

Install the RunComfy CLI and generate music with ACE Step using the appropriate text-to-audio or inpaint/outpaint route.

Frequently Asked Questions about ace-step

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate music from text using open-weight models?

Generate music from text using open-weight models by providing tag-based inputs for genre, mood, and instruments. ACE Step processes these tags via the RunComfy CLI to create base audio tracks reproducibly using seed fields.

What is music inpainting and outpainting for audio editing?

Music inpainting edits specific time-bound segments within an existing track, while outpainting expands audio bidirectionally. ACE Step supports both workflows to enable targeted edits and seamless expansion of generated audio content.

Do I need the RunComfy CLI to run ACE Step music generation?

Yes, the RunComfy CLI is required to run ACE Step music generation. You must install the CLI environment to access the text-to-audio, inpaint, and outpaint routes that drive the open-weights models.

Can I use multilingual lyrics for text-to-audio generation?

Yes, you can use multilingual lyrics for text-to-audio generation. The ACE Step 1.5 variant supports multilingual lyrics input alongside tag-driven genre and mood controls to create diverse audio content.

How do I ensure reproducible results across multiple music generations?

Ensure reproducible music generation results by setting a fixed seed value in the input fields. ACE Step uses this seed alongside your specified tags, lyrics, and duration parameters to consistently generate the same audio output.