acestep

Generate music tracks and isolate audio stems via RunPod inference.

1|Updated Apr 11, 2026
One-click install
npx skills add https://github.com/shige1014-dev/backup-OpenMontage --skill acestep-shige1014-dev
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: acestep
Source: https://github.com/shige1014-dev/backup-OpenMontage/tree/main/.agents/skills/acestep
Command: npx skills add https://github.com/shige1014-dev/backup-OpenMontage --skill acestep-shige1014-dev

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Many video projects need custom background music, vocal tracks, covers, or isolated stems but lack an affordable, configurable pipeline to generate and extract these assets; acestep provides fast, configurable music generation and stem separation tailored for video production workflows.

Core Features & Use Cases

  • Text-to-music generation: Create complete instrumental or vocal-backed tracks from descriptive prompts with control over duration, BPM, and key.
  • Covers and style transfer: Perform cover-style transformations from a reference audio file and tune cover strength to balance fidelity versus creativity.
  • Stem extraction: Isolate vocals, drums, bass, guitar, and other stems for remixing, cleanup, or mixing under narration in a video timeline.
  • Use Case: Produce an upbeat product demo soundtrack, extract vocal stems for synch with voiceover, and generate a short CTA jingle for social clips.

Quick Start

Generate a 30 second upbeat background track by running the music generator with a clear style prompt, set the duration to 30 seconds, and save the output to a file.

Frequently Asked Questions about acestep

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate background music with specific BPM and duration for video production?

Background music generation for video production allows you to create instrumental or vocal-backed tracks by setting descriptive prompts alongside configurable parameters for BPM, key, and exact duration to fit your timeline.

Can I extract isolated audio stems like vocals and drums from an existing track?

Yes, audio stem extraction isolates individual components such as vocals, drums, bass, and guitar from a track, enabling you to remix, clean up, or mix specific stems under video narration.

Do I need a RunPod API key to generate AI music and audio stems?

Yes, you must configure a RUNPOD_API_KEY and RUNPOD_ACESTEP_ENDPOINT_ID in your environment to run the RunPod inference required for AI music generation and stem separation.

What is the best way to create a cover version from a reference audio file?

Creating a cover version involves applying cover-style transfer from a reference audio file, where you can tune the cover strength parameter to balance fidelity to the original versus creative deviation.

How does text-to-music generation handle lyrics input for vocal tracks?

Text-to-music generation processes lyrics input alongside descriptive style prompts to produce complete vocal-backed tracks, allowing you to control musical elements like BPM and key simultaneously.

Can I use this for producing short jingles for social media clips?

Yes, you can produce short CTA jingles for social clips by configuring the music generator with a clear style prompt and setting the duration to match your specific video length requirements.