What problem does it solve? Producing original music, fixing a bad section of a track, or extending a short hook into a full song normally requires studio time or expensive commercial APIs. This Skill gives you tag-driven music generation, time-range inpainting, and bidirectional outpainting through StepFun-AI's open-weights ACE Step model at $0.0002β0.0003 per second of audio. ## Core Features & Use Cases - Text-to-audio generation: Compose 5β240 second stereo tracks from comma-separated genre, mood, and instrument tags, with optional structured lyrics using [Verse]/[Chorus]/[Bridge] markers. The ACE Step 1.5 endpoint supports vocals in 50+ languages. - Audio inpainting: Regenerate a specific time range inside an existing track (e.g., replace a weak chorus from 20β40 s) without re-rendering the whole song. - Audio outpainting: Extend a track bidirectionally by adding an intro before and an outro after, turning a 30 s hook into a 2 minute cut. - Use Case: A game developer needs background loop beds for a level. They run the base text-to-audio endpoint with tags like "seamless loop, consistent groove, 90 BPM" at 60β120 s per track, iterating cheaply before committing to final renders. ## Quick Start Ask the AI to generate a 90-second lo-fi hip-hop instrumental track with vinyl crackle and rhodes piano using ACE Step on RunComfy.