What problem does it solve? It turns plain-language descriptions of sound scenes—dialogue, ambient noise, sound effects, background music—into generated audio without requiring audio editing tools or manual synthesis. ## Core Features & Use Cases - Text-to-Audio (T2A): Generate audio purely from a text description, from a single spoken line to a complex multi-character scene with environment sounds and music. - Audio-to-Audio (A2A): Generate audio using up to 3 reference audio clips to control character voice timbre, with automatic normalization of reference mentions into @音频N markers. - Duration Control: Optionally specify a target duration of 1-120 seconds when the user explicitly requests it. - Use Case: A user provides a café scene script with two characters and reference voice clips; the skill normalizes the references, calls the audio generation tool once with the full scene, and delivers the resulting audio URL. ## Quick Start Ask the assistant to generate an audio clip by describing the scene, for example: generate a sound of a rainy night with distant dog barks and rain hitting a tin roof.