What problem does it solve? Creating and refining short video clips usually means regenerating from scratch for every small change. This Skill lets you generate 3-10 second 720p videos with synthesized audio and then edit them conversationally — changing lighting, removing objects, or altering on-screen text — while keeping everything else intact. ## Core Features & Use Cases - Stateful Conversational Editing: Each generation returns an interaction_id; pass it back with a short edit prompt like "Make the phone invisible. Keep everything else the same." to refine the same clip in layers without regenerating. - Reference-Image Binding: Attach local images and bind them to roles in the prompt with <FIRST_FRAME> and <IMAGE_REF_N> tags to control subjects, styles, and products. - Timecoded Beats and Rendered Text: Schedule multi-beat scenes with [0-3s] timecode syntax and render word-by-word on-screen text with synthesized audio. - Use Case: Generate a product demo clip, review it, then iteratively adjust the lighting, swap the background, and update the sign text across three follow-up edits — all on the same underlying video. ## Quick Start Use the gemini_omni_video tool to generate a 6-second clip of a cat on a sunny windowsill, then edit it to make the lighting more dramatic while keeping everything else the same.