What problem does it solve?
Creating a consistent, professional-looking cooking tutorial video from just one photo of a person is hard: faces drift between frames, kitchens change, and multi-step recipes lose continuity. This Skill solves that by first building a composite production reference sheet (character sheet, kitchen environment, and a 9-panel action board) and then animating the full sequence with dual image references so identity and setting stay locked.
Core Features & Use Cases
- Composite Reference Sheet Generation: Uses gpt-image-v2-edit to produce a 3840x2160 board combining character views, kitchen location reference, and a 9-step action plan.
- Dual-Reference Video Generation: Uses bytedance-seedance-2-0-reference-to-video-fast with the original photo as identity anchor and the reference sheet as narrative guide, including generated audio.
- Configurable Output: Customize dish, kitchen style, outfit, duration (10-15s), aspect ratio (16:9 or 9:16 for Reels), and resolution (480p/720p).
- Use Case: A food content creator uploads their photo and gets a polished 15-second pasta-making tutorial video with consistent face, outfit, and kitchen, ready to post or re-render vertically for Reels.
Quick Start
Ask the agent to turn your photo into a 15-second cooking tutorial video by providing the photo URL and optionally specifying the dish, kitchen style, and outfit.