What problem does it solve? Creating a convincing viral "Kiss Cam moment" video requires coordinating image generation, reference-likeness preservation, and video animation with native audio — a multi-step pipeline that is easy to get wrong (identity drift, over-acting subjects, lip-synced announcer dialogue). This Skill encodes a validated two-call pipeline that produces a spectator-POV Jumbotron still and a 15-second in-arena clip from just two subject photos. ## Core Features & Use Cases - Two-stage pika pipeline: Generates a spectator-POV Madison Square Garden Jumbotron still with gpt-image-2, then animates it into a 15-second 1080p clip with Kling v3-omni using first-frame locking. - Style-agnostic subject anchoring: Both subjects are anchored purely through reference images, preserving likeness and visual style (photoreal, 3D toy, illustrated avatar) without textual descriptions. - Native audio with off-screen PA announcer: Produces crowd reactions and PA-announcer commentary while keeping on-screen subjects silent, using calibrated prompt and negative-prompt anchors. - Use Case: A user provides two photos — one of themselves and one of a friend's illustrated avatar — and receives a fan-filmed-style Kiss Cam video of the two sharing a kiss on the MSG Jumbotron, ready to post to TikTok or Instagram. ## Quick Start Make me a kiss cam moment using photo-a.jpg as subject A and photo-b.jpg as subject B.