What problem does it solve?
This Skill centralizes multimodal media generation so users can produce images, videos, and audio from text or existing media without managing individual model APIs or orchestration details.
Core Features & Use Cases
- Multimodal Generation: Text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo), text-to-speech (CSM-1B), and video-to-audio (ThinkSound).
- Model Discovery & Job Management: Search and find models, run generate jobs, check async results, cancel jobs, and estimate costs via the fal.ai MCP.
- Use Case: Quickly iterate on visual drafts with Nano Banana 2, produce production-grade images with Nano Banana Pro, generate short cinematic clips with Seedance or Veo, and create narration or sound design with CSM-1B and ThinkSound.
Quick Start
Generate a photorealistic product image using the Nano Banana Pro model from the prompt "professional product photo of wireless headphones on marble surface".