fal-ai-media

Generate images, videos, and audio via fal.ai MCP tools.

Updated Mar 20, 2026
One-click install
npx skills add https://github.com/KanakMalpani/General-Private-Skills --skill fal-ai-media-kanakmalpani
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fal-ai-media
Source: https://github.com/KanakMalpani/General-Private-Skills/tree/main/skills/fal-ai-media
Command: npx skills add https://github.com/KanakMalpani/General-Private-Skills --skill fal-ai-media-kanakmalpani

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Unified media generation across images, videos, and audio using fal.ai MCP, enabling consistent, model-driven outputs without switching tools.

Core Features & Use Cases

  • Image generation from text prompts using Nano Banana models.
  • Video generation from prompts or inputs using Seedance, Kling, and Veo 3.
  • Audio generation including text-to-speech with CSM-1B and ThinkSound workflows, plus sample code for ElevenLabs integration.
  • Video-to-audio extraction and synchronization for multimedia projects.

Quick Start

Generate a 5-second video of a futuristic city at dusk using fal.ai MCP.

Frequently Asked Questions about fal-ai-media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate video and audio from text prompts using fal.ai?

Generate video and audio from text using fal.ai MCP by configuring the server and utilizing provided tools to access models like Seedance, Kling, Veo 3, and CSM-1B for multimedia synthesis.

What is the best way to create unified media assets without switching tools?

The best way to create unified media assets is using fal.ai MCP for model-driven generation across images, video, and audio, enabling consistent outputs without switching tools during rapid multimedia production.

Does fal.ai MCP support text-to-speech synthesis and video-to-audio extraction?

Yes, fal.ai MCP supports text-to-speech synthesis using CSM-1B and ThinkSound workflows, alongside video-to-audio extraction and synchronization for comprehensive multimedia projects.

Can I use ElevenLabs integration for audio generation with fal.ai MCP?

Yes, you can use ElevenLabs integration for audio generation; the Skill includes sample code for ElevenLabs workflows alongside native CSM-1B text-to-speech synthesis capabilities.

How do I start generating an image from text using Nano Banana models?

Start generating images from text by configuring your fal.ai MCP server, then use the supplied MCP tools like search and generate to execute text-to-image generation with Nano Banana models.