fal-ai-media

Generate images, videos, and audio via fal.ai MCP.

Updated Jul 27, 2026
One-click install
npx skills add https://github.com/kouiso/designdiff --skill fal-ai-media-kouiso
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fal-ai-media
Source: https://github.com/kouiso/designdiff/tree/main/.claude/skills/fal-ai-media
Command: npx skills add https://github.com/kouiso/designdiff --skill fal-ai-media-kouiso

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a unified interface for generating various media types (images, videos, audio) using advanced AI models from fal.ai, simplifying complex media creation workflows.

Core Features & Use Cases

  • Image Generation: Create images from text prompts using models like Nano Banana 2 and Pro. Supports editing and style transfer.
  • Video Generation: Generate videos from text or image inputs with models like Seedance, Kling, and Veo 3, including options for audio.
  • Audio Generation: Produce speech from text (CSM-1B) or generate audio from video content (ThinkSound).
  • Use Case: A marketing team needs to create a short promotional video with a custom voiceover and background music for a new product launch. This Skill can generate the video, synthesize the voiceover, and create ambient audio.

Quick Start

Use the fal-ai-media skill to generate an image of a futuristic cityscape at sunset in a cyberpunk style.

Frequently Asked Questions about fal-ai-media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images, video, and audio from text using fal.ai?

Media generation from text using fal.ai requires configuring the fal.ai MCP server and an API key to enable text-to-image, text-to-video, and text-to-speech functionalities across various models.

What AI models are available for image and video generation via the fal.ai MCP?

Available AI models for image generation include Nano Banana 2 and Pro, while video generation supports Seedance, Kling, and Veo 3, allowing text or image inputs to create dynamic video content.

Do I need an API key to use fal.ai for text-to-speech and audio generation?

Yes, an API key is required. You must configure the fal.ai MCP server with your API key to synthesize speech from text using CSM-1B or generate ambient audio from video content using ThinkSound.

Can I generate a promotional video with voiceover and background music in one workflow?

You can create a promotional video by generating video from text or images, synthesizing a custom voiceover with text-to-speech, and generating ambient background audio to complete the media generation workflow.

Does fal-ai-media support editing and style transfer for generated images?

Yes, image generation supports editing and style transfer. You can create images from text prompts using models like Nano Banana 2 and Pro, and apply modifications to achieve your desired visual style.

What are the prerequisites for setting up the fal.ai MCP server for media generation?

To use fal.ai media generation, you need a valid fal.ai API key and must configure the fal.ai MCP server in your environment to enable the unified interface for generating images, videos, and audio.