fal-ai-media

Generate images, videos, and audio from text via fal.ai MCP models.

1|Updated Mar 18, 2026
One-click install
npx skills add https://github.com/xxih/ai-harness-zh --skill fal-ai-media-xxih
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fal-ai-media
Source: https://github.com/xxih/ai-harness-zh/tree/main/references/translations/everything-claude-code/docs/zh-CN/skills/fal-ai-media
Command: npx skills add https://github.com/xxih/ai-harness-zh --skill fal-ai-media-xxih

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It eliminates the need to manually search for and operate multiple AI media models, providing a single unified interface to generate images, videos, and audio from text or existing media.

Core Features & Use Cases

  • Unified Generation: Supports text‑to‑image, text/image‑to‑video, and text‑to‑speech through fal.ai’s MCP models such as Nano Banana, Seedance, and CSM‑1B.
  • Cost Estimation: Built‑in tool to estimate generation costs before running expensive models.
  • Media Editing: Allows image uploads for in‑painting or using images as prompts for video creation.
  • Use Case Example: A marketer can instantly create a promotional video from a script and a few reference images, or a developer can generate audio narration for documentation.

Quick Start

Request the fal.ai media skill to generate a photorealistic product image of a wireless headset on marble.

Frequently Asked Questions about fal-ai-media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images, video, and audio from text prompts using fal.ai?

To generate media from text prompts using fal.ai, you use a unified interface supporting text-to-image, text-to-video, and text-to-speech. This requires a configured fal.ai MCP server with an API key to process your textual prompts into media content.

Can I estimate generation costs before running fal.ai media models?

Yes, you can estimate generation costs before running fal.ai media models. The skill includes a built-in cost estimation tool that evaluates optional parameters, allowing you to preview expenses before executing expensive generation tasks.

How do I use an existing image for video creation or in-painting with fal.ai?

You can use an existing image for video creation or in-painting with fal.ai by uploading it as a prompt. The media editing feature accepts image uploads to guide text-to-video generation or perform modifications on the original media.

Do I need a fal.ai MCP server to generate audio narration and marketing materials?

Yes, you need a configured fal.ai MCP server with an API key to generate audio narration and marketing materials. This environment setup is required to access models like CSM-1B for audio and Seedance for video generation.

What is the best way to create a promotional video from a script and reference images?

The best way to create a promotional video from a script and reference images is using a unified generation interface. By applying fal.ai media models, you can combine text scripts and uploaded reference images to instantly generate visual marketing content.

Which fal.ai models support text-to-video and text-to-speech generation?

Models supporting text-to-video and text-to-speech generation include Seedance and CSM-1B. These fal.ai MCP models process textual inputs to produce video and audio content, while Nano Banana handles text-to-image generation tasks.