fal-audio

Transcribe audio, translate speech, synthesize voices, and clone voices.

51|10|Updated Oct 22, 2025
One-click install
npx skills add https://github.com/JosiahSiegel/claude-plugin-marketplace --skill fal-audio-josiahsiegel
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fal-audio
Source: https://github.com/JosiahSiegel/claude-plugin-marketplace/tree/main/plugins/fal-ai-master/skills/fal-audio
Command: npx skills add https://github.com/JosiahSiegel/claude-plugin-marketplace --skill fal-audio-josiahsiegel

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill streamlines audio-to-text and text-to-speech tasks, enabling efficient transcription, translation, subtitle generation, and voice cloning.

Core Features & Use Cases

  • Speech-to-Text (STT): Transcribe audio/video to text with high accuracy using Whisper models.
  • Translation: Translate spoken content from various languages into English.
  • Text-to-Speech (TTS): Synthesize natural-sounding speech using F5-TTS, ElevenLabs, Kokoro, or XTTS.
  • Voice Cloning: Replicate specific voices for custom audio generation.
  • Subtitle Generation: Automatically create SRT subtitles from audio or video.

Quick Start

Use the fal-audio skill to transcribe the audio from the provided URL into text.

Frequently Asked Questions about fal-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio to text and generate SRT subtitles from a video?

You can transcribe audio to text and generate SRT subtitles by processing audio or video URLs with Whisper speech-to-text models. This provides accurate transcripts and automatically creates subtitle files for your media.

Can I clone a voice and use it for text-to-speech generation?

Yes, you can clone a voice for text-to-speech generation using F5-TTS, ElevenLabs, Kokoro, or XTTS. These models replicate specific voices to synthesize natural-sounding custom speech from text input.

What is the best way to translate spoken audio from another language into English?

The best way to translate spoken audio into English is using speech-to-text transcription with translation support. This processes spoken content from various languages and directly translates it into English text.

Does this audio processing approach support multiple text-to-speech models like F5-TTS and XTTS?

Yes, this audio processing approach supports multiple text-to-speech models including F5-TTS, ElevenLabs, Kokoro, and XTTS. It integrates STT and TTS endpoints to handle diverse audio generation and transcription needs.

Do I need to provide audio files directly, or can I use URLs for speech-to-text transcription?

You can use URLs for speech-to-text transcription instead of providing audio files directly. The system processes media from provided URLs to extract audio and transcribe it into text using Whisper models.