fish-audio

Generate speech, transcribe audio, and create voice narration with Fish Audio models.

199|25|Updated Jan 25, 2026
One-click install
npx skills add https://github.com/refly-ai/refly-skills --skill fish-audio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fish-audio
Source: https://github.com/refly-ai/refly-skills/tree/main/skills/fish-audio
Command: npx skills add https://github.com/refly-ai/refly-skills --skill fish-audio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill automates the creation and processing of audio content, enabling efficient text-to-speech conversion, audio transcription, and voice narration generation.

Core Features & Use Cases

  • Text-to-Speech: Convert written text into natural-sounding speech across multiple languages.
  • Audio Transcription: Transcribe spoken words from audio files into written text.
  • Voice Narration: Generate high-quality voiceovers for various applications.
  • Use Case: A content creator needs to generate an audio version of a blog post for their podcast and also transcribe an interview. This Skill can handle both tasks.

Quick Start

Use the fish-audio skill to convert the text 'Hello, world!' into speech using the default voice.

Frequently Asked Questions about fish-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech for multiple languages?

Audio transcription converts spoken words from provided audio files into written text. You supply the audio file, and the system processes the audio content to extract and output the text transcription.

Can I generate voice narration for a blog post?

Yes, voice narration generates high-quality voiceovers from written text. You input your blog post text, and the system outputs an audio version suitable for podcasts or other content applications.

How do I transcribe an audio file to text?

Audio transcription processes provided audio files to convert spoken words into written text. You supply the audio file, and the system extracts the spoken content to output a text transcription.

Does the text-to-speech model support specific voice styles?

The text-to-speech model supports specific voice styles for high-quality voice output. You can specify the desired voice style when inputting your text to generate customized audio narration.

What is the best way to automate audio content creation and processing?

Automating audio content creation involves using AI models for text-to-speech conversion, audio transcription, and voice narration generation. This handles both generating audio from text and transcribing audio files into text.