fish-audio

Generate expressive audio clips with Fish Audio S2 TTS using bracket emotion tags.

1.0k|158|Updated Feb 7, 2026
One-click install
npx skills add https://github.com/vellum-ai/vellum-assistant --skill fish-audio-vellum-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fish-audio
Source: https://github.com/vellum-ai/vellum-assistant/tree/main/skills/fish-audio
Command: npx skills add https://github.com/vellum-ai/vellum-assistant --skill fish-audio-vellum-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Recording and producing expressive audio content such as narration, memos, and announcements can be time-consuming and error-prone. Fish Audio TTS with bracket emotion tags enables rapid generation of expressive audio clips on demand, improving consistency and efficiency.

Core Features & Use Cases

  • Bracket emotion tags provide expressive prosody for narration, memos, podcasts, and alerts.
  • API-based generation using Fish Audio S2 Pro with configurable endpoint and key management.
  • Quick assembly workflow for multiple clips using standard audio tools to produce coherent voice content.

Quick Start

Generate a sample narration using Fish Audio S2 TTS with bracket emotion tags.

Frequently Asked Questions about fish-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate expressive audio narration with emotion tags?

Generate expressive audio narration by using bracket emotion tags within your text input to control prosody. The Fish Audio S2 TTS engine processes these tags to produce spoken content with varied emotional delivery.

Can I use bracket tags to control voice prosody for text-to-speech memos?

Yes, you can use bracket emotion tags to control voice prosody for text-to-speech memos. This feature allows you to inject expressive emotional cues directly into your generated voice memos.

What do I need to set up before generating audio clips on demand?

Before generating audio clips on demand, you need a Fish Audio API key stored in a credential store and a configured API endpoint. The system uses these to access the S2 Pro TTS service and output mp3 files.

How does API-based audio generation handle multiple audio clips?

API-based audio generation handles multiple audio clips by utilizing a quick assembly workflow with standard audio tools. This process combines individual TTS outputs into a single, coherent piece of voice content.

What is the default audio output format for generated voice announcements?

The default audio output format for generated voice announcements is mp3. The Fish Audio S2 TTS engine outputs all narration, memos, and alerts as mp3 files by default.