fish-audio

Convert text into expressive MP3 speech using Fish Audio S2 TTS.

2|2|Updated May 13, 2026
One-click install
npx skills add https://github.com/autonomy-cloud/kairos-interface --skill fish-audio-autonomy-cloud
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: fish-audio
Source: https://github.com/autonomy-cloud/kairos-interface/tree/main/skills/fish-audio
Command: npx skills add https://github.com/autonomy-cloud/kairos-interface --skill fish-audio-autonomy-cloud

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It solves the problem of producing natural-sounding spoken audio quickly, without needing a voice actor or manually editing performances.

Core Features & Use Cases

  • Generate expressive TTS clips on demand using Fish Audio S2 with bracket emotion and prosody tags, making messages sound tailored rather than flat.
  • Create narration and multi-part audio by generating multiple segments and combining them with silence-aware stitching for smooth pacing.
  • Use for real spoken content such as voice memos, announcements, podcast intros, narration, and dramatic readings with consistent voice control.

Quick Start

Ask your assistant to generate an MP3 narration from the text "Good morning everyone! [excited] Today we launch something new." and deliver it as a file attachment.

Frequently Asked Questions about fish-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text into expressive MP3 narration with controllable emotion?

You can create expressive TTS audio by adding bracket emotion tags like [excited] directly into your text. The Fish Audio S2 API processes these tags to generate MP3 clips with controlled emotion and prosody for natural-sounding narration.

Can I generate multi-part audio segments and stitch them together for podcast intros?

Yes, you can generate multiple TTS segments and combine them using silence-aware stitching to create smooth pacing. This enables seamless multi-part audio generation for podcast intros, announcements, and dramatic readings with consistent voice control.

Do I need a voice reference ID and API key to use Fish Audio S2 for voice memos?

Yes, generating voice memos requires a configured voice reference ID and secure API key credentials. These authenticate your Fish Audio S2 TTS API calls and ensure the spoken audio maintains a consistent, specific voice profile.

What's the best way to add prosody control to spoken audio announcements?

The best way to add prosody control to audio announcements is by using bracket emotion tags within your text. The TTS engine interprets these tags to adjust delivery style, ensuring announcements sound natural rather than flat.

Does audio generation support output formats other than MP3 for voice delivery?

The Skill requires an output format such as MP3 to deliver the generated spoken audio clips. MP3 is the specified format for producing natural-sounding voice memos, narration, and announcements through the Fish Audio S2 API.