mofa-fm

Synthesize speech from text using preset or cloned voices via ominix-api.

11|12|Updated Feb 28, 2026
One-click install
npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-fm
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mofa-fm
Source: https://github.com/mofa-org/mofa-skills/tree/main/mofa-fm
Command: npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-fm

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

MoFA FM provides high-quality text-to-speech with preset voices and optional custom voice cloning via ominix-api, enabling scalable audio production without manual recording.

Core Features & Use Cases

  • Preset voices for quick TTS: Vivian, Serena, Ryan, etc.
  • Custom voice cloning: upload a reference clip to create a named voice for downstream TTS.
  • Voice management: save, list, and delete custom voices; set a default voice.
  • Language, emotion prompts, and speed control to shape output for different contexts.
  • Output formats and local file delivery: WAV/MP3, with generated audio ready for delivery or integration.

Quick Start

Call fm_tts with the full text to generate speech and receive the audio file.

Frequently Asked Questions about mofa-fm

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech with a custom voice on macOS?

To convert text to speech on macOS, you can use custom voice cloning by uploading a reference audio clip to create a named voice, then synthesize speech from text using that cloned voice via the ominix-api.

Can I manage and save cloned voices locally for future text-to-speech generation?

Yes, local voice management allows you to save, list, and delete custom cloned voices. You can set a default voice to streamline future text-to-speech generation tasks.

What output formats are supported when generating audio from text?

Text-to-speech generation supports WAV and MP3 output formats. Generated audio files are delivered locally to your macOS environment for immediate playback or integration.

How do I adjust emotion and speed for voice-over text-to-speech?

You can shape voice-over output by applying language settings, emotion prompts, and speed control parameters during speech synthesis to match different contextual requirements.

Does text-to-speech generation work without manual recording for accessibility read-alouds?

Yes, preset voices like Vivian, Serena, and Ryan enable quick text-to-speech for accessibility read-alouds, providing scalable audio production without manual recording.

Do I need a specific environment to run voice cloning and TTS synthesis?

Yes, voice cloning and TTS synthesis require a macOS environment to function, as local voice management and audio file delivery are enforced through macOS system operations.