mofa-podcast

Convert topics or scripts into multi-speaker podcast MP3s with TTS voices.

11|12|Updated Feb 28, 2026
One-click install
npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-podcast
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: mofa-podcast
Source: https://github.com/mofa-org/mofa-skills/tree/main/mofa-podcast
Command: npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-podcast

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

Generating professional multi-speaker podcasts from topic briefs or texts without expensive recording setups or editing.

Core Features & Use Cases

  • Script-driven podcast generation from topic or provided text with emotion cues and music guidelines.
  • Supports built-in voices plus custom/cloned voices via mofa-fm, with per-segment timeline assembly and BGM.
  • Outputs ready-to-publish MP3s and segment audio for post-production workflows, scalable for episodes and shows.

Quick Start

Provide a topic or markdown script to generate a complete multi-speaker podcast using built-in or cloned voices.

Frequently Asked Questions about mofa-podcast

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a multi-speaker podcast from a text script?

You generate a multi-speaker podcast by providing a topic or an approved markdown script, which the system transforms into a scripted audio show using TTS voices, emotion cues, and music guidelines.

Can I use cloned voices for podcast audio assembly?

Yes, you can use cloned voices for podcast audio assembly. The system supports built-in voices and custom voice cloning via mofa-fm to assign distinct voices to 1-5 speakers per segment.

Do I need FFmpeg to output podcast MP3 files?

You do not strictly need FFmpeg to output podcast audio. The system outputs a final MP3 file when FFmpeg is available, but it defaults to generating WAV files if FFmpeg is missing from your environment.

How many speakers can I include in a single TTS podcast generation?

You can include between 1 and 5 speakers in a single TTS podcast generation. The system supports multi-speaker configurations suitable for short-form podcasts, interviews, talk shows, and educational narrations.

What is the best way to add background music and emotion cues to TTS audio?

The best way to add background music and emotion cues is by using an approved markdown script format. The system processes these script guidelines to drive per-segment timeline assembly and BGM integration during audio generation.

What audio formats are outputted for post-production podcast workflows?

The audio formats outputted for post-production workflows are MP3 and WAV. The system generates ready-to-publish MP3s by default and provides individual segment audio files to allow further editing.