podcast-generation

Convert text content into two-host podcast scripts and synthesized audio.

Updated Apr 7, 2026
One-click install
npx skills add https://github.com/easyspace-ai/minote --skill podcast-generation-easyspace-ai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: podcast-generation
Source: https://github.com/easyspace-ai/minote/tree/main/skills/public/podcast-generation
Command: npx skills add https://github.com/easyspace-ai/minote --skill podcast-generation-easyspace-ai

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires requests, and includes scripts (resource) components.

What problem does it solve?

Converts text content into podcast-ready scripts and synthesized audio, enabling rapid creation of professional podcasts from written material.

Core Features & Use Cases

  • Converts articles, reports, or documentation into a two-host podcast format with natural dialogue.
  • Generates a structured JSON script and synthesizes audio via text-to-speech.
  • Produces an optional transcript for accessibility and post-production.
  • Supports English and Chinese content and end-to-end podcast workflow from script to final MP3.

Quick Start

Provide the source text and language, and request a complete podcast audio file with optional transcript.

Frequently Asked Questions about podcast-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text articles into a natural two-host podcast?

You can convert text articles into a natural two-host podcast by providing source text and language to generate a structured JSON script and synthesize multi-voice audio via text-to-speech. This orchestrates the end-to-end workflow from script to final MP3.

Does the text-to-speech workflow support English and Chinese podcast generation?

Yes, the text-to-speech workflow supports English and Chinese podcast generation. You provide the source text and language to generate a structured JSON script and synthesize multi-voice audio in either language.

How does multi-voice TTS generation work for podcast scripts?

Multi-voice TTS generation works by converting your source text into a structured JSON script formatted as natural dialogue, then synthesizing the audio with distinct voices for each host. This creates a two-host podcast MP3 from your written material.

Do I need to provide a script or just raw text to generate podcast audio?

You only need to provide raw text and specify the language to generate podcast audio. The workflow automatically creates the structured JSON script and synthesizes the multi-voice TTS audio for you.

What limitations should I expect when converting reports to podcast audio?

When converting reports to podcast audio, you should expect the workflow to require the requests dependency and process content strictly in English or Chinese. The output is a structured JSON script, synthesized multi-voice TTS audio, and an optional transcript.