ai-podcast-creation

Generate AI podcasts and audio content with text-to-speech and media merging.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/RomainGRAS42/Procedio-AI --skill ai-podcast-creation
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-podcast-creation
Source: https://github.com/RomainGRAS42/Procedio-AI/tree/main/.agents/skills/ai-podcast-creation
Command: npx skills add https://github.com/RomainGRAS42/Procedio-AI --skill ai-podcast-creation

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill automates the creation of AI-powered podcasts, audiobooks, and voice content, simplifying audio production workflows.

Core Features & Use Cases

  • Text-to-Speech: Generate audio from text using various high-quality AI voices (Kokoro TTS, DIA TTS).
  • Multi-Voice Conversations: Create dialogues and interviews with distinct AI speakers.
  • AI Music Generation: Add intro/outro music and background tracks.
  • Audio Merging & Editing: Combine audio segments, apply crossfades, and mix background music.
  • Use Case: Quickly produce a podcast episode by generating a script with an LLM, synthesizing voiceovers for hosts, adding background music, and merging all elements into a final audio file.

Quick Start

Use the infsh CLI to run the kokoro-tts app with the text "Welcome to the AI Frontiers podcast." and the voice "am_michael".

Frequently Asked Questions about ai-podcast-creation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a podcast from text using AI voiceover?

To generate a podcast from text using AI voiceover, provide a text script to the text-to-speech engine, select an AI voice like am_michael, and the workflow synthesizes the audio. It uses tools like Kokoro TTS to convert written scripts into spoken audio content.

Can I create multi-voice conversations for an AI podcast?

Yes, you can create multi-voice conversations for an AI podcast by assigning distinct AI speakers to different dialogue parts. The Skill synthesizes interviews and dialogues using multiple distinct voices to simulate host and guest interactions.

What is the best way to add background music to AI-generated audio?

The best way to add background music to AI-generated audio is through the built-in audio merging tools. The Skill generates AI music for intro and outro tracks, then mixes and applies crossfades to combine these background tracks with the synthesized voiceover.

Does the text-to-speech workflow support full episode production pipelines?

Yes, the text-to-speech workflow supports full episode production pipelines. It automates the entire process from generating scripts with an LLM and synthesizing voiceovers to adding background music and merging all elements into a final audio file.

Can I use this Skill to produce audiobooks and audio newsletters?

Yes, you can use this Skill to produce audiobooks and audio newsletters. It facilitates automated workflows for various voice content formats, converting written text into synthesized speech for long-form audio distribution.

What are the limitations of using AI voices for audio content generation?

The limitations of using AI voices for audio content generation depend on the available TTS models like Kokoro TTS and DIA TTS. While the Skill automates voice synthesis and audio merging, complex emotional inflections or non-standard audio editing may require manual post-processing.