dialogue-audio

Generate multi-speaker dialogue audio with Dia TTS via inference.sh CLI.

688|95|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/inference-sh/skills --skill dialogue-audio-inference-sh
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: dialogue-audio
Source: https://github.com/inference-sh/skills/tree/main/tools/audio/dialogue-audio
Command: npx skills add https://github.com/inference-sh/skills --skill dialogue-audio-inference-sh

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the creation of realistic multi-speaker dialogue audio, simplifying the production of podcasts, audiobooks, and character voiceovers.

Core Features & Use Cases

  • Multi-Speaker Synthesis: Generates audio with up to two distinct speakers using Dia TTS.
  • Emotion & Pacing Control: Allows fine-tuning of delivery through punctuation, non-speech sounds, and sentence structure.
  • Conversation Flow: Supports various dialogue patterns like interviews, tutorials, and debates.
  • Post-Production Assistance: Offers guidance on volume normalization, background music integration, and segment merging.
  • Use Case: Generate dialogue for a podcast episode featuring an interviewer and an expert guest, ensuring natural conversational flow and distinct voices.

Quick Start

Use the dialogue-audio skill to generate a two-speaker conversation about a new feature.

Frequently Asked Questions about dialogue-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate multi-speaker dialogue audio for a podcast?

To generate multi-speaker dialogue audio for a podcast, this Skill uses the Dia TTS model to synthesize up to two distinct speaker voices from tagged text inputs. It structures conversation patterns and controls pacing through punctuation to produce realistic conversational flows.

Can I control emotion and pacing in text-to-speech voiceover generation?

Yes, you can control emotion and pacing in text-to-speech voiceover generation by using specific punctuation, non-speech cues, and sentence structuring within the Dia TTS model. These text formatting techniques fine-tune the delivery and emotional tone of the synthesized speech.

What is the maximum number of speakers supported for audio generation?

The maximum number of speakers supported for audio generation is two. The Dia TTS model synthesizes distinct voices for up to two speakers using specific speaker tags within the text input to differentiate the dialogue.

How do I normalize volume and merge audio segments after TTS generation?

To normalize volume and merge audio segments after TTS generation, this Skill provides post-production guidance for integrating background music and combining multiple dialogue parts. These instructions help achieve consistent audio quality across the final conversation output.

Does the Dia TTS model support non-speech sounds in dialogue synthesis?

Yes, the Dia TTS model supports non-speech sounds in dialogue synthesis by interpreting specific text-based cues. This feature allows the generation of realistic audio that includes natural conversational elements alongside the spoken voiceover lines.