What problem does it solve? Creating realistic multi-speaker audio for podcasts, audiobooks, and explainers normally requires voice actors or complex TTS configuration. This Skill generates natural two-person conversations with Dia TTS via the inference.sh CLI, controlling emotion, pacing, and turn-taking through simple text markup. ## Core Features & Use Cases - Speaker Tag Control: Uses [S1] and [S2] tags to alternate between two consistent voices within a generation. - Emotion & Pacing Control: Interprets punctuation (., !, ?, ...) and non-speech cues like (laughs), (sighs), and (whispers) for expressive delivery. - Conversation Patterns: Provides templates for interviews, tutorials, debates, and explainers, plus post-production merging of segments and background music. - Use Case: A content creator writes a short script for a product explainer video, tags each line with [S1] or [S2], adds emotion cues, and generates a ready-to-use dialogue MP3 in one command. ## Quick Start Ask the AI to generate a two-speaker dialogue audio clip with Dia TTS from a short script using [S1] and [S2] speaker tags.