What problem does it solve?
Users often need to switch between multiple disconnected tools for text-to-speech generation, audio transcription, format conversion, and audio editing, leading to inconsistent results and wasted time managing separate workflows for common audio tasks.
Core Features & Use Cases
- Text-to-Speech & Transcription: Generate natural speech from text using cloud or local TTS engines, and transcribe audio files to timestamped text with OpenAI Whisper.
- Audio Format Conversion: Convert between all common audio formats (MP3, WAV, FLAC, AAC, etc.) with customizable bitrate, sample rate, and channel settings.
- Audio Editing & Podcast Production: Trim, concatenate, and normalize audio loudness, plus assemble full podcast episodes with intro/outro segments and professional-level loudness standards.
- Use Case: A podcaster can record raw interview segments, transcribe them automatically, combine the segments with intro/outro music, normalize the loudness to streaming standards, and export a ready-to-publish MP3 file all in one consistent workflow.
Quick Start
Use the audio-gen skill to transcribe the attached meeting recording 'team-standup.mp3' to a text file and convert the audio to a 128kbps MP3 format optimized for sharing.