What problem does it solve?
This Skill removes the friction of building voice-enabled workflows by letting you convert text to spoken audio and transcribe spoken audio back into text with a single, reusable interface.
Core Features & Use Cases
- Text to Speech: Generate playable audio from text for narration, voiceovers, or agent responses.
- Speech to Text: Transcribe audio files or upstream audio references into text for analysis, summaries, or downstream automation.
- Multi-provider support: Choose between providers such as OpenAI, ElevenLabs, Deepgram, Groq, and Sarvam AI depending on quality, latency, language, or diarization needs.
- Use Case: Build a voice loop that listens to a recording, converts it to text, reasons over the transcript, and speaks back a response.
Quick Start
Connect the speech skill to your workflow and configure the provider credential, then send it either text for synthesis or an audio file for transcription.