What problem does it solve?
Manual transcription of audio and video content is slow, expensive, and prone to human error, especially for long recordings, multi-speaker meetings, or content in less common languages. This Skill automates accurate transcription with optional speaker labels and timestamps, reducing transcription time from hours to minutes.
Core Features & Use Cases
- Multi-Language Batch Transcription: Convert pre-recorded audio or video files up to 10 hours long to text, with support for 90+ languages and custom keyterms to boost accuracy for domain-specific jargon like product names or technical terms.
- Speaker Diarization & Timestamped Output: Automatically identify different speakers and generate word-level timestamps for use in subtitles, meeting notes, or video editing workflows.
- Real-Time Live Transcription: Stream live audio with ultra-low latency (~150ms) for use cases like live captioning, voice agent backends, or real-time meeting note taking.
Quick Start
Use the speech-to-text skill to transcribe the attached team meeting recording 'q3-sync.mp3' and output a timestamped transcript with speaker labels separated by speaker ID.