What problem does it solve? Converting spoken content in audio or video files into accurate, structured text is slow and error-prone when done manually, especially for meetings, interviews, and subtitle creation across multiple languages and speakers. ## Core Features & Use Cases - Batch Transcription: Transcribe audio and video files (MP3, WAV, MP4, and more) with word-level timestamps, speaker diarization for up to 32 speakers, and support for 90+ languages. - Real-Time Streaming: Stream live audio with ~150ms latency using partial and committed transcripts, with manual or VAD-based commit strategies for microphone input. - Accuracy Controls: Use keyterm prompting for jargon and product names, language hints, and entity detection for PII/PCI content. - Use Case: Transcribe a recorded team meeting with speaker labels, then generate timestamped subtitles for a published video using the word-level timing output. ## Quick Start Transcribe the attached audio file 'meeting.mp3' to text with speaker diarization and word-level timestamps using ElevenLabs Scribe v2.