What problem does it solve?
This Skill removes the manual effort of listening back to recordings and turning speech into usable text, making audio and video content searchable, editable, and ready for downstream workflows.
Core Features & Use Cases
- Batch transcription: Convert uploaded audio or video files into full transcripts with language detection, timestamps, diarization, and optional entity or audio-event tagging.
- Real-time transcription: Stream microphone or server audio for partial and committed transcripts, with commit strategies that support live conversation, capture, and playback workflows.
- Production transcription workflows: Generate subtitles, meeting notes, interview transcripts, podcast text, and speaker-labeled transcripts from supported media formats.
Quick Start
Transcribe the attached audio or video file with ElevenLabs Scribe v2 and return the transcript text with timestamps and speaker labels if available.