What problem does it solve?
This Skill turns audio and video recordings into accurate, time-aligned transcripts and caption drafts while reducing the manual work of speaker labeling, transcript cleanup, privacy review, and delivery preparation.
Core Features & Use Cases
- Speech-to-Text Production: Transcribe audio and video with Scribe v2 using language hints, word or character timestamps, audio-event tags, and curated keyterms.
- Speaker and Caption Workflows: Handle diarization, isolated multichannel recordings, speaker turns, edit preparation, SRT and VTT drafts, and accessibility caption review.
- Privacy and Delivery Guardrails: Plan entity detection and redaction, webhook processing, artifact custody, retention controls, consent checks, provenance tracking, and human QA.
- Use Case: Prepare a searchable, speaker-labeled podcast transcript with word timings and keyterms, then create reviewed caption drafts and candidate clips from the final recording.
Quick Start
Use the ElevenLabs Scribe skill to transcribe the attached recording with word timestamps, speaker labels, relevant keyterms, preserved raw JSON, and a caption QA checklist.