elevenlabs-scribe

Transcribe audio and video into time-aligned transcripts with speaker metadata.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill elevenlabs-scribe
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: elevenlabs-scribe
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/speech-to-text/elevenlabs-scribe
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill elevenlabs-scribe

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill turns audio and video recordings into accurate, time-aligned transcripts and caption drafts while reducing the manual work of speaker labeling, transcript cleanup, privacy review, and delivery preparation.

Core Features & Use Cases

  • Speech-to-Text Production: Transcribe audio and video with Scribe v2 using language hints, word or character timestamps, audio-event tags, and curated keyterms.
  • Speaker and Caption Workflows: Handle diarization, isolated multichannel recordings, speaker turns, edit preparation, SRT and VTT drafts, and accessibility caption review.
  • Privacy and Delivery Guardrails: Plan entity detection and redaction, webhook processing, artifact custody, retention controls, consent checks, provenance tracking, and human QA.
  • Use Case: Prepare a searchable, speaker-labeled podcast transcript with word timings and keyterms, then create reviewed caption drafts and candidate clips from the final recording.

Quick Start

Use the ElevenLabs Scribe skill to transcribe the attached recording with word timestamps, speaker labels, relevant keyterms, preserved raw JSON, and a caption QA checklist.

Frequently Asked Questions about elevenlabs-scribe

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio recordings with speaker labels and word timestamps?

Transcribe audio recordings using Scribe v2 speech-to-text to generate production-ready outputs with word-level timestamps, diarization, speaker labeling, and audio-event tags. Configure language hints and curated keyterms to improve accuracy across podcasts, interviews, and customer calls.

Can I generate SRT and VTT caption drafts from video recordings?

Generate SRT and VTT caption drafts from video recordings by processing speech-to-text outputs with character timestamps and speaker turns. Run accessibility caption review and edit preparation workflows to refine the drafts before final delivery.

Does speech-to-text transcription support privacy redaction and entity detection?

Speech-to-text transcription supports privacy redaction through entity detection, consent checks, and retention controls. Plan entity redaction and provenance tracking during transcription to ensure compliance when handling sensitive customer calls or interviews.

What is the best way to handle diarization for multi-speaker podcast transcripts?

Handle diarization for multi-speaker podcast transcripts by using isolated multichannel recordings, speaker turn labeling, and curated keyterms. Preserve raw JSON outputs and apply human quality assurance to verify speaker labels and transcript accuracy.

How do I process webhook audio inputs for automated transcription pipelines?

Process webhook audio inputs by routing recordings through Scribe v2 transcription with configured language hints and timestamp settings. Manage artifact custody, provenance tracking, and retention controls to maintain output integrity across automated pipelines.

When should I apply human quality assurance to speech-to-text transcripts?

Apply human quality assurance to speech-to-text transcripts when preparing accessibility captions, privacy redactions, or critical transcript content. Human review ensures speaker labels, keyterms, and timing metadata meet production-ready standards before localization handoff.