speech-to-text

Transcribe audio and video to text using ElevenLabs Scribe v2.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/hanzoai/video --skill speech-to-text-hanzoai
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/hanzoai/video/tree/main/.agents/skills/speech-to-text
Command: npx skills add https://github.com/hanzoai/video --skill speech-to-text-hanzoai

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Transcribe audio to text accurately from speeches, podcasts, and meetings to enable searchable transcripts, captions, and analysis.

Core Features & Use Cases

  • High-accuracy transcription with ElevenLabs Scribe v2 across 90+ languages, with optional speaker diarization and timestamps.
  • Subtitle and caption generation for videos, meetings, and lectures, plus searchable transcripts for content indexing.
  • Real-world example: Convert a 60-minute meeting recording into a timestamped transcript with speaker labels for minutes and action items.

Quick Start

Upload an audio file to generate a full transcript with optional timestamps.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio to text with timestamps and speaker labels?

You can transcribe audio to text with word-level timestamps and speaker diarization using the ElevenLabs Scribe v2 model, which processes audio files to generate labeled transcripts for meetings and videos.

What is the best way to generate subtitles from an audio or video file?

Generating subtitles from audio or video files is handled by the speech-to-text transcription process, converting spoken content into accurate text captions across 90+ languages for content indexing and accessibility.

Do I need an API key to transcribe audio or generate meeting transcripts?

Yes, you need an ElevenLabs API key and internet access to run the Scribe v2 transcription model, which converts your audio recordings into searchable text transcripts.

Can I transcribe multilingual audio recordings into text?

Multilingual audio transcription is supported across 90+ languages, allowing you to convert diverse speech recordings into accurate text transcripts using the Scribe v2 model.

Does speech-to-text transcription work for real-time audio and batch processing?

The transcription skill supports both real-time and batch audio processing, enabling you to transcribe live speech or pre-recorded audio files into text with timestamps and diarization.

What are the limitations of using Scribe v2 for audio transcription?

Scribe v2 audio transcription requires internet access and a valid ElevenLabs API key, meaning it cannot operate offline and depends on external connectivity to process speech-to-text conversions.