speech-to-text

Transcribe audio and video content into searchable text using the ElevenLabs Scribe v2 API.

Updated Mar 25, 2026
One-click install
npx skills add https://github.com/almazom/agents_slash_skills --skill speech-to-text-almazom
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/almazom/agents_slash_skills/tree/main/speech-to-text
Command: npx skills add https://github.com/almazom/agents_slash_skills --skill speech-to-text-almazom

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Transcribing audio or video to text can be time-consuming and error-prone. This Skill provides accurate transcription with language support, diarization, and timestamps to speed up content processing and searchability.

Core Features & Use Cases

  • Transcribes audio and video content into reliable text with optional speaker diarization and timestamp data.
  • Supports batch transcription and real-time streaming for subtitles, meeting transcripts, and media archives.
  • Detects language hints and provides word-level timestamps for precise subtitles and searchable transcripts.

Quick Start

Transcribe an audio file to text using the ElevenLabs Scribe v2 model.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files into text with timestamps and speaker diarization?

Transcribe audio and video content into accurate text with word-level timestamps and speaker diarization using the ElevenLabs Scribe v2 API. It processes meetings, interviews, and media archives while detecting language hints to speed up content searchability.

Can I generate real-time streaming subtitles for video content?

Real-time streaming subtitles for video content are supported to provide immediate transcription. The Skill handles live meeting transcripts and media workloads, converting spoken audio into searchable text as the content plays.

Do I need an ElevenLabs API key to convert speech to text?

Converting speech to text requires an ElevenLabs API key and internet access. The Skill uses the Scribe v2 model to process audio and video inputs, applying language detection and diarization for accurate transcription outputs.

What is the best way to transcribe long-form interviews and meetings?

The best way to transcribe long-form interviews and meetings is using batch transcription with the Scribe v2 API. It provides reliable text with optional timestamp data and diarization, ensuring precise subtitles and searchable transcripts for extended media workloads.

Does speech-to-text transcription work with batch processing for media archives?

Speech-to-text transcription supports batch processing for media archives, converting large volumes of audio and video into searchable text. It applies language hints and speaker diarization to organize extensive content libraries efficiently.

Why use diarization and language hints for audio transcription?

Diarization and language hints improve audio transcription accuracy by separating distinct speakers and identifying spoken languages. This ensures precise word-level timestamps for subtitles and creates searchable, structured transcripts for meetings and interviews.