speech-to-text

Transcribe audio and video files into text using ElevenLabs Scribe v2.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/MythologIQ/Zo-Qore --skill speech-to-text-mythologiq
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/MythologIQ/Zo-Qore/tree/main/.claude/commands/scripts/custom/speech-to-text
Command: npx skills add https://github.com/MythologIQ/Zo-Qore --skill speech-to-text-mythologiq

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill converts spoken audio into written text, making audio and video content searchable, editable, and accessible.

Core Features & Use Cases

  • High-Accuracy Transcription: Utilizes ElevenLabs Scribe v2 for precise transcription across 90+ languages.
  • Advanced Features: Supports speaker diarization, word-level timestamps, and keyterm prompting for specialized vocabulary.
  • Use Case: Transcribe a recorded lecture to create study notes, generate subtitles for a video, or convert a meeting recording into a searchable transcript.

Quick Start

Use the speech-to-text skill to transcribe the audio file 'meeting_recording.mp3'.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe an audio file to text with speaker diarization?

You can transcribe audio with speaker diarization using this Skill, which leverages the ElevenLabs Scribe v2 model to identify individual speakers and generate accurate text transcripts across over 90 languages.

Can I get word-level timestamps when transcribing video files?

Yes, you can get word-level timestamps when transcribing video files. The Skill uses ElevenLabs Scribe v2 to provide precise timing data for each spoken word alongside the generated text transcript.

How do I transcribe specialized vocabulary and keyterms accurately?

To accurately transcribe specialized vocabulary, this Skill supports keyterm prompting, a feature that guides the ElevenLabs Scribe v2 model to correctly recognize and output specific domain terms during audio transcription.

Does speech to text transcription support Python and JavaScript APIs?

Yes, speech to text transcription supports Python, JavaScript, and cURL APIs, allowing developers to integrate batch and real-time audio processing directly into their applications and workflows.

What is the best way to generate subtitles for a recorded lecture?

The best way to generate subtitles for a recorded lecture is using this Skill, which converts spoken audio into searchable text with word-level timestamps, making your video content editable and accessible.

What languages are supported for audio transcription?

Audio transcription supports over 90 languages through the ElevenLabs Scribe v2 model, ensuring high-accuracy conversion of spoken audio content into written text regardless of the input language.