speech-to-text

Convert spoken audio and video into text with timestamps and speaker labels.

Updated Apr 2, 2026
One-click install
npx skills add https://github.com/gyalamanch001a/pur-new --skill speech-to-text-gyalamanch001a
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/gyalamanch001a/pur-new/tree/main/.agents/skills/speech-to-text
Command: npx skills add https://github.com/gyalamanch001a/pur-new --skill speech-to-text-gyalamanch001a

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill removes the manual burden of turning speech in audio or video into accurate written text, helping you capture conversations, meetings, lectures, and media without retyping.

Core Features & Use Cases

  • Batch Transcription: Convert uploaded audio or video files into text with language detection, timestamps, speaker diarization, and optional keyterm prompting.
  • Real-Time Streaming: Transcribe live microphone or server audio with partial and committed transcripts for interactive applications.
  • Subtitles and Notes: Generate subtitle-ready text and meeting notes from interviews, webinars, podcasts, and classroom recordings.
  • Use Case: A team records a customer call and uses this Skill to produce a searchable transcript with speaker labels, timestamps, and highlighted product terms.

Quick Start

Use the speech-to-text skill to transcribe the attached audio file into clean text with timestamps and speaker labels.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe an audio file into text with timestamps and speaker labels?

To transcribe audio into text with timestamps and speaker labels, use this Skill to process your files via ElevenLabs Scribe v2. It supports batch transcription with diarization, language detection, and keyterm prompting for accurate multilingual output.

What is the best way to generate subtitles from a video recording?

The best way to generate subtitles from video recordings is using a speech-to-text tool that outputs timestamped text. This Skill converts video audio into subtitle-ready text, supporting lectures, podcasts, and webinars with language detection.

Do I need an ElevenLabs API key to use real-time speech-to-text streaming?

Yes, real-time speech-to-text streaming requires an ELEVENLABS_API_KEY environment variable and ElevenLabs Scribe v2. This setup transcribes live microphone or server audio into partial and committed transcripts for interactive applications.

Can I transcribe live streaming audio into partial transcripts for interactive applications?

Transcribing live streaming audio into partial transcripts is supported for interactive applications. This Skill processes live microphone or server audio, delivering both partial and committed transcripts in real-time using ElevenLabs Scribe v2.

Does speech-to-text support keyterm prompting for specialized vocabulary?

Speech-to-text transcription supports optional keyterm prompting to accurately capture specialized vocabulary. This helps correctly transcribe specific product terms, names, or technical jargon during meetings, interviews, and lectures.

Why use diarization for meeting notes transcription?

Diarization identifies and labels individual speakers in meeting notes transcription, making the text searchable and organized. This Skill applies speaker diarization to audio files, clarifying who said what in conversations and interviews.