speech-to-text

Convert audio and video recordings into multilingual transcripts with timestamps.

1|Updated May 5, 2026
One-click install
npx skills add https://github.com/Galaxies-dev/voice-notes --skill speech-to-text-galaxies-dev
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/Galaxies-dev/voice-notes/tree/main/.agents/skills/speech-to-text
Command: npx skills add https://github.com/Galaxies-dev/voice-notes --skill speech-to-text-galaxies-dev

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of transforming audio and video recordings into precise, searchable text, simplifying transcription tasks.

Core Features & Use Cases

  • Accurate Transcription: Convert audio files into text with support for 90+ languages, speaker diarization, and word-level timestamps.
  • Multiple Formats: Handle diverse audio and video file types up to 3GB or 10 hours in length.
  • Use Case: Transcribe long meetings or interviews to generate subtitles, searchable archives, or detailed minutes.

Quick Start

Use the speech-to-text skill to transcribe an audio file and obtain a text transcript with timestamps.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe an audio recording with multiple speakers into text?

To transcribe audio recordings with multiple speakers, this Skill provides speaker diarization to identify who is speaking, word-level timestamps, and supports 90+ languages. It processes files up to 3GB or 10 hours long.

Can I generate subtitles from a video file automatically?

Yes, you can generate subtitles from video files automatically. The Skill supports diverse audio and video file types, converting spoken content into text transcripts with timestamps suitable for subtitle creation.

Does audio transcription work with long meetings and interviews?

Audio transcription works effectively for long meetings and interviews, handling recordings up to 10 hours or 3GB in size to generate detailed minutes, searchable archives, and accurate text outputs.

How does multilingual speech-to-text handle different languages?

Multilingual speech-to-text handles different languages by utilizing ElevenLabs Scribe v2 models, supporting over 90 languages to accurately convert spoken audio into searchable text transcripts.

Do I need an API key to convert spoken audio into text?

Yes, you need an API key and internet access to convert spoken audio into text. The transcription process utilizes ElevenLabs Scribe v2 models, requiring external API connectivity for batch and real-time scenarios.