speech-to-text

Convert audio to text with speaker diarization and word-level timestamps.

3|1|Updated May 29, 2026
One-click install
npx skills add https://github.com/het8802/OpenNolan --skill speech-to-text-het8802
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/het8802/OpenNolan/tree/main/.agents/skills/speech-to-text
Command: npx skills add https://github.com/het8802/OpenNolan --skill speech-to-text-het8802

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires elevenlabs, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill solves the problem of converting spoken language into written text, making it easier to transcribe audio content, create subtitles, and process spoken data.

Core Features & Use Cases

  • High Accuracy Transcription: Offers state-of-the-art accuracy with support for over 90 languages.
  • Speaker Diarization: Identifies and labels each speaker in multi-speaker audio.
  • Word-Level Timestamps: Provides precise timestamps for each word in the transcription.
  • Use Case: Ideal for transcribing meetings, interviews, and creating subtitles for videos.

Quick Start

Use the speech-to-text skill to transcribe the audio file 'meeting_recording.mp3'.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert audio to text with multiple speakers?

Audio transcription with speaker diarization identifies and labels each speaker in multi-speaker recordings. This Skill supports transcribing meetings and interviews while providing word-level timestamps for accurate subtitle creation.

Do I need an ElevenLabs API key to transcribe audio?

Yes, an ElevenLabs API key is required to perform audio transcription. The Skill requires internet access and this API dependency to deliver state-of-the-art speech recognition accuracy across over 90 languages.

What is the best way to create subtitles from an audio file?

The best way to create subtitles from an audio file is using speech recognition with word-level timestamps. This provides precise timing for each transcribed word, making it ideal for generating accurate video subtitles.

Does this speech recognition tool support languages other than English?

Yes, this speech recognition tool supports transcribing audio in over 90 languages. It provides high accuracy transcription for audio content, making it suitable for diverse international meeting notes and subtitle creation.

How to transcribe a meeting recording into written notes?

To transcribe a meeting recording into written notes, provide the audio file to the speech-to-text Skill. It converts the spoken audio into text with high accuracy, utilizing speaker diarization to label each participant.