speech-to-text

Transcribe spoken audio to text using the ElevenLabs Scribe API.

Updated May 23, 2026
One-click install
npx skills add https://github.com/xingBaGan/FANovelist --skill speech-to-text-xingbagan
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/xingBaGan/FANovelist/tree/main/src/openharness/openmontage/.claude/skills/speech-to-text
Command: npx skills add https://github.com/xingBaGan/FANovelist --skill speech-to-text-xingbagan

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires elevenlabs, and includes scripts (resource) and references (resource) components.

What problem does it solve?

The Skill simplifies the process of transcribing audio content into text, which can be used for creating subtitles, transcribing meetings, or archiving spoken content for reference.

Core Features & Use Cases

  • Accurate Transcription: Uses ElevenLabs' Scribe v2 to transcribe speech to text with high accuracy.
  • Multilingual Support: Transcription available in 90+ languages, including speaker diarization and word-level timestamps.
  • Use Case: Automate the transcription of customer service calls for post-analysis and record-keeping.

Quick Start

Convert an audio file 'call_recording.wav' to text by running the 'speech-to-text' command with the 'audio.mp3' input.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe an audio file to text automatically?

To transcribe an audio file to text automatically, run the speech-to-text command with your input file. This utilizes the ElevenLabs Scribe v2 API to accurately convert spoken audio into readable text.

Does AI transcription support multiple languages and speaker identification?

AI transcription supports over 90 languages and includes advanced features like speaker diarization and word-level timestamps. This allows for accurate meeting transcription and separation of multiple speakers.

Do I need an ElevenLabs API key to use speech recognition for meeting transcription?

Yes, you need an ElevenLabs API key for speech recognition. The Skill depends on the ElevenLabs Scribe API to process spoken language and automate transcription workflows for meetings or customer service calls.

Can I use audio-to-text conversion for creating subtitles from call recordings?

You can use audio-to-text conversion for creating subtitles or archiving spoken content. It automates the transcription of customer service calls and recordings into readable text for post-analysis.

What is the best way to automate transcription of customer service calls?

The best way to automate transcription of customer service calls is using the ElevenLabs Scribe API. It provides high-accuracy speech-to-text processing, suitable for post-analysis and record-keeping workflows.