whisper

Transcribe multilingual audio into text across 99 languages with Python.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/travelinman1013/leroys-agent --skill whisper-travelinman1013
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: whisper
Source: https://github.com/travelinman1013/leroys-agent/tree/main/optional-skills/mlops/whisper
Command: npx skills add https://github.com/travelinman1013/leroys-agent --skill whisper-travelinman1013

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Whisper automates the conversion of speech to text across 99 languages, reducing manual transcription effort and enabling searchable transcripts for accessibility, indexing, and multilingual analytics.

Core Features & Use Cases

  • Multilingual speech recognition across 99 languages
  • Transcription and translation to English
  • Model size options and GPU acceleration for faster processing
  • Use cases include podcast transcription, video captions, meeting notes, and archival of multilingual audio

Quick Start

Load a Whisper model in Python and transcribe an audio file to obtain text.

Frequently Asked Questions about whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe multilingual audio into text using Python?

To transcribe multilingual audio into text using Python, load a Whisper model and pass the audio file to obtain accurate text across 99 languages. It supports language-specific transcription options for desktop and cloud workflows.

Can I translate foreign language audio to English during speech recognition?

Yes, you can translate foreign language audio to English during speech recognition. The transcription workflow supports optional translation to English alongside multilingual transcription across 99 languages.

Does multilingual speech recognition support GPU acceleration for faster processing?

Multilingual speech recognition supports GPU acceleration for faster processing. It offers model size options and GPU acceleration to optimize transcription speed for podcast transcripts, video captions, and meeting notes.

What is the best way to generate podcast transcripts and video captions automatically?

The best way to generate podcast transcripts and video captions automatically is using multilingual speech recognition. It automates converting speech to text across 99 languages, reducing manual transcription effort for searchable, accessible archives.

Do I need GPU hardware to run speech recognition for meeting notes?

You do not need GPU hardware to run speech recognition for meeting notes. GPU acceleration is optional for faster processing, but the Python-based model loading works across desktop and cloud workflows without mandatory GPU requirements.

How many languages does Whisper speech recognition support for audio archiving?

Whisper speech recognition supports 99 languages for audio archiving. It provides multilingual speech recognition and optional English translation, ideal for searchable transcripts and multilingual analytics across diverse audio sources.