whisper

Transcribe spoken audio into text across 99 languages using openai-whisper.

1|Updated Jun 25, 2026
One-click install
npx skills add https://github.com/Signmanal/VIGIL --skill whisper-signmanal
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: whisper
Source: https://github.com/Signmanal/VIGIL/tree/main/optional-skills/mlops/whisper
Command: npx skills add https://github.com/Signmanal/VIGIL --skill whisper-signmanal

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Manual transcription of audio content is time-consuming, especially for multilingual recordings, podcasts, or non-English meetings. This Skill automates speech-to-text conversion, cutting transcription work from hours to minutes without manual typing.

Core Features & Use Cases

  • Multilingual Transcription: Convert speech in 99 languages to accurate text, including low-resource languages with specialized model support.
  • English Translation: Automatically translate non-English audio content to English text for cross-team collaboration.
  • Timestamped Output: Generate time-stamped transcripts for subtitles, meeting notes, or audio editing workflows.
  • Use Case: A podcast producer can use this Skill to transcribe a 1-hour Spanish-language episode, translate it to English, and export SRT subtitle files for global audiences in under 10 minutes.

Quick Start

Use the whisper skill to transcribe the audio file 'team-standup.mp3' and output a timestamped English transcript.

Frequently Asked Questions about whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe multilingual audio recordings to text?

To transcribe multilingual audio, this Skill processes spoken recordings and outputs written text across 99 supported languages. It automates speech-to-text conversion, eliminating manual typing for podcasts, meetings, and non-English audio content.

Can I translate non-English speech to English text automatically?

Yes, automatic English translation converts non-English speech to English text. This Skill handles multilingual audio processing and provides English translation output, enabling cross-team collaboration for global audiences.

Do I need ffmpeg and Python to run Whisper for speech-to-text conversion?

Yes, speech-to-text conversion requires the openai-whisper Python library and ffmpeg for audio decoding. Optional GPU acceleration can be configured for faster inference when processing large audio files.

How do I generate timestamped transcripts for subtitle files?

To generate timestamped transcripts, this Skill processes audio files and outputs time-stamped text. This format is ideal for creating SRT subtitle files, meeting notes, or audio editing workflows for global audiences.

What is the best way to transcribe a 1-hour podcast quickly?

Automated speech-to-text with optional GPU acceleration is the best way to transcribe a 1-hour podcast. This Skill processes long-form audio and generates accurate text transcripts in under 10 minutes, significantly faster than manual typing.

Does speech-to-text work with low-resource languages?

Speech-to-text works with low-resource languages through specialized model support. This Skill handles multilingual audio processing across 99 supported languages, ensuring accurate transcription even for less commonly spoken languages.