whisper

Transcribe audio into text across 99 languages and translate to English.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/valentinuuiuiu/vikarma --skill whisper-valentinuuiuiu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: whisper
Source: https://github.com/valentinuuiuiu/vikarma/tree/main/hermes_agent/skills/mlops/models/whisper
Command: npx skills add https://github.com/valentinuuiuiu/vikarma --skill whisper-valentinuuiuiu

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires openai-whisper, transformers, torch, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of converting spoken language into written text, providing accurate transcription and translation capabilities across multiple languages.

Core Features & Use Cases

  • Multilingual Speech Recognition: Supports 99 languages, offering flexibility for global use.
  • Robust Transcription: Accurately transcribes speech-to-text for various audio sources.
  • Translation to English: Offers translation capabilities for non-English audio into English.
  • Use Case: Ideal for transcription of podcasts, meetings, and multilingual content.

Quick Start

Use the whisper skill to transcribe the audio file 'meeting.mp3'.

Frequently Asked Questions about whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files into text for podcast episodes and meetings?

Yes, this speech recognition service provides multilingual transcription across 99 languages and automatically translates non-English audio into English text.

Can I use multilingual speech recognition to translate non-English audio to English text?

Yes, this multilingual speech recognition service supports 99 languages and provides translation capabilities to convert non-English audio input directly into English text.

Does speech recognition with OpenAI Whisper require PyTorch and Transformers installed?

Yes, running speech recognition with OpenAI Whisper requires PyTorch and Transformers installed, as these dependencies provide the underlying machine learning framework for processing audio.

What is the best way to process multilingual audio for meeting notes?

To transcribe audio files into text, you can use this speech recognition service to convert spoken language from podcasts and meetings into accurate written transcripts.

Are there limitations when using speech recognition for transcription across 99 languages?

While supporting 99 languages for transcription, speech recognition accuracy may vary depending on audio quality, background noise, and the specific dialects present in the input audio.