whisper

Transcribe speech to text across 99 languages and translate to English.

Updated Apr 30, 2026
One-click install
npx skills add https://github.com/photonics-dhl/Hermes --skill whisper-photonics-dhl
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: whisper
Source: https://github.com/photonics-dhl/Hermes/tree/main/hermes-home/skills/mlops/models/whisper
Command: npx skills add https://github.com/photonics-dhl/Hermes --skill whisper-photonics-dhl

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires openai-whisper, transformers, torch, and includes references (resource) components.

What problem does it solve?

Transcribe speech to text across 99 languages and offer translations to English, enabling rapid transcription, captioning, and multilingual audio analysis.

Core Features & Use Cases

  • Automatic multilingual transcription for podcasts, meetings, videos, and lectures.
  • On-demand translation to English for cross-language workflows.
  • Model-size options and offline processing for robust performance.

Quick Start

Install the whisper package and run a transcription on your audio file to obtain a text transcript.

Frequently Asked Questions about whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe multilingual audio files to text?

To transcribe multilingual audio to text, you can use this Skill to convert speech across 99 languages into written transcripts. It processes podcasts, meetings, and lectures automatically using openai-whisper and torch.

Can I translate speech to English during audio transcription?

Yes, you can translate speech to English during audio transcription. This Skill provides on-demand translation to English for cross-language workflows, handling both multilingual transcription and English translation tasks.

Do I need Python and torch to run whisper for speech recognition?

Yes, you need Python dependencies including openai-whisper, transformers, and torch to run speech recognition. Optional GPU acceleration is supported to increase processing speed for large audio files.

What is the best way to generate captions for podcasts and meetings?

The best way to generate captions for podcasts and meetings is using automatic multilingual speech recognition. This Skill transcribes spoken audio directly into text transcripts, supporting multiple model-size options for robust offline processing.

Does offline speech recognition work for multilingual video transcription?

Yes, offline speech recognition works for multilingual video transcription. This Skill processes audio locally using model-size options, generating accurate text transcripts across 99 languages without requiring continuous internet access.