whisper

Transcribe spoken audio into timestamped text across 99 languages.

Updated Jun 25, 2026
One-click install
npx skills add https://github.com/davpatel605-beep/hermusagent --skill whisper-davpatel605-beep
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: whisper
Source: https://github.com/davpatel605-beep/hermusagent/tree/main/backend/vendor/hermes/optional-skills/mlops/whisper
Command: npx skills add https://github.com/davpatel605-beep/hermusagent --skill whisper-davpatel605-beep

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill solves the challenge of converting spoken audio into accurate text across many languages, reducing manual transcription effort for recordings, meetings, and media.

Core Features & Use Cases

  • Speech Recognition: Transcribes audio into text with support for 99 languages and multilingual speech processing.
  • Translation and Language Processing: Converts non-English speech into English text and identifies spoken languages for audio workflows.
  • Use Case: Imagine you have hours of podcast recordings or meeting audio. Use this Skill to generate searchable transcripts, subtitles, and translated text automatically.

Quick Start

Use the whisper skill to transcribe the attached audio file and generate a timestamped transcript.

Frequently Asked Questions about whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files into text across multiple languages?

To transcribe audio into text across multiple languages, you can use speech recognition processing to convert spoken recordings into accurate text. This approach supports up to 99 languages and handles multilingual speech processing automatically.

Can I generate timestamped transcripts for podcast recordings and meeting notes?

Yes, you can generate timestamped transcripts for podcast recordings and meeting notes. Speech-to-text processing provides configurable timestamps alongside the transcribed text, creating searchable output for automated transcription workflows.

Does speech translation work for converting non-English audio into English text?

Speech translation does work for converting non-English audio into English text. The transcription process identifies spoken languages in the audio file and applies language processing to translate the speech directly into English text.

What is the best way to create subtitles from multilingual audio?

The best way to create subtitles from multilingual audio is through automated speech-to-text transcription. This process detects the spoken language, transcribes the audio, and generates text that can be formatted into subtitle files for media workflows.

Do I need to manually configure language detection for multilingual speech processing?

You do not need to manually configure language detection for multilingual speech processing. The transcription workflow identifies the spoken language automatically, though you can select configurable model sizes to match your audio processing requirements.