What problem does it solve?
Whisper solves the problem of turning spoken audio into readable text, including multilingual transcription and translation to English, so you can document meetings, podcasts, or media without manual typing.
Core Features & Use Cases
- Multilingual speech-to-text (99 languages): Transcribe audio in many languages and automatically identify the spoken language.
- Transcription and translation: Produce either a faithful transcription or an English translation of non-English audio.
- Timestamps and refinement: Generate word/segment timestamps and improve accuracy for technical content by providing an initial prompt.
Example use case: Convert a recorded multilingual team meeting into structured notes by transcribing the audio and translating the result to English for faster review.
Quick Start
Use the whisper skill to transcribe and translate your audio file to English, and return the generated text.