What problem does it solve?
Whisper removes the manual burden of converting spoken audio into usable text, making it easier to capture meetings, interviews, lectures, podcasts, and multilingual recordings with high accuracy.
Core Features & Use Cases
- Multilingual speech recognition: Transcribe audio in 99 languages and detect the spoken language automatically.
- Translation and timestamps: Translate non-English speech into English and generate segment or word-level timing for subtitles and review.
- Practical workflows: Use it for podcast transcription, meeting notes, subtitle creation, noisy audio cleanup, and batch processing of audio files.
- Model flexibility: Choose from tiny through large model sizes to balance speed, quality, and hardware requirements.
Quick Start
Use the whisper skill to transcribe the attached audio file, detect its language, and return a clean transcript with timestamps.