What problem does it solve?
Whisper converts spoken audio into readable text, handling multilingual transcription and optional translation to English so you can avoid manual playback and note-taking.
Core Features & Use Cases
- Multilingual speech-to-text: Transcribe 99 languages with timestamps for segment-level and (optionally) word-level timing.
- Translation to English: Convert non-English speech into English text for cross-language understanding and summaries.
- Robust usability: Supports model size selection (tiny to large/turbo), language specification for speed, and prompts to improve accuracy on technical or domain-specific terms.
Use case example: Transcribe a multilingual podcast and automatically translate it to English, then save the result as SRT subtitles for sharing or indexing.
Quick Start
Use the whisper skill to transcribe the attached file 'episode.mp3' and translate it to English.