What problem does it solve?
Whisper solves the problem of converting spoken audio into accurate, readable transcripts and translations, so you can capture meetings, podcasts, and multilingual conversations without manual typing.
Core Features & Use Cases
- Multilingual Speech-to-Text: Transcribe audio in 99 languages with automatic or specified language handling.
- Translation to English: Convert non-English speech into English text for consistent understanding across teams.
- Timestamps & Tuning: Produce segment and word-level timestamps and improve quality using model selection, language hints, and initial prompts.
Use cases include transcribing podcasts and interviews, generating meeting notes, translating multilingual recordings for documentation, and preparing audio for downstream search or analysis (e.g., turning recordings into text for summaries).
Quick Start
Use the whisper skill to transcribe the attached audio file 'audio.mp3' and return the full text.