What problem does it solve?
Manual transcription of audio content is time-consuming, costly, and often fails with multilingual speech, noisy audio, or large volumes of content, making it difficult to turn podcasts, meetings, and interviews into usable, searchable text.
Core Features & Use Cases
- Multilingual Speech Recognition: Accurately transcribe audio in 99 languages, including low-resource languages, with support for multiple model sizes to balance speed and accuracy.
- Translation to English: Automatically translate non-English audio recordings into English text for global accessibility.
- Use Case: A content team can use this skill to transcribe dozens of podcast episodes, generate subtitles, and create meeting notes from recorded calls without manual effort.
Quick Start
Use the whisper skill to transcribe the attached audio file 'team-standup.mp3' into English text with word-level timestamps.