openai-whisper

Transcribe audio files into text with timestamps and optional English translation.

Updated May 17, 2026
One-click install
npx skills add https://github.com/tiankong0101-byte/skills-registry --skill openai-whisper-tiankong0101-byte
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: openai-whisper
Source: https://github.com/tiankong0101-byte/skills-registry/tree/main/skills/openai-whisper
Command: npx skills add https://github.com/tiankong0101-byte/skills-registry --skill openai-whisper-tiankong0101-byte

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill turns spoken audio into written text, removing the manual effort of listening, pausing, and transcribing recordings by hand.

Core Features & Use Cases

  • Speech-to-Text Transcription: Convert meeting recordings, voice memos, interviews, and other audio files into readable text.
  • Timestamped Output: Generate transcripts with timestamps for review, editing, or quote extraction.
  • Translation to English: Translate non-English audio into English text when needed.
  • Batch Processing: Handle multiple audio files in one workflow for higher-volume transcription needs.

Quick Start

Ask the skill to transcribe the attached audio file into text with timestamps if needed.

Frequently Asked Questions about openai-whisper

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files into text with timestamps?

To transcribe audio files into text with timestamps, you can use this Skill to convert meeting recordings, voice memos, or interviews into readable, time-stamped written transcripts automatically.

Do I need an OpenAI API key to convert speech to text?

Yes, you need an OpenAI API key and network access to api.openai.com to convert speech to text, as the transcription and translation processes rely on OpenAI's Whisper infrastructure.

Can I translate non-English audio into English text?

Yes, you can translate non-English audio into English text. The Skill processes spoken audio from common formats and outputs the translated English written content directly.

What is the best way to handle batch transcription for multiple audio files?

The best way to handle batch transcription for multiple audio files is using the Skill's batch processing feature, which transcribes higher volumes of meeting recordings or voice memos in one workflow.

Does audio transcription work for common audio formats like meeting recordings?

Yes, audio transcription works for common audio formats. It is specifically designed to process meeting recordings, voice memos, and interviews, turning them into accurate written text.

Are there limitations when processing audio for speech-to-text conversion?

Limitations for speech-to-text conversion include requiring active network access to api.openai.com and a valid OpenAI API key, as the Skill cannot perform offline transcription or batch processing without them.