transcribe

Convert audio and video files into markdown transcripts with speaker labels.

28|Updated Jan 3, 2026
One-click install
npx skills add https://github.com/ma08/botfiles --skill transcribe-ma08
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: transcribe
Source: https://github.com/ma08/botfiles/tree/main/codex/skills/transcribe
Command: npx skills add https://github.com/ma08/botfiles --skill transcribe-ma08

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

Manually transcribing audio and video files is slow, tedious, and prone to errors, especially for long recordings like meetings, interviews, or media content.

Core Features & Use Cases

  • Flexible Input Support: Transcribe single audio/video files or batch process entire directories of recordings at once.
  • Accuracy Optimization: Add custom domain context and specialized terminology to improve transcription accuracy for industry-specific or team-specific content.
  • Multiple Output Formats: Get both raw API response files and readable markdown transcripts with speaker labels for easy sharing and use.
  • Use Case Example: Transcribe a full day's worth of team standup recordings with your product's custom terminology correctly recognized, all in one batch run.

Quick Start

Use the transcribe skill to convert your audio or video recording into an accurate, readable text transcript.

Frequently Asked Questions about transcribe

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe meeting recordings into text with speaker labels?

Audio transcription works by converting speech from single or multiple audio and video files into accurate text transcripts. It processes recordings in batch and outputs both raw API response files and formatted markdown transcripts with speaker labels.

Can I batch transcribe multiple audio and video files at once?

Yes, batch transcription of multiple audio and video files is supported by processing entire directories of recordings at once. This enables converting large collections of media content into formatted text transcripts with a single run.

How do I improve speech-to-text accuracy for specialized terminology?

You can improve speech-to-text accuracy for specialized terminology by providing custom domain context and specialized vocabulary as inputs. This helps the transcription engine correctly recognize industry-specific terms in your audio or video content.

What is the best way to convert interviews to text transcripts?

The best way to convert interviews to text transcripts is using automated speech-to-text technology that applies custom domain context. This skill outputs both raw API data and readable markdown files with speaker labels for easy sharing and use.

What output formats do I get from video transcription?

Video transcription yields multiple output formats, including both raw API response files and readable markdown transcripts. The markdown transcripts include speaker labels, making the digitized media content easy to read and share.