Transcription Automation

Convert audio and video recordings into transcripts, subtitles, and meeting notes.

1|Updated May 18, 2026
One-click install
npx skills add https://github.com/hmzainjamil/claude-office-skills --skill transcription-automation-hmzainjamil
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Transcription Automation
Source: https://github.com/hmzainjamil/claude-office-skills/tree/main/transcription-automation
Command: npx skills add https://github.com/hmzainjamil/claude-office-skills --skill transcription-automation-hmzainjamil

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill removes the manual burden of converting meetings, interviews, lectures, podcasts, and videos into accurate transcripts, subtitles, and notes.

Core Features & Use Cases

  • Speech-to-Text Transcription: Convert audio and video into readable text with timestamps and confidence details.
  • Speaker Diarization: Identify and label different speakers in meetings and conversations for clearer records.
  • Subtitle Generation: Produce SRT and VTT captions for videos, webinars, and published media.
  • Meeting and Content Processing: Generate summaries, action items, searchable archives, show notes, and multilingual outputs.
  • Use Case: A team records a client meeting and uses this Skill to create a transcript, assign speakers, extract action items, and publish subtitles for the recording.

Quick Start

Transcribe the attached meeting recording into a speaker-labeled transcript with timestamps, a concise summary, and action items.

Frequently Asked Questions about Transcription Automation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio and video recordings into text with timestamps?

Transcribe audio and video recordings into readable text by applying speech-to-text conversion with timestamps and confidence details. This process transforms meetings, podcasts, and lectures into structured, searchable text outputs.

How do I generate SRT and VTT subtitle files for a video?

Generate SRT and VTT subtitle files by applying timestamped subtitle formatting to your media. This produces accurate captions for videos, webinars, and published media, exporting structured text outputs for accessibility.

Can I identify and label different speakers in meeting transcripts?

Identify and label different speakers in meeting transcripts using speaker diarization. This technique distinguishes individual speakers in conversations, producing clearer records and accurately structured meeting notes.

Does this transcription approach support multilingual media processing?

This transcription approach supports multilingual media processing by applying speech-to-text conversion and translation to spoken audio. It handles diverse language inputs to generate multilingual outputs, summaries, and searchable text archives.

What is the best way to extract action items and summaries from recorded meetings?

Extract action items and summaries from recorded meetings by applying transcription, cleanup, and summarization to the audio. This converts raw spoken content into concise summaries, structured action items, and searchable archives.

What file formats can I export my transcripts and subtitles in?

Export transcripts and subtitles in txt, srt, vtt, or json formats. These structured export options allow you to integrate transcribed text outputs, meeting notes, and generated captions seamlessly into various external platforms.