media-transcription-report

Convert audio and video files into timestamped Markdown, TeX, and PDF reports.

1|Updated Jun 3, 2026
One-click install
npx skills add https://github.com/lachlanchen/LazySkills --skill media-transcription-report
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media-transcription-report
Source: https://github.com/lachlanchen/LazySkills/tree/main/skills/media-transcription-report
Command: npx skills add https://github.com/lachlanchen/LazySkills --skill media-transcription-report

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires whisper, whisperx, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill automates the process of transcribing audio or video into Markdown, TeX, and PDF reports, providing a neutral third-person perspective suitable for documentation and analysis.

Core Features & Use Cases

  • Transcription: Converts audio and video files into timestamped Markdown transcripts.
  • WhisperX Diarization: Offers speaker diarization for more detailed transcripts.
  • Report Generation: Creates Markdown, TeX, and PDF reports from transcripts and chat notes.
  • Use Case: Ideal for researchers or content creators who need to document audio/visual content in a structured format.

Quick Start

Use the media-transcription-report skill to transcribe the latest audio file and generate a neutral third-person report in Markdown format.

Frequently Asked Questions about media-transcription-report

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate timestamped transcripts from audio and video files?

Timestamped transcripts are generated by converting audio and video files into Markdown, TeX, and PDF reports using Whisper or WhisperX for transcription. The Skill processes your media input and outputs structured documents with precise timestamps.

Can I use WhisperX for speaker diarization in audio transcription?

Yes, WhisperX supports speaker diarization for audio transcription when you provide optional WhisperX credentials. This feature identifies individual speakers within the generated timestamped transcripts, adding detailed speaker attribution to your Markdown and TeX reports.

What is the best way to convert video transcription output into a TeX report?

The best way to convert video transcription output into a TeX report is through automated generation. The Skill directly transforms transcribed media content into structured TeX documents alongside Markdown and PDF formats from a single processing run.

Do I need WhisperX credentials to transcribe audio into Markdown?

You do not need WhisperX credentials to transcribe audio into Markdown, as standard Whisper transcription is supported without them. Credentials are only required optionally for utilizing the advanced speaker diarization features provided by WhisperX.

Does this audio transcription tool support PDF report generation?

Yes, this audio transcription tool supports PDF report generation alongside Markdown and TeX outputs. It automatically creates neutral, third-person perspective PDF documents from your transcribed media files.

When should I use speaker diarization for video transcription reports?

You should use speaker diarization for video transcription reports when you need to identify and distinguish individual speakers in multi-person conversations. This requires WhisperX credentials and provides detailed speaker attribution for your documentation.