transcribe-refiner

Clean raw caption files into readable transcripts with timestamps removed.

1|Updated Dec 27, 2025
One-click install
npx skills add https://github.com/praxstack/skills-and-personas --skill transcribe-refiner
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: transcribe-refiner
Source: https://github.com/praxstack/skills-and-personas/tree/main/skills/transcribe-refiner
Command: npx skills add https://github.com/praxstack/skills-and-personas --skill transcribe-refiner

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Clean and reconstruct raw auto-generated captions (Zoom, YouTube, Teams, Google Meet, Otter.ai, etc.) into readable, coherent transcripts. Use when the user provides raw caption files (.txt, .vtt, .srt), meeting transcripts with timestamps and speaker tags, or asks to clean up/refine a transcript. Handles: timestamp removal, speaker tag normalization, filler word removal, broken sentence reconstruction, transcription error correction, paragraph formation. Preserves every piece of substantive content while removing noise.

Core Features & Use Cases

  • Smart cleanup of noisy captions into coherent transcripts.
  • Speaker normalization and sectioning for Q&A or lectures.
  • Preservation of all substantive content with noise removal and structure.

Quick Start

Use this skill to clean up a raw caption file (txt, vtt, srt) into a clean, readable transcript with timestamps removed and speakers normalized.

Frequently Asked Questions about transcribe-refiner

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I clean up raw auto-generated captions from Zoom or YouTube into a readable transcript?

To clean up raw auto-generated captions, the transcribe-refiner processes files by removing timestamps, normalizing speaker tags, eliminating filler words, and reconstructing broken sentences. It handles .txt, .vtt, and .srt formats to produce coherent, readable transcripts with zero substantive content loss.

Can I normalize speaker tags and remove timestamps from Otter.ai or Google Meet transcripts?

Yes, you can normalize speaker tags and remove timestamps from Otter.ai or Google Meet transcripts. The tool reconstructs fragmented sentences and corrects transcription errors while preserving all substantive content, making it ideal for educational lectures and corporate meetings.

What is the best way to convert messy VTT or SRT caption files into coherent text?

The best way to convert messy VTT or SRT caption files into coherent text is using a dedicated caption refinement process. This involves stripping timestamp noise, merging broken sentences, and applying paragraph formation to structure the raw auto-generated captions cleanly.

Does transcript cleanup preserve the original content when removing filler words and transcription errors?

Transcript cleanup preserves the original content completely when removing filler words and transcription errors. The refinement process guarantees zero content loss by targeting only noise—such as redundant speaker tags and timestamps—while keeping all substantive spoken material intact.

Are there limitations when reconstructing broken sentences in raw meeting transcripts?

There are no specific limitations mentioned for reconstructing broken sentences in raw meeting transcripts. The process is designed to handle noisy captions from platforms like Teams and YouTube, effectively correcting transcription errors and forming coherent paragraphs without losing substantive content.