What problem does it solve? Manually transcribing video or audio recordings is slow and error-prone, and producing properly time-coded subtitle files in multiple formats adds even more work. This Skill automates the full pipeline from raw media upload to finished transcript. ## Core Features & Use Cases - Multi-format transcription: Convert any uploaded video or audio file into structured JSON (with segment timings), SRT, and VTT outputs hosted on permanent CDN URLs. - Speaker diarization: Automatically label who is speaking, with the option to disable it for clean single-speaker output. - Language control: Auto-detect the language or set it explicitly to avoid misdetection on short or noisy clips. - Use Case: You have a recorded meeting and need to know who said what. Upload the file, run transcription with diarization enabled, poll the job status, and read the JSON output grouped by speaker with exact time codes. ## Quick Start Ask your agent to transcribe the attached interview video and return the transcript with speaker labels and an SRT subtitle file.