What problem does it solve?
Converts spoken audio into readable, timestamped text and optionally labels speakers so teams can quickly review meetings, interviews, and recorded conversations without manual transcription.
Core Features & Use Cases
- Fast transcription: Defaults to a lightweight transcribe model for quick text output.
- Speaker diarization: Use a diarize model to produce diarized_json with speaker segments and known-speaker hints.
- Robust CLI workflow: Validates API key and file size, supports chunking for long audio, and writes outputs to organized job directories.
- Use Case: Transcribe a 45-second interview, include speaker labels from reference audio, and export diarized JSON for downstream analysis.
Quick Start
Transcribe the attached meeting audio, produce diarized_json with speaker labels, and save the transcript to output/transcribe/meeting.