What problem does it solve?
Speech transcribe turns recorded audio or video into usable text you can search, quote, and act on, without manually listening and typing.
Core Features & Use Cases
- Provider-agnostic transcription: Uses the native audio_transcribe tool with optional provider selection (auto, OpenAI, Deepgram, or AssemblyAI) for practical reliability.
- Diarization and timestamps: Produces speaker-attributed segments (when diarization is enabled) and supports word/segment timestamps for analysis or review.
- Transcript artifacts and operational visibility: Persists transcript text and segment JSON as workspace artifacts while accounting for duration and cost, enabling controlled sharing.
- Use Case: You need a call recap with speaker-by-speaker notes and word-level timing for compliance review, generated from a single uploaded audio file.
Quick Start
Ask for the transcription by saying: transcribe the attached audio clip with diarization and word timestamps, and include the detected language.