speech-to-text

Convert spoken audio into searchable transcripts with timestamps and speaker labels.

4|Updated Mar 10, 2026
One-click install
npx skills add https://github.com/JigsawStack/interfaze-skills --skill speech-to-text-jigsawstack
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-to-text
Source: https://github.com/JigsawStack/interfaze-skills/tree/main/skills/speech-to-text
Command: npx skills add https://github.com/JigsawStack/interfaze-skills --skill speech-to-text-jigsawstack

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Speech-to-text removes the manual effort of listening to recordings and rewriting spoken content, especially for meetings, interviews, and voice notes.

Core Features & Use Cases

  • Accurate transcription from audio: Transcribe audio files, voice notes, podcasts, and meeting recordings into readable text.
  • Speaker diarization and timestamps: Add speaker-labeled chunks and timestamps to make transcripts usable for review and indexing.
  • Language detection and translation: Identify the spoken language and translate the transcript when needed.

Quick Start

Upload or link your audio file and ask for a transcript with speaker labels and timestamps.

Frequently Asked Questions about speech-to-text

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files with speaker labels and timestamps?

To transcribe audio with speaker labels and timestamps, provide an audio file or URL to generate a searchable transcript. The output includes structured text with speaker diarization and segment-level timing for easy indexing.

Can I automatically detect the spoken language and translate the transcript?

Language detection and translation are supported during audio transcription. The process identifies the spoken language and can output translated text with corresponding language codes alongside the original transcript.

What is the best way to transcribe meeting recordings into structured text?

Transcribing meeting recordings into structured text requires uploading the audio file or linking a URL. The output provides full text with timestamped chunks and optional speaker diarization to improve review and usability.

Does speaker diarization work for podcasts and voice notes?

Speaker diarization works for podcasts, voice notes, and interviews by adding speaker-labeled chunks to the transcript. This segments the audio by speaker, making the content searchable and easier to review.

How do I get structured schema outputs from a podcast audio transcription?

Structured schema outputs from podcast audio transcription are generated by processing the audio input via file or URL. The result is a full text transcript with timestamped chunks and optional translated text with language codes.