Speech

Transcribe audio to text with Whisper and synthesize speech with 11Labs.

2|2|Updated Feb 12, 2026
One-click install
npx skills add https://github.com/bishwashere/cowCode --skill speech-bishwashere
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Speech
Source: https://github.com/bishwashere/cowCode/tree/main/skills/speech
Command: npx skills add https://github.com/bishwashere/cowCode --skill speech-bishwashere

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Transcribe audio content into text for easy search, note-taking, and archiving, and convert text into natural-sounding speech to respond or narrate content.

Core Features & Use Cases

  • Transcription: convert audio files into text using Whisper for accurate transcripts.
  • Synthesis: generate speech from text using 11Labs with configurable voice options.
  • Real-world scenarios: transcribe meetings or voice notes, summarize content, and reply in spoken form within chats.

Quick Start

To begin, provide an audio file to transcribe or supply text to synthesize; the skill will perform transcription or speech synthesis accordingly.

Frequently Asked Questions about Speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio files to text for meetings and voice notes?

To transcribe audio to text, provide your audio file and the skill uses Whisper to generate accurate transcripts for meetings, voice notes, and archiving.

What is the best way to generate natural-sounding speech from text?

To generate spoken audio from text, supply your text and the skill uses 11Labs synthesis to create natural-sounding speech with configurable voice and language options.

Can I use 11Labs to reply with spoken audio in chat applications?

Yes, you can use 11Labs speech synthesis to generate spoken audio replies within chats, enabling you to narrate content or respond in spoken form alongside transcribing audio.

Does Whisper transcription support configurable language options?

Whisper transcription processes audio into searchable text, and the skill provides configurable language and output options to handle various transcription and synthesis workflows.

When do I need audio processing for accessibility workflows?

You need audio processing for accessibility workflows when converting audio into text for screen readers, or generating spoken audio from text to assist users requiring voice-based content delivery.