voicemode

Manage local speech-to-text and text-to-speech services for voice conversations.

1.3k|183|Updated Jun 8, 2025
One-click install
npx skills add https://github.com/mbailey/voicemode --skill voicemode
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: voicemode
Source: https://github.com/mbailey/voicemode/tree/main/.
Command: npx skills add https://github.com/mbailey/voicemode --skill voicemode

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires uv, ffmpeg, openai, sounddevice, webrtcvad, psutil, pyyaml, simpleaudio, gcc, portaudio, libasound2-dev, libportaudio2, pulseaudio, node, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill eliminates the tedious manual input and reading required for AI interactions, allowing you to communicate naturally through speech. It automates the setup and management of voice services, saving time and reducing the complexity of integrating AI into your workflow.

Core Features & Use Cases

  • Natural Voice Conversations: Engage in fluid, real-time spoken dialogues with AI assistants like Claude Code, making interactions intuitive and fast.
  • Voice Service Management: Easily start, stop, and monitor local Speech-to-Text (STT) and Text-to-Speech (TTS) services (e.g., Whisper.cpp, Kokoro) directly through AI commands.
  • Automated Configuration: Simplifies setup by managing voice preferences, audio formats, and provider selection, ensuring optimal performance without manual tweaking.
  • Use Case: Instead of typing out complex coding queries or reading lengthy AI responses, simply say "Claude, find all Python files modified last week and summarize their changes," and hear the response, allowing you to focus on coding.

Quick Start

Start a voice conversation with Claude. claude converse