eachlabs-voice-audio

Convert text to speech and transcribe audio with speaker diarization.

28|5|Updated Feb 9, 2026
One-click install
npx skills add https://github.com/eachlabs/skills --skill eachlabs-voice-audio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: eachlabs-voice-audio
Source: https://github.com/eachlabs/skills/tree/main/skills/eachlabs-voice-audio
Command: npx skills add https://github.com/eachlabs/skills --skill eachlabs-voice-audio

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a comprehensive suite of tools for voice and audio manipulation, including text-to-speech, speech-to-text, and voice conversion, simplifying complex audio tasks.

Core Features & Use Cases

  • Text-to-Speech (TTS): Generate natural-sounding speech from text using various models like ElevenLabs and Kling.
  • Speech-to-Text (STT): Transcribe audio files with options for diarization (speaker identification) and timestamps using models like Whisper and Wizper.
  • Voice Conversion & Cloning: Transform voices or clone them using models like RVC and ElevenLabs.
  • Audio Utilities: Merge audio with video and perform general audio conversions.
  • Use Case: You need to add a voiceover to a video in multiple languages, transcribe a meeting with speaker labels, or create a custom AI voice for your brand.

Quick Start

Use the eachlabs-voice-audio skill to convert the text 'Hello, world!' into speech using the elevenlabs-text-to-speech model.

Frequently Asked Questions about eachlabs-voice-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio with speaker diarization and timestamps?

To transcribe audio with speaker labels, speech-to-text models like Whisper and Wizper process audio files to generate text output that includes speaker diarization and timestamps.

Can I generate natural speech from text using ElevenLabs?

Text-to-speech generation using ElevenLabs converts written text into natural-sounding speech, supporting multiple languages for voiceovers and diverse audio generation needs.

What's the best way to convert or clone a custom voice?

Voice conversion and voice cloning transform or replicate voices using models like RVC and ElevenLabs, allowing you to create a custom AI voice for your brand or project.

Do I need an API key to run text-to-speech and voice conversion predictions?

An API key is required for authentication to run text-to-speech and voice conversion predictions, integrating directly with the EachLabs API for execution.

How do I merge generated audio with an existing video file?

Audio utilities within the voice and audio processing suite allow you to merge generated audio tracks with existing video files and perform general audio format conversions.