What problem does it solve?
This Skill provides a comprehensive suite of tools for voice and audio manipulation, including text-to-speech, speech-to-text, and voice conversion, simplifying complex audio tasks.
Core Features & Use Cases
- Text-to-Speech (TTS): Generate natural-sounding speech from text using various models like ElevenLabs and Kling.
- Speech-to-Text (STT): Transcribe audio files with options for diarization (speaker identification) and timestamps using models like Whisper and Wizper.
- Voice Conversion & Cloning: Transform voices or clone them using models like RVC and ElevenLabs.
- Audio Utilities: Merge audio with video and perform general audio conversions.
- Use Case: You need to add a voiceover to a video in multiple languages, transcribe a meeting with speaker labels, or create a custom AI voice for your brand.
Quick Start
Use the eachlabs-voice-audio skill to convert the text 'Hello, world!' into speech using the elevenlabs-text-to-speech model.