voice-audio-engineer

Generate TTS, clone voices, and produce podcasts with ElevenLabs workflows.

181|30|Updated Nov 16, 2025
One-click install
npx skills add https://github.com/erichowens/some_claude_skills --skill voice-audio-engineer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: voice-audio-engineer
Source: https://github.com/erichowens/some_claude_skills/tree/main/.claude/skills/voice-audio-engineer
Command: npx skills add https://github.com/erichowens/some_claude_skills --skill voice-audio-engineer

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Expert in voice synthesis, speech processing, and vocal production using ElevenLabs and professional audio techniques. Focuses on high-quality TTS, voice cloning, dialogue processing, podcast/audiobook production, and voice UI design with robust loudness control and de-essing.

Core Features & Use Cases

  • Text-to-speech generation with brand-appropriate voice options
  • Voice cloning for brand-consistent personas
  • Speech-to-speech transformation and dialogue mixing
  • Podcast/audiobook production workflows (editing, mastering, LUFS)
  • Voice UI integration for conversational agents
  • Loudness normalization and sibilance control (LUFS, de-essing)
  • Transcription and speech-to-text for indexing and analytics

Quick Start

Use voice-audio-engineer to set up TTS, clone a voice, and configure a podcast workflow with LUFS targets.

Frequently Asked Questions about voice-audio-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate natural-sounding speech from text using TTS?

Text-to-speech generation converts written content into spoken audio with brand-appropriate voice selection. ElevenLabs-based TTS produces high-quality output suitable for podcasts, audiobooks, and voice interfaces by applying voice cloning and transformation to match your brand identity.

What's the best way to normalize audio loudness for podcast distribution?

Loudness normalization targets LUFS ranges: -16 to -19 for podcasts and -14 for streaming platforms, with true peak management at -1 dBTP. This Skill applies compression, de-essing, and high-pass filtering at 80-100 Hz to meet technical broadcast standards consistently.

Can I clone a voice and use it across multiple audio projects?

Voice cloning creates brand-consistent personas from source audio, enabling speech-to-speech transformation and dialogue mixing across projects. This approach maintains vocal identity while applying professional processing like sibilance control and background-noise isolation.

How do I produce a podcast with professional audio mastering?

Podcast production workflows combine transcription, speaker diarization, and mastering with loudness targets and de-essing applied. This Skill integrates MCP tool workflows for editing, LUFS compliance, and noise isolation to deliver broadcast-ready audio.

Does voice synthesis work for building conversational voice UIs?

Voice UI integration uses TTS and speech-to-text capabilities to power conversational agents with natural dialogue processing. ElevenLabs models support real-time synthesis and speaker recognition for interactive applications.

What audio processing handles background noise and vocal clarity?

Audio processing combines high-pass filtering, de-essing for sibilance, background-noise isolation, and compression with configurable knee and attack/release settings. These techniques isolate clean vocal tracks and ensure professional clarity across speech content.