audio-producer-agent

Automate single-voice audio production with TTS, music, and FFmpeg assembly.

24|2|Updated Dec 30, 2025
One-click install
npx skills add https://github.com/michaelboeding/skills --skill audio-producer-agent
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audio-producer-agent
Source: https://github.com/michaelboeding/skills/tree/main/skills/audio-producer-agent
Command: npx skills add https://github.com/michaelboeding/skills --skill audio-producer-agent

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill automates the production of single-voice audio content, eliminating repetitive manual tasks and speeding up creative workflows.

Core Features & Use Cases

  • Orchestrates text-to-speech, background music, and final audio assembly for consistent, branded audio output.
  • Use cases include audiobooks, voiceovers for videos or presentations, audio ads, jingles, sonic logos, meditation guides, and soundscapes.
  • Real-world example: produce a 60-second voiceover with ambient music and a closing tagline.

Quick Start

Use the audio-producer-agent skill to generate a 60-second VO with background music.

Frequently Asked Questions about audio-producer-agent

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate text-to-speech and music production for voiceovers?

Text-to-speech and music automation for voiceovers combines TTS services like Gemini TTS with music generation and FFmpeg assembly to produce polished audio end-to-end. This Skill orchestrates all three components, eliminating manual layering and delivering consistent, branded output in a single workflow.

Can I use this to generate audiobooks with background music?

Yes. This Skill automates audiobook production by integrating TTS for narration, optional background music layers via Lyria, and FFmpeg-based assembly. It handles single-voice content with consistent pacing and optional ambient soundscapes.

What TTS and music generation services does this support?

This Skill requires Gemini TTS or a compatible TTS service for voice generation and Lyria for background music. Both services require API keys configured before use. FFmpeg handles final audio assembly and mixing.

Do I need FFmpeg installed to produce audio with this Skill?

Yes. FFmpeg is required for audio assembly and mixing. The Skill uses FFmpeg to layer TTS output with optional music tracks and produce the final polished audio file.

What audio formats and use cases does single-voice audio production work for?

Single-voice audio production suits audiobooks, voiceovers, jingles, audio ads, sonic logos, meditation guides, and soundscapes. The Skill delivers end-to-end generation optimized for branded, consistent single-narrator content.

What are the limitations of automating audio production this way?

This Skill focuses on single-voice content with optional music layers. Multi-speaker dialogue, complex sound design beyond ambient music, and real-time audio processing fall outside its scope. TTS and music generation quality depend on API service limits and API key quotas.