text-to-speech

Generates customizable speech audio from text input via HeyGen's Starfish TTS API.

602|12|Updated Apr 22, 2026
One-click install
npx skills add https://github.com/video-production-buddy/video-production-buddy --skill text-to-speech-video-production-buddy
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: text-to-speech
Source: https://github.com/video-production-buddy/video-production-buddy/tree/main/.agents/local/skills/text-to-speech
Command: npx skills add https://github.com/video-production-buddy/video-production-buddy --skill text-to-speech-video-production-buddy

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill removes the friction of manual voiceover recording and disconnected text-to-speech tools, letting you generate high-quality, customizable speech audio directly within your video production workflow without needing audio recording equipment or separate software.

Core Features & Use Cases

  • Customizable Speech Generation: Convert text to natural-sounding audio with full control over voice selection, speed, pitch, and pause placement for voiceovers, narration, or podcast content.
  • Voice Discovery: Browse and filter HeyGen Starfish TTS voices by language, gender, and supported features to find the perfect match for your project's tone and audience.
  • Use Case: Quickly generate a professional voiceover for a marketing demo video, or create multilingual audio snippets for global social media content, all without leaving your AI assistant.

Quick Start

Use the text-to-speech skill to generate a 1-minute English narration track for your product demo script using a warm, conversational female voice with 1.1x speed.

Frequently Asked Questions about text-to-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate voiceover audio from a script without manual recording?

You can convert text to natural speech audio instantly using the text-to-speech function, which eliminates the need for manual voiceover recording by generating high-quality narration directly from your script.

How do I find natural-sounding voices for multilingual narration?

You can browse and filter HeyGen Starfish TTS voices by language and gender to find the perfect natural-sounding match for your multilingual narration or podcast content.

Can I adjust the speed and pitch of TTS audio for my video production workflow?

Yes, you can customize speech generation by adjusting speed and pitch, placing pauses, and selecting specific voices directly within your video production workflow without external audio tools.

Do I need external audio tools to create podcast content from text?

No, you do not need external audio tools to create podcast content from text. The Skill generates high-quality speech audio with pause control and timestamped output directly within your existing environment.

How do I get timestamped audio output for synchronizing voiceovers to video?

You get timestamped audio output for synchronizing voiceovers to video by generating speech audio from text input, which provides timestamped output to assist with synchronization tasks.

What are the limitations of using text-to-speech for video production?

The limitation is that voice generation is restricted to the available HeyGen Starfish TTS API voices, meaning you cannot import custom recorded audio or manually edit audio waveforms within the tool.