text-to-speech

Generate speech audio from text using HeyGen Starfish TTS.

Updated Apr 2, 2026
One-click install
npx skills add https://github.com/gyalamanch001a/pur-new --skill text-to-speech-gyalamanch001a
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: text-to-speech
Source: https://github.com/gyalamanch001a/pur-new/tree/main/.agents/skills/text-to-speech
Command: npx skills add https://github.com/gyalamanch001a/pur-new --skill text-to-speech-gyalamanch001a

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill removes the friction of producing spoken audio from written text, letting you create narration, voiceovers, and audio previews without recording a human voice.

Core Features & Use Cases

  • Standalone speech generation: Convert scripts, announcements, and product copy into downloadable audio.
  • Voice control: Choose voices and adjust speed, pitch, and locale for more natural results.
  • Production workflows: Generate audio for podcasts, demos, training materials, and caption-synced content using timestamps and break tags.

Quick Start

Ask the skill to turn your text into speech audio with a chosen voice, speed, and pitch, then return the downloadable audio URL.

Frequently Asked Questions about text-to-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate natural-sounding speech audio from text for a voiceover?

To generate natural-sounding speech audio from text, you provide your script along with voice settings like speed, pitch, and locale. The Skill uses HeyGen Starfish TTS to convert the text into a downloadable audio URL for your voiceover.

Do I need a HeyGen API key to convert text to speech?

Yes, you need a HeyGen API key to convert text to speech. The Skill uses HeyGen Starfish TTS and its /v1/audio endpoints to process your text, apply voice settings, and generate the final narration audio.

Can I create caption-synced audio and get word timestamps for my narration?

Yes, you can create caption-synced audio for production workflows. The Skill supports using word timestamps and break tags through the HeyGen /v1/audio endpoints to generate precisely timed narration for podcasts and demos.

What is the best way to produce multilingual narration without recording a human voice?

The best way to produce multilingual narration without a human voice is using this text to speech Skill. It lets you choose different voices and adjust the locale setting, converting your written scripts into natural spoken audio for global audiences.

Can I adjust the speed and pitch of generated TTS audio?

Yes, you can adjust the speed and pitch of generated TTS audio. When requesting speech generation, you specify your desired voice, speed, and pitch parameters alongside your text to control the final audio output's tone and pacing.

Does text to speech work for generating standalone podcast clips and product demos?

Yes, text to speech works for generating standalone podcast clips and product demos. The Skill removes manual narration friction, allowing you to convert scripts and product copy directly into downloadable audio files.