amazon-polly

Select Amazon Polly engines, voices, SSML, and S3 workflows for text-to-speech audio.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill amazon-polly-calesthio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: amazon-polly
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/text-to-speech/amazon-polly
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill amazon-polly-calesthio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps production teams create reliable AWS-hosted narration without guessing which Polly engine, operation, voice, SSML feature, or timing workflow fits the deliverable.

Core Features & Use Cases

  • Engine and Operation Selection: Choose among Standard, Neural, Long-form, and Generative engines with compatible synchronous, streaming, or S3-backed asynchronous workflows.
  • Production Timing and Pronunciation: Create speech marks for captions, word highlighting, viseme-based lip sync, and SSML cues while managing lexicons, multilingual terms, and supported markup.
  • Production Delivery and Governance: Plan audio formats, mastering, quotas, pricing, IAM, S3 custody, privacy, rights, consent, disclosure, and final QA for videos, audiobooks, training, accessibility audio, avatars, and podcasts.
  • Use Case: Create a three-minute product training narration with consistent product-name pronunciation, exact word-level captions, S3 storage, and a documented audio and compliance QA pass.

Quick Start

Ask the Amazon Polly skill to create production-ready narration for your approved script, including engine and voice selection, SSML, pronunciation lexicons, speech-mark timing, S3 storage, rights notes, and final QA.

Frequently Asked Questions about amazon-polly

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate text to speech audio with precise word-level captions for video narration?

Text to speech production for video narration requires selecting a compatible engine and generating speech marks alongside the audio to drive exact word-level captions. This workflow outputs production-ready voice audio synchronized with precise timing metadata.

What is SSML and how does it improve voice narration pronunciation?

SSML is a markup language used to control pronunciation, pacing, and emphasis in voice narration. By authoring SSML cues and applying pronunciation lexicons, you ensure consistent articulation of specific terms like product names across multilingual media outputs.

Can I use asynchronous S3 synthesis for high-volume text to speech generation?

Asynchronous S3 synthesis supports high-volume text to speech generation by offloading processing to AWS storage workflows. You select the appropriate Polly engine, submit the text, and retrieve production-ready audio files directly from S3 custody upon completion.

How do I get viseme timing data for lip sync avatar animations?

Viseme timing data for lip sync avatar animations is retrieved by requesting speech marks during text to speech synthesis. The workflow outputs precise viseme timestamps that map visual mouth shapes to the generated audio track.

Which Amazon Polly engine should I choose for long-form audiobook production?

Long-form audiobook production benefits from the Long-form or Generative Polly engines, which deliver expressive, natural-sounding voice narration. Engine selection depends on balancing the required vocal expressiveness against AWS quotas and pricing constraints.

What are the limitations when using standard text to speech for accessibility audio?

Standard text to speech engines may lack the natural expressiveness needed for engaging accessibility audio, and usage is bound by AWS quotas, IAM permissions, and privacy consent requirements. You must document rights and compliance QA passes to ensure delivery governance.