google-cloud-speech

Plan Google Cloud speech workflows for transcription, captions, narration, and synthetic voices.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill google-cloud-speech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: google-cloud-speech
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/speech-and-voice/google-cloud-speech
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill google-cloud-speech

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps media-production teams turn speech audio into accurate transcripts, captions, narration, and consented synthetic voice assets while managing provider limits, production requirements, and data-governance risks.

Core Features & Use Cases

  • Transcription and Captions: Plan Google Cloud Speech-to-Text workflows for short clips, long-form media, live captions, reviewed transcripts, SRT, and WebVTT.
  • Narration and Voice Synthesis: Select and use appropriate Google Cloud Text-to-Speech voice families for narration, dialogue prototypes, accessibility audio, and localization.
  • Production and Governance: Coordinate audio preparation, script adaptation, segmentation, QC, pricing and quota checks, regional deployment, retention, IAM, consent, rights, and synthetic-voice disclosure.
  • Use Case: Create a reviewed transcript and caption package for a long training video, then adapt and synthesize a localized narration track with documented voice, model, region, and review decisions.

Quick Start

Use the Google Cloud Speech skill to create a production plan for transcribing the attached video, generating reviewed English captions, and preparing a Spanish narration workflow.

Frequently Asked Questions about google-cloud-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I plan a speech-to-text workflow for long-form video transcription and captioning?

You can plan speech-to-text workflows for long-form media by selecting the appropriate Google Cloud Speech-to-Text models to generate reviewed transcripts, SRT, and WebVTT captions. The workflow requires deliberate selection of API methods, regions, and quotas to manage production limits.

Can I use Google Cloud Text-to-Speech for narrating and localizing media production?

Yes, you can use Google Cloud Text-to-Speech voice families for narration, dialogue prototypes, and accessibility audio. It supports voice synthesis workflows for localization, requiring deliberate selection of voice models, regions, and consent documentation for synthetic voices.

What is the best way to manage synthetic voice rights and data governance for speech workflows?

Managing synthetic voice rights requires coordinating consent, IAM, retention, and synthetic-voice disclosure within your speech workflows. You must document voice, model, region, and review decisions to handle data-governance risks and speech-related rights properly.

Does the Google Cloud Speech API support live event captioning and SRT output?

Yes, the Google Cloud Speech API supports live event captioning and outputs reviewed transcripts in SRT and WebVTT formats. Planning these workflows involves selecting specific Speech-to-Text methods, checking regional quotas, and applying human quality review.

How do I prepare audio and scripts for dubbing and voice synthesis?

Preparing for dubbing and voice synthesis involves coordinating audio preparation, script adaptation, and segmentation before generation. You must execute pricing checks, regional deployment, and human quality review to adapt and synthesize localized narration tracks.