What problem does it solve?
This Skill helps media-production teams turn speech audio into accurate transcripts, captions, narration, and consented synthetic voice assets while managing provider limits, production requirements, and data-governance risks.
Core Features & Use Cases
- Transcription and Captions: Plan Google Cloud Speech-to-Text workflows for short clips, long-form media, live captions, reviewed transcripts, SRT, and WebVTT.
- Narration and Voice Synthesis: Select and use appropriate Google Cloud Text-to-Speech voice families for narration, dialogue prototypes, accessibility audio, and localization.
- Production and Governance: Coordinate audio preparation, script adaptation, segmentation, QC, pricing and quota checks, regional deployment, retention, IAM, consent, rights, and synthetic-voice disclosure.
- Use Case: Create a reviewed transcript and caption package for a long training video, then adapt and synthesize a localized narration track with documented voice, model, region, and review decisions.
Quick Start
Use the Google Cloud Speech skill to create a production plan for transcribing the attached video, generating reviewed English captions, and preparing a Spanish narration workflow.