What problem does it solve?
This Skill helps transform scripts and recorded audio into production-ready narration, transcripts, captions, translations, and bounded audio conversations while selecting the correct OpenAI API route and preserving review, rights, privacy, and artifact-custody requirements.
Core Features & Use Cases
- Text-to-Speech Production: Generate directed narration and spoken product content with voice, pacing, pronunciation, format, audition, and approval controls.
- Speech-to-Text Workflows: Transcribe interviews, tutorials, and recordings with model routing for accuracy, speaker diarization, vocabulary prompts, and word or segment timestamps.
- Audio API Routing: Choose between request-based speech, transcription, translation, audio-capable chat, and separate realtime voice workflows based on the task and time shape.
- Production Guardrails: Manage consent, synthetic-voice disclosure, privacy, output formats, versioned artifacts, checksums, manifests, QA, and repair paths.
- Use Case: Create a polished SaaS demo narration in WAV, generate a compact review copy, then transcribe the approved final audio into timed captions for editorial delivery.
Quick Start
Use the openai-audio skill to generate a warm, pronunciation-accurate WAV narration from the provided script and transcribe the approved final audio for word-level captions.