openai-audio

Generate and validate OpenAI speech, transcriptions, captions, translations, and audio conversations.

123|21|Updated Jul 11, 2026
One-click install
npx skills add https://github.com/calesthio/generative-media-skills --skill openai-audio-calesthio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: openai-audio
Source: https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/speech-and-voice/openai-audio
Command: npx skills add https://github.com/calesthio/generative-media-skills --skill openai-audio-calesthio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps transform scripts and recorded audio into production-ready narration, transcripts, captions, translations, and bounded audio conversations while selecting the correct OpenAI API route and preserving review, rights, privacy, and artifact-custody requirements.

Core Features & Use Cases

  • Text-to-Speech Production: Generate directed narration and spoken product content with voice, pacing, pronunciation, format, audition, and approval controls.
  • Speech-to-Text Workflows: Transcribe interviews, tutorials, and recordings with model routing for accuracy, speaker diarization, vocabulary prompts, and word or segment timestamps.
  • Audio API Routing: Choose between request-based speech, transcription, translation, audio-capable chat, and separate realtime voice workflows based on the task and time shape.
  • Production Guardrails: Manage consent, synthetic-voice disclosure, privacy, output formats, versioned artifacts, checksums, manifests, QA, and repair paths.
  • Use Case: Create a polished SaaS demo narration in WAV, generate a compact review copy, then transcribe the approved final audio into timed captions for editorial delivery.

Quick Start

Use the openai-audio skill to generate a warm, pronunciation-accurate WAV narration from the provided script and transcribe the approved final audio for word-level captions.

Frequently Asked Questions about openai-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate text to speech narration with OpenAI audio APIs?

Text to speech production with OpenAI audio APIs involves routing scripts through directed narration controls for voice, pacing, pronunciation, and output formats to create review-ready audio artifacts.

Can I transcribe audio with word-level timestamps and speaker diarization?

Speech-to-text workflows support transcribing recordings with model routing for accuracy, speaker diarization, vocabulary prompts, and word or segment timestamps to produce structured transcripts.

What is the best way to route continuous live voice sessions versus bounded audio tasks?

Bounded audio conversations use request-based speech, transcription, translation, or audio-capable chat APIs, while continuous live voice sessions route to separate realtime APIs based on task time shape.

Does OpenAI audio production handle consent, synthetic-voice disclosure, and privacy?

Production guardrails manage consent, synthetic-voice disclosure, privacy, output formats, versioned artifacts, checksums, manifests, QA, and repair paths to ensure rights and artifact custody.

How do I create timed captions from approved final audio files?

Transcribe approved final audio into timed captions by applying speech-to-text workflows with word-level or segment timestamps, yielding editorial-ready caption files for delivery.

What are the limitations of using OpenAI audio APIs for accessibility workflows?

Limitations include requiring correct endpoint and model selection, supported audio parameters, pronunciation direction, and rights review, while continuous live voice sessions need realtime APIs instead.