What problem does it solve?
This Skill helps agents design, deploy, and validate reliable speech workflows with NVIDIA Speech NIM instead of treating ASR, TTS, translation, and voice cloning as a single undifferentiated service.
Core Features & Use Cases
- Speech Pipeline Design: Compose standalone ASR, TTS, and NMT services for transcription, captioning, localization, voice agents, and speech-to-speech translation.
- Deployment and Integration: Select models, profiles, APIs, GPU configurations, runtime requirements, and self-hosted or hosted deployment paths.
- Production Guardrails: Apply consent, voice rights, privacy, licensing, observability, audio custody, latency, accuracy, loudness, and end-to-end QA checks.
- Use Case: Create a multilingual training-video localization workflow that transcribes source audio, protects product terminology during translation, synthesizes approved narration, and records deployment and review metadata.
Quick Start
Use the NVIDIA Speech NIM skill to design a production-ready workflow for the requested speech, voice, translation, or captioning task with API, deployment, rights, and QA recommendations.