nemotron-voice-agent-deploy

Deploy Nemotron Voice Agent on Workstation, Jetson Thor, or Cloud NIMs with WebRTC or WebSocket transport.

2.8k|332|Updated Feb 25, 2026
One-click install
npx skills add https://github.com/NVIDIA/skills --skill nemotron-voice-agent-deploy
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemotron-voice-agent-deploy
Source: https://github.com/NVIDIA/skills/tree/main/skills/nemotron-voice-agent/nemotron-voice-agent-deploy
Command: npx skills add https://github.com/NVIDIA/skills --skill nemotron-voice-agent-deploy

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Deploy Nemotron Voice Agent across Workstation (x86), Jetson Thor, or Cloud NIMs to enable real-time speech-to-speech using NVIDIA ASR, TTS, and LLM with WebRTC or WebSocket transport.

Core Features & Use Cases

  • Deployment across multiple platforms (Workstation, Jetson Thor, Cloud NIMs) with WebRTC or WebSocket transport.
  • Integrated ASR, TTS and LLM for live conversational capability across devices.
  • Use Case: Enterprises consistently provision edge devices for hands-free voice assistants in customer support or automation workflows.

Quick Start

Deploy the Nemotron Voice Agent by following the workspace, Jetson Thor, or cloud deployment steps in the nemotron-voice-agent-deploy directory.

Frequently Asked Questions about nemotron-voice-agent-deploy

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I deploy a real-time voice agent using NVIDIA ASR, TTS, and LLM?

To deploy a real-time voice agent, you configure NVIDIA ASR, TTS, and LLM services with WebRTC or WebSocket transport using Docker. The deployment process requires an NVIDIA API key and environment setup to enable live speech-to-speech capabilities.

Can I run the Nemotron Voice Agent on Jetson Thor or do I need an x86 workstation?

You can run the Nemotron Voice Agent on Jetson Thor, an x86 workstation, or Cloud NIMs. The deployment process includes automatic hardware detection and platform selection to configure the environment accordingly across supported devices.

What is needed to set up real-time speech-to-speech with WebRTC or WebSocket?

Setting up real-time speech-to-speech requires Docker, an NVIDIA API key, and proper environment configuration. The deployment guide provides steps to configure NVIDIA ASR, TTS, and LLM with your chosen WebRTC or WebSocket transport protocol.

Does the Nemotron Voice Agent support multilingual live conversations?

Yes, the Nemotron Voice Agent supports optional multilingual capabilities for live conversations. During environment setup, you can configure multilingual support and select specific models to enable speech-to-speech interactions in various languages.

What is the best way to provision edge devices for hands-free voice assistants?

The best way to provision edge devices for hands-free voice assistants is using Docker to deploy the Nemotron Voice Agent on platforms like Jetson Thor. This enables integrated ASR, TTS, and LLM for real-time conversational automation workflows.