nemotron-ultra

Retrieve NVIDIA Nemotron 3 Ultra architecture, training, evaluation, and release details from source files.

Updated Jul 30, 2026
One-click install
npx skills add https://github.com/Lhhiep-maxcode/Nemotron --skill nemotron-ultra
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nemotron-ultra
Source: https://github.com/Lhhiep-maxcode/Nemotron/tree/main/skills/nemotron-ultra
Command: npx skills add https://github.com/Lhhiep-maxcode/Nemotron --skill nemotron-ultra

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps you quickly answer technical questions about NVIDIA Nemotron 3 Ultra without wading through long reports, scattered docs, or release notes. It is designed for accurate model understanding, evaluation lookup, and deployment-oriented fact finding.

Core Features & Use Cases

  • Model identity and release status: Clarifies the 550B/55B architecture, staged availability, checkpoint variants, and intended use.
  • Architecture and training pipeline: Explains the hybrid Mamba-Attention MoE design, NVFP4 pretraining, long-context extension, SFT, RLVR, MOPD, and MTP boosting.
  • Evaluation and inference facts: Surfaces benchmark results, throughput claims, quantization details, serving regimes, and safety-related notes.
  • Use case: A researcher can ask for the exact Ultra post-training sequence or the best file for a specific benchmark, and get a concise, source-backed answer.

Quick Start

Ask for a concise, source-backed explanation of Nemotron 3 Ultra’s architecture, training pipeline, release status, or benchmark results.

Frequently Asked Questions about nemotron-ultra

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is the architecture of NVIDIA Nemotron 3 Ultra?

NVIDIA Nemotron 3 Ultra features a hybrid Mamba-Attention MoE design with 550B total and 55B active parameters, utilizing NVFP4 pretraining and long-context extension for efficient inference.

How do I find benchmark results for Nemotron 3 Ultra?

To find benchmark results for Nemotron 3 Ultra, query the skill to retrieve specific evaluation facts, throughput claims, and source-backed benchmark data from its internal reference files.

Does Nemotron 3 Ultra support NVFP4 quantization for inference?

Yes, Nemotron 3 Ultra supports NVFP4 quantization, utilizing NVFP4 pretraining and offering distinct NVFP4 checkpoint variants optimized for specific serving regimes and inference throughput.

What is the difference between base and post-trained Nemotron 3 Ultra checkpoints?

The difference between base and post-trained Nemotron 3 Ultra checkpoints lies in the post-training pipeline, which applies SFT, RLVR, MOPD, and MTP boosting to the base BF16 model.

How do I trace the Nemotron 3 Ultra post-training stages?

Trace the Nemotron 3 Ultra post-training stages by retrieving the exact sequence of SFT, RLVR, MOPD, and MTP boosting applied to the base model, distinguishing between BF16 and NVFP4 variants.

Are there specific limitations when analyzing Nemotron 3 Ultra deployment facts?

When analyzing Nemotron 3 Ultra deployment facts, precise source attribution is required to distinguish between base, BF16 post-trained, NVFP4, and GenRM serving claims without mixing checkpoint contexts.