Nemotron 3 Ultra
Nemotron 3 Ultra is an open-weight language model from NVIDIA, released on June 4, 2026. It is a mixture-of-experts model with 550B total parameters, of which 55B are active for each token. NVIDIA lists a context window of up to 1M tokens. The model takes text input.
- Developer
- NVIDIA
- Released
- June 4, 2026
- Total parameters
- 550B
- Active parameters per token
- 55B
- Architecture
- MoE
- Context window
- 1M
- Input
- Text
- License
- OpenMDW License Agreement v1.1
- Commercial use
- Yes
- Official weights, GB
- 1,121 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under the OpenMDW License Agreement v1.1, which allows commercial use.
The official weights on Hugging Face take about 1,121 GB in BF16. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 550B
- Active parameters, billions
- 55B
- Context window, tokens
- 1M
- License type
- Custom license
- License conditions
- OpenMDW 1.1: redistributions must include the agreement and origin notices, and the rights end for anyone who sues claiming the materials infringe.
- Weights precision
- BF16
- Official quantized or alternative versions
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
- License text
- raw.githubusercontent.com
- Sources
Page URL Hugging Face model card https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/raw/main/README.md License text on GitHub https://raw.githubusercontent.com/OpenMDW/OpenMDW/refs/heads/main/1.1/LICENSE.OpenMDW-1.1 Hugging Face weights index https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/raw/main/model.safetensors.index.json