NVIDIA Nemotron 3 Ultra 550B-A55B
Release in the NVIDIA Nemotron family · version 3 Ultra (550B-A55B), v1.0
Nemotron 3 Ultra is an NVIDIA language model with 550B total and 55B active parameters, built on a hybrid Mamba-2 and attention architecture with latent mixture-of-experts layers and multi-token prediction. Reasoning can be switched on or off through the chat template. NVIDIA released BF16 and NVFP4 checkpoints and a separate base checkpoint on Hugging Face.1
- Model hub: Model card (BF16) (external site: huggingface.co)
- License: OpenMDW License Agreement, version 1.1 (external site: raw.githubusercontent.com)
- Documentation: Nemotron 3 Ultra training recipe (external site: github.com)
Availability and license
Overall availability
Weights are downloadable from Hugging Face under the OpenMDW-1.1 license; the repository metadata showed no access gate when checked.12
Availability is separate from permission: read the license before using or redistributing.
OpenMDW License Agreement, version 1.1 (external site: raw.githubusercontent.com)13
Apache License 2.0 (Nemotron Developer Repository training recipes) (external site: raw.githubusercontent.com)5
OpenMDW-1.1 grants free permission to use, modify, and share the model materials without restriction, including under copyright, patent, database, and trade secret rights. Redistributors must include a copy of the agreement and keep applicable origin notices. All rights end for anyone who sues claiming the materials infringe a patent or copyright, unless responding to a suit brought against them first. It imposes no conditions on outputs, disclaims warranties, and leaves clearing third-party rights to the user. The SPDX list (version 3.29.0) includes OpenMDW-1.0 but not version 1.1, so no SPDX ID is recorded.34
Model-disclosure tier
The model parameters for this release can be downloaded by the public. License terms may still restrict use, redistribution, or commercial use.
The weights are under a license that is not on the rubric's OSI-approved list. Read its terms before use.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | BF16 and NVFP4 post-trained checkpoints and a BF16 base checkpoint are published on Hugging Face.12 |
| Inference codeIs code for running the model published? | Public | The model card gives deployment instructions for vLLM, SGLang, and TensorRT-LLM.1 |
| Training codeIs the code used to train the model published? | Public | The Apache-2.0 Nemotron Developer Repository provides pretraining and SFT recipes built on Megatron-Bridge and a representative RL/distillation pass using NeMo RL.65 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Partial | The model card lists the pre- and post-training datasets. Major portions are released in Hugging Face collections, some requiring access approval, while several third-party and NVIDIA datasets are listed as private and not publicly accessible.17 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Partial | NVIDIA's recipe covers pretraining and SFT, but states that the full multi-iteration RL and on-policy distillation pipeline is not reproduced because the intermediate teacher checkpoints it depends on have not been released. It also omits the 1M-token long-context pretraining phase because that data is not open-source.6 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Public | The model card publishes benchmark results and points to the NeMo Evaluator SDK for replication; it says some benchmarks used internal scaffolding not yet open-sourced.1 |
What it is useful for
The model card lists agentic workflows, long-context analysis, tool use, multilingual reasoning, and retrieval-augmented generation, and states the model is ready for commercial and non-commercial use.1
Run and use notes
- The model card documents serving with vLLM, SGLang, and TensorRT-LLM, and lists NVIDIA Hopper (H100, H200), Grace Blackwell (GB200, GB300), and Blackwell (B200, B300) GPUs as test hardware.1
- For the BF16 checkpoint, the model card states a minimum of 8x GB200/B200/GB300/B300, 16x H100, or 8x H200 GPUs, and points to the separate NVFP4 checkpoint for a smaller footprint.1
Organization context
Provenance and derivatives
Post-trained by NVIDIA (supervised fine-tuning, reinforcement learning, and multi-domain on-policy distillation) from its own pre-trained Nemotron 3 Ultra base checkpoint. The model card says post-training synthetic data was generated with teacher models, including third-party open models such as GPT-OSS-120B.1
- Derived from: NVIDIA-Nemotron-3-Ultra-550B-A55B-Base-BF16 (external site: huggingface.co) — NVIDIA's pre-trained base checkpoint for this release.
Other releases in the NVIDIA Nemotron family
- NVIDIA Nemotron 3 Super 120B-A12BModel-disclosure tier (USASI rubric v0.1): Open-stack
- NVIDIA Nemotron 3.5 Lightning 30B-A3BModel-disclosure tier (USASI rubric v0.1): Open-stack
U.S. eligibility
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.