← Back to Top 10
NVIDIA

Nemotron-3-Ultra

NVIDIA's frontier 550B LatentMoE Mamba-2 hybrid with 1M context, Multi-Token Prediction, and high-stakes reasoning.

Generic Info

  • Publisher: NVIDIA
  • Release Date: June 2026
  • Parameters: 550B Total (55B active per token)
  • Architecture: LatentMoE - Mamba-2 + MoE + Transformer Hybrid with Multi-Token Prediction (MTP)
  • Context Window: 1,000,000 tokens (1M native)
  • License: OpenMDW License v1.1 (Open Weights)
  • Key Capabilities: Linear-Time State Space Attention, High-Stakes RAG, Thinking Mode (`enable_thinking=True`), Complex Enterprise Agent Workflows

Nemotron-3-Ultra represents NVIDIA's groundbreaking architecture combining state-space models (Mamba-2) with sparse LatentMoE and Multi-Token Prediction (MTP). By replacing quadratic attention layers with linear state-space operators in alternating blocks, Nemotron-3-Ultra achieves unprecedented inference speed across 1-million-token contexts while maintaining the highest frontier accuracy in enterprise reasoning, tool use, and multi-document RAG.

Hello World Guide

Run Nemotron-3-Ultra using Hugging Face transformers or NVIDIA TensorRT-LLM.

Python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

messages = [
    {"role": "system", "content": "You are Nemotron-3-Ultra, NVIDIA's frontier enterprise agent model."},
    {"role": "user", "content": "Formulate a zero-hallucination document synthesis pipeline across 500k context tokens."}
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True
)

inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)

Industry Usage

Million-Token Enterprise RAG

Hybrid Mamba-2 architecture allows sustained 1M context evaluation with linear memory scaling and zero quadratic slowdown.

Multi-Token Prediction Throughput

MTP delivers up to 2.5x generation speedups, drastically cutting inference latency in interactive developer environments.

Autonomous Enterprise Agents

Designed for mission-critical orchestration, executing multi-step tool calls, compliance audits, and security validation.