0
MODEL SIGNAL · NVIDIA

Nemotron-3-Ultra-550B-A55B-NVFP4

A frontier-scale 550-billion parameter hybrid model from NVIDIA featuring a 1-million token context window and NVFP4 precision.

CATEGORYGeneral
CONTEXT1000000
RELEASEDJune 4, 2026
Key Features
  • 550 billion total parameters with 55 billion active parameters per token
  • Hybrid LatentMoE Mamba-2 + MoE + Attention architecture with Multi-Token Prediction (MTP)
  • Up to 1,000,000 token context length
  • Pretrained and quantized NVFP4 precision checkpoint

Provider announcement →

Read the Model Signal report →

MODEL SIGNAL

Nemotron-3-Ultra-550B-A55B-NVFP4

NVIDIA’s 550-billion parameter hybrid model combines Mamba-2, MoE, and NVFP4 precision to push frontier-scale reasoning efficiency.

Bottom line

NVIDIA has released Nemotron-3-Ultra-550B-A55B-NVFP4, a frontier-scale reasoning model that balances massive scale with high sparsity. Confirmed primary specifications detail a 550-billion total parameter count utilizing only 55 billion active parameters per token. Officially released on June 4, 2026, the model is built on a Hybrid LatentMoE Mamba-2 + MoE + Attention architecture, features Multi-Token Prediction (MTP), and natively supports a 1,000,000-token context window as a pretrained and quantized NVFP4 precision checkpoint.

Signal

The clearest signal is NVIDIA's structural advancement in mixed-architecture deployment. By integrating Mamba-2 with traditional Mixture of Experts and Attention mechanisms into a LatentMoE hybrid, the model establishes a distinct path designed to optimize the compute-to-reasoning ratio over ultra-long contexts. The native NVFP4 checkpoint represents a strong directional indicator for the industry: packaging massive 500B+ class models into highly efficient footprints. The operator read is a deliberate focus on serving efficiency for massive 1-million-token workloads without sacrificing the logical depth of a frontier parameter base.

Noise

While moving Hugging Face routing telemetry indicates immediate ecosystem availability and visibility, any assumptions regarding first-token latency, throughput limits, or exact operational costs derived from these early snapshots should be treated as noise. These metrics are fluctuating launch-window data points and do not represent authoritative specs or validated local deployment baselines.

Where it fits

With its 1-million-token context ceiling and high-sparsity activation profile, this model is built for heavy-duty, long-running agentic reasoning applications where high token volume intersects with complex logic. The emerging pattern suggests Nemotron-3-Ultra fits best in sophisticated enterprise environments already equipped for advanced NVFP4 quantization, targeting massive document ingestion workloads that would typically overwhelm smaller, standard MoE deployments.

Model Signal · Signal + Noise · Isaiah Steinfeld