Nemotron 3 Super 120B A12B
NVIDIA · 2026-03-11 · 123.6B parameters
Nemotron 3 Super 120B A12B is NVIDIA's open-weight, post-trained 120B-class text model for reasoning and agentic workloads, released March 11, 2026; the released BF16 checkpoint contains exactly 123,611,012,096 parameters, while NVIDIA reports only rounded active counts (12.7B including embeddings, 12.1B excluding them). It introduces LatentMoE and multi-token prediction to the hybrid Mamba-2/attention architecture, offers switchable thinking, and supports up to 1,048,576 tokens, although the repository configuration and serving examples default to 262,144. Its base was pretrained over a 25-trillion-token horizon with a mixed NVFP4/BF16/MXFP8 recipe, then underwent supervised fine-tuning and reinforcement learning, so the final checkpoint's exact cumulative training-token exposure is undisclosed.