Nemotron 3 Nano 30B A3B
NVIDIA · 2025-12-15 · 31.6B parameters
Nemotron 3 Nano 30B A3B is an open-weight NVIDIA MoE language model, trained from scratch as a unified model for reasoning and non-reasoning tasks; the released BF16 checkpoint contains exactly 31,577,937,344 parameters. Its 52-layer hybrid architecture uses no positional embeddings and combines 23 Mamba-2, 23 MoE, and 6 GQA layers, with each MoE layer activating 6 of 128 routed experts and using shared-expert capacity. After a reported 25T-token main pretraining run plus a reported 121B-token long-context stage, it underwent SFT, RLVR, and RLHF; reasoning can be toggled, and it supports up to 1M tokens although the Hugging Face config defaults to 262,144 due to VRAM requirements.