Nemotron 3 Nano 4B
NVIDIA · 2026-03-16 · 4B parameters
Nemotron 3 Nano 4B is NVIDIA's open-weight, text-only small language model compressed from NVIDIA-Nemotron-Nano-9B-v2 with the Nemotron Elastic framework and released as a BF16 checkpoint for both reasoning and non-reasoning tasks. The checkpoint contains exactly 3,973,556,832 parameters, and its 42-layer hybrid architecture comprises 21 Mamba-2, four grouped-query-attention, and 17 MLP layers. Reasoning is switchable through the chat template, and the released configuration supports a 262,144-token context window.