Llama 3.1 Nemotron Nano 4B V1.1
NVIDIA · 2025-05-20 · 4.5B parameters
Llama 3.1 Nemotron Nano 4B V1.1 is NVIDIA's open-weight, dense 4B-class reasoning model derived from Llama 3.1 8B through NVIDIA's Minitron compression process; the released checkpoint contains exactly 4,512,746,496 parameters. Released on May 20, 2025, it was post-trained for reasoning, chat, math, code, RAG, and tool calling, supports switchable reasoning through the system prompt, fits on a single RTX GPU, and has a 131,072-token context window.