Llama 3.3 Nemotron Super 49B V1
NVIDIA · 2025-03-18 · 49.9B parameters
Llama 3.3 Nemotron Super 49B V1 is NVIDIA's open-weight reasoning model derived from Meta's Llama 3.3 70B Instruct and released on March 18, 2025; the released BF16 checkpoint contains exactly 49,867,145,216 parameters. Its 80-block decoder uses NAS-selected variable FFN widths, skips attention in 31 blocks, and uses grouped-query attention in the other 49. It was post-trained for switchable reasoning and non-reasoning chat, RAG, and tool calling, and supports a 131,072-token context window.