Llama 3.3 Nemotron Super 49B V1.5
NVIDIA · 2025-07-25 · 49.9B parameters
Llama 3.3 Nemotron Super 49B V1.5 is NVIDIA's upgraded open-weight reasoning and chat model derived from Meta's Llama 3.3 70B Instruct, released July 25, 2025; its BF16 checkpoint contains exactly 49,867,145,216 parameters. NAS and block-wise distillation produce a nonuniform 80-layer architecture whose 49 attention-bearing blocks use GQA while 31 skip attention; NVIDIA's launch post reports single-H100 deployment. It supports switchable reasoning, RAG, and tool calling with a 131,072-token context window.