Llama 3.1 Nemotron Ultra 253B V1
NVIDIA · 2025-04-07 · 253.4B parameters
Llama 3.1 Nemotron Ultra 253B V1 is NVIDIA's open-weight, dense decoder-only reasoning model derived from Meta's Llama 3.1 405B Instruct through NAS, attention removal, variable FFNs, and FFN fusion; the released checkpoint contains exactly 253,401,268,224 parameters. Released on April 7, 2025, it accepts and emits text, supports a 131,072-token context, and uses a system prompt to switch between reasoning and non-reasoning modes for chat, RAG, tool calling, and agentic workloads.