Llama 3.1 Tülu 3 405B
Allen Institute for AI (Ai2) · 2025-01-30 · 405.9B parameters
Llama 3.1 Tülu 3 405B is Ai2's final RLVR checkpoint in its 405B Tülu 3 post-training sequence, built from Meta's Llama 3.1 405B base and publicly launched on January 30, 2025. Ai2 applied SFT, DPO, and then RLVR on the MATH training split. The repository retains Llama 3.1's 131,072-token configured limit, although Ai2 notes that its Tülu 3 post-training data were relatively short, with most samples under 2,048 tokens and no long multi-turn data.