K2 Think V2
LLM360 · 2026-01-27 · 72.6B parameters
K2 Think V2 is LLM360's 70B open-weight general reasoning model, built on K2-V2-Instruct and post-trained with two stages of reinforcement learning with verifiable rewards (RLVR) using GRPO. Its chat template defaults to high reasoning effort, while low and medium settings remain available but were not evaluated in the official release. The checkpoint supports up to 262,144 tokens through 2× YaRN scaling from a 131,072-token serving context, and official evaluations report strong results on AIME 2025, HMMT 2025, and GPQA-Diamond.