Atlas

Benchmarks

← All benchmarks

BLXBench Reasoning category

Code · 2026-04-25

Reasoning category component of the BLXBench public leaderboard. The source describes this category as arithmetic, symbolic steps, and structured problem-solving; scores are category score percentages.

Top models (higher is better)

ModelScore
Opus 4.879.5
GPT-5.576.2
Opus 4.775.6
Claude Fable 575.4
GPT-5.3-Codex74.7
Kimi K2.7 Code73.6
Granite 4.1 8B73.4
DeepSeek-V4-Pro73.2
MiMo-V2.573.2
Grok Build 0.172.3
MiniMax M371.7
MiMo-V2.5-Pro71.5
Qwen3.6-Flash-2026-04-1670.4
Qwen3.7-Max69.9
Ring 2.6 1T69.6
CoBuddy69.3
Mistral Small 469.2
DeepSeek-V4-Flash69.1
Qwen3.7-Plus68.0
Kimi K2.667.8
Mistral Medium 3.567.8
Nemotron 3 Nano 30B A3B67.5
Step 3.7 Flash67.1
Nemotron 3 Super 120B A12B66.5
North Mini Code66.5
Grok 4.363.6
GLM-5.263.2
GLM-5.160.4
MiniMax M2.759.5
Nemotron 3 Nano Omni 30B A3B51.9
Gemini 3.5 Flash50.5
Gemini 3.1 Flash-Lite49.9
Loading Atlas data…