Atlas

Benchmarks

← All benchmarks

BLXBench Coding category

Code · 2026-05-10

Coding category component of the BLXBench public leaderboard. The source describes this category as implementation-focused coding tasks with structured correctness checks; scores are category score percentages.

Top models (higher is better)

ModelScore
Claude Fable 598.5
Opus 4.798.5
Opus 4.898.5
GPT-5.3-Codex98.5
GPT-5.598.5
Ring 2.6 1T96.8
DeepSeek-V4-Pro96.2
MiMo-V2.5-Pro96.2
Qwen3.7-Max95.8
Qwen3.7-Plus95.8
Grok Build 0.195.4
Kimi K2.695.4
GLM-5.294.6
Gemini 3.1 Flash-Lite94.3
MiMo-V2.593.9
DeepSeek-V4-Flash92.3
MiniMax M392.3
GLM-5.192.0
Grok 4.391.2
Nemotron 3 Super 120B A12B90.0
Qwen3.6-Flash-2026-04-1690.0
Kimi K2.7 Code89.7
Mistral Small 489.3
Mistral Medium 3.588.9
Gemini 3.5 Flash86.2
CoBuddy75.5
Nemotron 3 Nano 30B A3B73.6
Granite 4.1 8B64.8
Nemotron 3 Nano Omni 30B A3B62.8
MiniMax M2.746.0
North Mini Code28.0
Step 3.7 Flash25.7
Loading Atlas data…