Atlas

Benchmarks

← All benchmarks

BLXBench Cost category

Code · 2026-05-10

Cost category component of the BLXBench v2 public leaderboard. It scores cost-aware correctness and efficient API spend per successful task; higher percentage scores indicate better category performance.

Top models (higher is better)

ModelScore
Opus 4.892.7
GPT-5.3-Codex92.7
Grok Build 0.192.7
Nemotron 3 Super 120B A12B92.7
Opus 4.791.3
GPT-5.591.3
Granite 4.1 8B90.7
Grok 4.390.7
MiMo-V2.590.7
Mistral Small 490.7
CoBuddy90.0
DeepSeek-V4-Pro90.0
Mistral Medium 3.590.0
Qwen3.6-Flash-2026-04-1690.0
Qwen3.7-Plus90.0
Nemotron 3 Nano 30B A3B89.3
Qwen3.7-Max89.3
GLM-5.188.7
Claude Fable 588.0
Gemini 3.1 Flash-Lite88.0
Kimi K2.688.0
MiniMax M388.0
North Mini Code88.0
DeepSeek-V4-Flash87.3
MiMo-V2.5-Pro87.3
Nemotron 3 Nano Omni 30B A3B85.3
Ring 2.6 1T84.8
GLM-5.284.7
Kimi K2.7 Code82.7
MiniMax M2.778.0
Step 3.7 Flash66.0
Gemini 3.5 Flash56.7
Loading Atlas data…