Atlas

Benchmarks

← All benchmarks

BLXBench Debugging category

Code · 2026-04-25

Debugging category component of the BLXBench public leaderboard. The source describes this category as bug fixes, edge conditions, and minimal patch accuracy; scores are category score percentages.

Top models (higher is better)

ModelScore
Qwen3.7-Max80.2
Kimi K2.679.9
GPT-5.3-Codex79.6
GPT-5.577.0
MiMo-V2.576.5
DeepSeek-V4-Pro76.2
GLM-5.175.8
Grok Build 0.174.7
CoBuddy73.9
MiMo-V2.5-Pro73.8
North Mini Code72.9
DeepSeek-V4-Flash71.9
Grok 4.371.2
Nemotron 3 Super 120B A12B71.2
Qwen3.7-Plus71.0
MiniMax M369.9
GLM-5.269.7
Mistral Medium 3.569.6
Qwen3.6-Flash-2026-04-1669.0
Gemini 3.1 Flash-Lite68.1
Mistral Small 467.9
Nemotron 3 Nano 30B A3B67.9
Granite 4.1 8B67.2
Step 3.7 Flash62.4
Nemotron 3 Nano Omni 30B A3B61.9
Claude Fable 558.2
Kimi K2.7 Code53.1
Ring 2.6 1T50.8
MiniMax M2.745.5
Opus 4.745.1
Opus 4.841.3
Gemini 3.5 Flash25.8
Loading Atlas data…