Atlas

Benchmarks

← All benchmarks

SWE-bench Verified (Vals 1–4 hour split)

Code · 2026-06-17

Vals difficulty split containing SWE-bench Verified tasks estimated to take 1 to 4 hours for humans. It preserves the duration split rather than aggregating all rows under the parent benchmark.

Top models (higher is better)

ModelScore
GPT-5.6 Sol97.6
Claude Fable 592.9
Kimi K392.9
Claude Opus 590.5
GPT-5.6 Luna85.7
Sonnet 576.2
Opus 4.873.8
Grok 4.573.8
GLM-5.266.7
Opus 4.764.3
GPT-5.6 Terra64.3
Inkling57.1
GPT-5.3-Codex54.8
Composer 2.552.4
Gemini 3.5 Flash52.4
Muse Spark 1.152.4
Sonnet 4.650.0
Gemini 3.6 Flash50.0
GPT-5.450.0
GPT-5.550.0
Kimi K2.7 Code50.0
MiniMax M347.6
DeepSeek-V4-Pro45.2
GLM-5.145.2
Opus 4.642.9
Gemini 3.1 Pro Preview42.9
Gemini 3 Pro Preview42.9
Opus 4.540.5
Gemini 3.5 Flash-Lite40.5
Kimi K2.640.5
MiniMax M2.140.5
Muse Spark40.5
Gemini 3 Flash Preview38.1
GPT-5.238.1
Grok 4.338.1
MiniMax M2.538.1
Nemotron 3 Ultra 550B A55B38.1
Qwen3.6 Plus (2026-04-02)38.1
Qwen3.7-Max38.1
GLM-4.735.7
Loading Atlas data…