Atlas

Benchmarks

← All benchmarks

SWE-bench Verified (Vals 15 min–1 hour split)

Code · 2026-06-17

Vals difficulty split containing SWE-bench Verified tasks estimated to take 15 minutes to 1 hour for humans. It preserves the duration split rather than aggregating all rows under the parent benchmark.

Top models (higher is better)

ModelScore
Claude Opus 596.9
GPT-5.6 Sol95.0
Claude Fable 594.6
Kimi K394.6
GPT-5.6 Luna92.0
Opus 4.888.1
Grok 4.583.5
Muse Spark 1.182.4
GLM-5.281.6
GPT-5.581.2
Gemini 3.6 Flash79.3
Opus 4.778.9
Composer 2.578.9
Gemini 3.1 Pro Preview77.8
Gemini 3.5 Flash77.8
Sonnet 577.4
GPT-5.476.2
DeepSeek-V4-Pro75.9
Opus 4.675.5
GPT-5.6 Terra75.5
Kimi K2.7 Code75.5
Sonnet 4.675.1
Inkling75.1
Opus 4.574.7
MiMo-V2.5-Pro74.7
Kimi K2.674.3
Gemini 3 Pro Preview73.6
Gemini 3.5 Flash-Lite73.2
GPT-5.3-Codex73.2
GLM-5.172.8
MiniMax M2.172.8
MiniMax M372.8
Muse Spark72.8
GPT-5.272.4
Gemini 3 Flash Preview72.0
MiniMax M2.772.0
GPT-5.4 Mini70.9
MiniMax M2.570.9
Qwen3.6 Plus (2026-04-02)70.9
MiMo-V2.570.5
Loading Atlas data…