Atlas

Benchmarks

← All benchmarks

SWE-bench Verified - >4 Hours

Code · 2024-08-13

The >4 hours split of SWE-bench Verified. This child benchmark separates a source-reported subtask or subtrack from the parent aggregate so scores at different grains do not share one benchmark_slug.

Top models (higher is better)

ModelScore
Claude Fable 5100.0
Claude Opus 5100.0
GPT-5.6 Sol100.0
Opus 4.766.7
Opus 4.866.7
Sonnet 566.7
Composer 2.566.7
GLM-5.266.7
GPT-5.566.7
GPT-5.6 Luna66.7
Grok 4.566.7
Kimi K366.7
Muse Spark 1.166.7
Haiku 4.533.3
Opus 4.633.3
Sonnet 4.633.3
DeepSeek-V3.233.3
Gemini 3.1 Pro Preview33.3
Gemini 3.5 Flash33.3
Gemini 3.5 Flash-Lite33.3
Gemini 3.6 Flash33.3
Gemini 3 Flash Preview33.3
Gemini 3 Pro Preview33.3
GLM-4.733.3
GLM-5.133.3
GLM-533.3
GPT-5.2-Codex33.3
GPT-5.233.3
GPT-5.3-Codex33.3
GPT-5.4 Nano33.3
GPT-5.6 Terra33.3
GPT-533.3
GPT-5 Mini33.3
Grok 4.2033.3
Grok 4.333.3
Inkling33.3
Kimi K2.7 Code33.3
Kimi K2 Thinking33.3
MiMo-V2.5-Pro33.3
MiniMax M2.133.3
Loading Atlas data…