Atlas

Benchmarks

← All benchmarks

SWE-Marathon

Code · 2026-06-05

SWE-Marathon evaluates agents on 20 executable, ultra-long-horizon tasks spanning software engineering and adjacent technical domains. Scores report the percentage of tasks solved.

Top models (higher is better)

ModelScore
Kimi K342.0
Opus 4.840.0
GPT-5.6 Sol39.0
Claude Fable 535.0
Grok 4.529.0
Opus 4.716.0
GPT-5.514.0
GLM-5.213.0
Gemini 3.5 Flash7.0
DeepSeek-V4-Pro4.0
Gemini 3.1 Pro Preview4.0
GLM-5.11.0
Kimi K2.60.0
MiniMax M2.70.0
Loading Atlas data…