SWE-Marathon
Code · 2026-06-05
SWE-Marathon evaluates agents on 20 executable, ultra-long-horizon tasks spanning software engineering and adjacent technical domains. Scores report the percentage of tasks solved.
Top models (higher is better)
| Model | Score |
|---|---|
| Kimi K3 | 42.0 |
| Opus 4.8 | 40.0 |
| GPT-5.6 Sol | 39.0 |
| Claude Fable 5 | 35.0 |
| Grok 4.5 | 29.0 |
| Opus 4.7 | 16.0 |
| GPT-5.5 | 14.0 |
| GLM-5.2 | 13.0 |
| Gemini 3.5 Flash | 7.0 |
| DeepSeek-V4-Pro | 4.0 |
| Gemini 3.1 Pro Preview | 4.0 |
| GLM-5.1 | 1.0 |
| Kimi K2.6 | 0.0 |
| MiniMax M2.7 | 0.0 |