DeepSWE v1 (Best of up to 3 attempts)
Code · 2026-05-26
This DeepSWE v1 aggregation reports the best whole-benchmark score from up to three attempts on its 113 long-horizon software-engineering tasks. The aggregation qualifier is kept separate from the benchmark's standard pooled pass@1 and pass@4 metrics.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 70.0 |
| Macaron-V1-Venti | 58.4 |
| Opus 4.8 | 58.0 |
| GLM-5.2 | 54.9 |
| MiniMax M3 | 20.0 |
| Qwen3.7-Max | 18.0 |
| Gemini 3.1 Pro Preview | 10.0 |