DeepSWE v1 pass@4
Code · 2026-05-26
DeepSWE v1 pass@4 is the percentage of attempted tasks with at least one passing scored rollout among four nominal runs. Tasks with only excluded provider, verifier, or network errors do not enter the denominator.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 88.3 |
| Opus 4.7 | 85.8 |
| Opus 4.8 | 80.5 |
| GPT-5.4 | 77.0 |
| GLM-5.2 | 69.9 |
| Sonnet 4.6 | 61.9 |
| Gemini 3.5 Flash | 56.6 |
| Opus 4.6 | 50.4 |
| Kimi K2.6 | 48.7 |
| MiniMax M3 | 48.7 |
| GPT-5.4 Mini | 46.0 |
| MiMo-V2.5-Pro | 45.1 |
| Qwen3.7-Max | 41.6 |
| GLM-5.1 | 38.9 |
| Grok Build 0.1 | 29.2 |
| Gemini 3.1 Pro Preview | 24.8 |
| DeepSeek-V4-Pro | 18.6 |
| Gemini 3 Flash Preview | 15.0 |
| Qwen3.6 Plus (2026-04-02) | 9.7 |
| Haiku 4.5 | 0.9 |
| MiniMax M2.7 | 0.9 |