DeepSWE v1.1 pass@4
Code · 2026-06-14
DeepSWE v1.1 pass@4 is the percentage of attempted tasks with at least one passing scored rollout among four nominal runs. Tasks with only excluded provider, verifier, or network errors do not enter the denominator.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-5.5 | 90.3 |
| GPT-5.6 Luna | 90.3 |
| Claude Opus 5 | 89.4 |
| Kimi K3 | 89.4 |
| Claude Fable 5 | 88.5 |
| GPT-5.6 Terra | 88.5 |
| GPT-5.6 Sol | 86.7 |
| Opus 4.8 | 80.5 |
| Sonnet 5 | 79.6 |
| Muse Spark 1.1 | 79.6 |
| GPT-5.4 | 77.9 |
| Grok 4.5 | 77.9 |
| GLM-5.2 | 77.0 |
| Gemini 3.6 Flash | 76.1 |
| Gemini 3.5 Flash | 66.4 |
| Kimi K2.7 Code | 61.1 |
| Sonnet 4.6 | 56.6 |
| Gemini 3.1 Pro Preview | 28.3 |