DeepSWE v1.1
Code · 2026-06-14
DeepSWE v1.1 evaluates frontier coding agents on the same 113 original, long-horizon software-engineering tasks as v1, with committed patches graded in clean verifier containers and structured per-test reports. Pass@1 is the pooled pass rate over scored rollout attempts; context-window failures and agent timeouts count as failures, while provider, verifier, and network errors are excluded.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 73.6 |
| GPT-5.6 Sol | 72.7 |
| Claude Fable 5 | 70.0 |
| GPT-5.6 Terra | 69.6 |
| Kimi K3 | 69.0 |
| GPT-5.6 Luna | 67.2 |
| GPT-5.5 | 67.0 |
| Opus 4.8 | 59.0 |
| Sonnet 5 | 54.0 |
| Grok 4.5 | 54.0 |
| Muse Spark 1.1 | 53.3 |
| GPT-5.4 | 51.8 |
| GLM-5.2 | 46.2 |
| Laguna S 2.1 | 40.4 |
| Kimi K2.7 Code | 30.5 |
| Sonnet 4.6 | 29.9 |
| Gemini 3.1 Pro Preview | 12.0 |
| DeepSeek-V4-Pro | 9.0 |