SWE-Bench Pro V2 — Full
Code · 2026-09-22
Scale AI and Reflection's SWE-Bench Pro V2 public split contains 642 tasks across 11 repositories after dropping 89 invalid tasks and correcting contradictory instructions. Agent network access is restricted to model endpoints, web tools are disabled, and submitted patches are regraded on pristine images. Resolve rate requires both fixing the issue and passing regression tests.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 99.4 |
| Claude Fable 5.1 | 99.1 |
| Kimi K3 | 97.7 |
| GPT-6 Astra | 96.9 |
| GLM-5.3 | 95.6 |
| GPT-5.6 Sol | 95.5 |
| Gemini 3.8 Flash | 94.9 |
| Sonnet 5 | 93.2 |
| GPT-5.6 Terra | 92.4 |
| Inkling | 89.9 |