SWE-Bench Pro V2 — Hard
Code · 2026-09-22
The separately reported HARD subset of SWE-Bench Pro V2, using the revised public task set and locked evaluation protocol. The source does not state the subset's task count on the leaderboard.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Opus 5 | 98.0 |
| Claude Fable 5.1 | 92.2 |
| GPT-6 Astra | 90.2 |
| Sonnet 5 | 88.2 |
| Kimi K3 | 88.2 |
| GPT-5.6 Terra | 86.3 |
| GLM-5.3 | 84.3 |
| GPT-5.6 Sol | 82.4 |
| Gemini 3.8 Flash | 58.8 |
| Inkling | 56.9 |
| Haiku 4.5 | 25.5 |