SWE-bench Verified Anthropic-Compatible 489-Task Subset
Code · 2025-02-24
The 489-task subset of SWE-bench Verified whose golden solutions passed on Anthropic's internal infrastructure for the Claude 3.7 Sonnet evaluation. Eleven tasks from the 500-task parent benchmark are excluded, so scores on this subset are not directly comparable to full-set results.
Top models (higher is better)
| Model | Score |
|---|---|
| Claude Mythos Preview | 89.6 |
| Opus 4.8 | 87.1 |
| GPT-5.5 | 81.0 |
| Opus 4.6 | 79.0 |
| GLM-5.2 | 75.3 |
| GPT-5.4 Mini | 73.0 |
| Opus 4 | 66.7 |
| GPT-5 | 63.0 |
| DeepSeek-V3.1 | 54.8 |
| DeepSeek-R1-0528 | 44.6 |
| gpt-oss-120b | 42.6 |
| DeepSeek-R1 | 25.4 |