Atlas

Benchmarks

← All benchmarks

SWE-bench Verified Anthropic-Compatible 489-Task Subset

Code · 2025-02-24

The 489-task subset of SWE-bench Verified whose golden solutions passed on Anthropic's internal infrastructure for the Claude 3.7 Sonnet evaluation. Eleven tasks from the 500-task parent benchmark are excluded, so scores on this subset are not directly comparable to full-set results.

Top models (higher is better)

ModelScore
Claude Mythos Preview89.6
Opus 4.887.1
GPT-5.581.0
Opus 4.679.0
GLM-5.275.3
GPT-5.4 Mini73.0
Opus 466.7
GPT-563.0
DeepSeek-V3.154.8
DeepSeek-R1-052844.6
gpt-oss-120b42.6
DeepSeek-R125.4
Loading Atlas data…