Atlas

Benchmarks

← All benchmarks

SWE-Bench Pro V2 — Hard

Code · 2026-09-22

The separately reported HARD subset of SWE-Bench Pro V2, using the revised public task set and locked evaluation protocol. The source does not state the subset's task count on the leaderboard.

Top models (higher is better)

ModelScore
Claude Opus 598.0
Claude Fable 5.192.2
GPT-6 Astra90.2
Sonnet 588.2
Kimi K388.2
GPT-5.6 Terra86.3
GLM-5.384.3
GPT-5.6 Sol82.4
Gemini 3.8 Flash58.8
Inkling56.9
Haiku 4.525.5
Loading Atlas data…