Atlas

Benchmarks

← All benchmarks

FrontierSWE V2 Best@5

Code

FrontierSWE V2 mean of the best normalized reward per task across up to five graded trials. This is not a confidence-interval bound.

Top models (higher is better)

ModelScore
GPT-6 Astra73.0
Claude Opus 5.571.5
Claude Opus 560.7
GPT-5.6 Sol44.2
Grok 4.741.1
GLM-5.340.6
Kimi K337.6
Grok 4.637.3
Gemini 3.8 Flash31.6
Gemini 3.7 Flash31.2
Muse Spark 1.218.4
Inkling10.8
Loading Atlas data…