Atlas

Benchmarks

← All benchmarks

FrontierSWE V2 Worst@5

Code

FrontierSWE V2 mean of the worst normalized reward per task across up to five graded trials. This is not a confidence-interval bound.

Top models (higher is better)

ModelScore
GPT-6 Astra55.1
Claude Opus 5.552.0
Claude Opus 539.1
GPT-5.6 Sol22.2
GLM-5.317.7
Grok 4.715.9
Kimi K313.9
Grok 4.612.9
Gemini 3.7 Flash10.3
Gemini 3.8 Flash9.5
Muse Spark 1.26.8
Inkling0.1
Loading Atlas data…