Atlas

Benchmarks

← All benchmarks

FrontierSWE V2

Code

34 ultra-long-horizon coding, scientific computing and research tasks with normalized rewards from 0 to 1, averaged and expressed as percentages. Proximus harness, 20-hour task budget and five planned trials. V2 revises task QA and scoring relative to V1; incomplete graded trial counts are retained in observation notes.

Top models (higher is better)

ModelScore
GPT-6 Astra65.5
Claude Opus 5.562.3
Claude Opus 552.0
GPT-5.6 Sol32.2
GLM-5.330.2
Grok 4.729.5
Kimi K325.9
Grok 4.625.3
Gemini 3.7 Flash20.3
Gemini 3.8 Flash19.6
Muse Spark 1.212.0
Inkling4.1
Loading Atlas data…