Atlas

Benchmarks

← All benchmarks

FrontierCode 1.1 Extended - Score

Code · 2026-07-07

FrontierCode 1.1 Extended weighted score across the full set of 150 production-code tasks. The 1.1 methodology adds fair-internet-use instructions and verification, and relaxes overly strict blocker criteria; flagged runs receive zero.

Top models (higher is better)

ModelScore
Claude Fable 564.9
Claude Opus 563.6
GPT-5.6 Sol60.6
Opus 4.859.6
GPT-5.556.7
Grok 4.556.6
Sonnet 556.2
GPT-5.6 Terra55.8
GPT-5.6 Luna55.1
SWE-1.754.6
Opus 4.753.9
Kimi K2.7 Code45.4
Composer 2.540.8
GLM-5.240.0
DeepSeek-V4-Pro31.0
MiniMax M328.5
Inkling24.8
Qwen3.7-Plus21.8
SWE-1.620.5
Loading Atlas data…