Atlas

Benchmarks

← All benchmarks

FrontierCode 1.1 Main - Score

Code · 2026-07-07

FrontierCode 1.1 Main weighted score on the 100 hardest tasks in the 150-task Extended set. The 1.1 methodology adds fair-internet-use instructions and verification, and relaxes overly strict blocker criteria; flagged runs receive zero.

Top models (higher is better)

ModelScore
Claude Fable 553.5
Claude Opus 553.4
GPT-5.6 Sol47.5
Opus 4.846.5
GPT-5.543.0
Sonnet 542.7
Grok 4.542.4
SWE-1.742.3
GPT-5.6 Terra41.3
GPT-5.6 Luna39.8
Opus 4.738.5
Kimi K2.7 Code30.1
Composer 2.525.6
GLM-5.224.5
DeepSeek-V4-Pro17.6
MiniMax M314.7
Inkling14.0
Qwen3.7-Plus10.2
SWE-1.69.4
Loading Atlas data…