Atlas

Benchmarks

← All benchmarks

FrontierCode 1.0 Extended - Score

Code · 2026-06-08

FrontierCode 1.0 Extended weighted score across the full set of 150 production-code tasks. Maintainer-authored rubrics cover correctness, regression safety, test quality, scope, and code quality; a trial that fails any blocker criterion receives a zero score.

Top models (higher is better)

ModelScore
Claude Fable 561.6
Opus 4.851.8
GPT-5.544.8
Opus 4.743.9
Kimi K2.7 Code39.9
GLM-5.239.0
Sonnet 538.8
Kimi K2.637.0
GPT-5.4 Mini36.0
Gemini 3.1 Pro Preview34.2
Sonnet 4.633.6
Kimi K2.522.7
Qwen3.6 Plus (2026-04-02)21.3
MiniMax M2.719.9
SWE-1.618.4
MiniMax M2.515.8
Gemini 3.1 Flash-Lite Preview14.6
Loading Atlas data…