Atlas

Benchmarks

← All benchmarks

FrontierCode 1.0 Diamond - Score

Code · 2026-06-08

FrontierCode 1.0 Diamond weighted score on the original benchmark's 50 hardest production-code tasks. Maintainer-authored rubrics cover correctness, regression safety, test quality, scope, and code quality; a trial that fails any blocker criterion receives a zero score. Cognition deprecated Diamond in FrontierCode 1.1.

Top models (higher is better)

ModelScore
Opus 4.813.4
GPT-5.56.3
Opus 4.75.2
Gemini 3.1 Pro Preview4.7
GPT-5.4 Mini4.6
Kimi K2.63.8
Sonnet 4.63.5
SWE-1.62.5
MiniMax M2.72.4
MiniMax M2.51.1
Kimi K2.51.0
Gemini 3.1 Flash-Lite Preview0.7
Loading Atlas data…