Atlas

Benchmarks

← All benchmarks

Capability

Indexes · 2026-08-03

Synthetic general capability score from the Atlas factor model. Higher scores indicate stronger estimated overall benchmark performance.

Top models (higher is better)

ModelScore
Claude Mythos Preview166.3
GPT-5.6 Sol Pro164.9
Claude Mythos 5164.2
Claude Fable 5164.2
Claude Opus 5163.8
GPT-5.5 Pro162.9
GPT-5.6 Terra Pro162.8
GPT-5.6 Sol162.6
Gemini 3 Deep Think161.4
GPT-5.6 Luna Pro161.3
GPT-5.4 Pro161.2
GPT-5.6 Terra160.9
Kimi K3160.4
GPT-5.5160.3
Opus 4.8160.2
GPT-5.6 Luna159.0
Grok 4.5158.9
DeepSeek-V4-Flash-0731158.6
Sonnet 5158.4
Opus 4.7158.4
Muse Spark 1.1158.0
GPT-5.4157.7
Opus 4.6156.8
Gemini 3.5 Flash156.8
GPT-5.3-Codex156.7
GPT-5.2 Pro156.6
GLM-5.2156.5
Gemini 3.6 Flash156.2
Gemini 3.1 Pro Preview155.1
Laguna S 2.1155.1
Gemini 2.5 Deep Think154.8
Sonnet 4.6154.6
Muse Spark154.6
Hy3154.5
Motif-3-Beta154.4
Composer 2.5154.4
GPT-5.2154.2
Nex-N2-Pro154.2
Qwen3.7-Max154.1
GPT-5.2-Codex153.8
Loading Atlas data…