Atlas

Benchmarks

← All benchmarks

GDPval-AA v2 (normalized)

Professional Work · 2026-04-18

GDPval-AA v2 is a normalized benchmark for real-world professional tasks across 44 occupations and 9 industries. Models use shell and web access to produce documents, slides, diagrams, and spreadsheets; higher scores mean better judged task performance.

Top models (higher is better)

ModelScore
Claude Opus 567.9
Claude Fable 562.2
GPT-5.6 Sol61.6
Kimi K359.4
Sonnet 555.0
Opus 4.854.5
GPT-5.6 Terra54.1
GPT-5.6 Luna54.1
DeepSeek-V4-Flash-073152.9
Grok 4.551.4
GLM-5.250.5
Opus 4.749.5
GPT-5.549.5
Gemini 3.6 Flash46.2
GPT-5.444.6
MiniMax M344.5
Sonnet 4.643.8
Muse Spark 1.143.8
Gemini 3.5 Flash42.1
DeepSeek-V4-Pro40.2
Qwen3.7-Max38.5
Inkling-Small38.4
JT-4.1 Flash 236B A21B38.3
MiMo-V2.5-Pro38.2
Motif-3-Beta37.9
GLM-5.137.8
Nex-N2-Pro37.5
Inkling36.8
Hy335.8
Grok Build 0.135.7
DeepSeek-V4-Flash34.5
Kimi K2.7 Code34.4
Kimi K2.634.4
Agnes 2.5 Pro Alpha33.6
GPT-5.4 Mini33.5
GLM-4.733.2
Nemotron 3 Ultra 550B A55B33.1
MiniMax M2.732.9
MiMo-V2.532.2
Muse Spark32.2
Loading Atlas data…