Atlas

Benchmarks

← All benchmarks

GDPval-AA v2.1

Professional Work

Artificial Analysis GDPval-AA v2.1, evaluating 220 professional tasks with shell and web access through Stirrup. Blind pairwise judgments produce Elo ratings anchored to DeepSeek V4.1 Flash (max) at 1600. This version is kept separate from GDPval-AA v2's earlier judging and rating scale.

Top models (higher is better)

ModelScore
Claude Opus 5.51846
Claude Fable 5.11735
Claude Opus 51708
Grok 4.71695
Muse Spark 1.31674
MiMo-V2.6-Pro1673
Qwen3.8 Max (0902)1668
GLM-5.31646
GLM-5.3 Flash1641
Grok 4.61632
Qwen3.8-Flash-Next1612
DeepSeek V4.1 Flash1600
Qwen3.8 2.4T A95B1597
Qwen3.8-Max1596
Claude Fable 51595
GPT-5.6 Sol1588
Step 5 Preview1566
Agnes 3.0 Flash (hosted)1550
GPT-6 Astra1542
DeepSeek V4 Flash Vision Exp1534
Kimi K31524
GPT-6 Sol1487
Muse Spark 1.21482
Sonnet 51449
GPT-5.6 Luna1443
DeepSeek V4 Pro 08131441
Opus 4.81438
GPT-5.6 Terra1432
DeepSeek-V4-Flash-07311427
Gemini 3.8 Flash1412
Qwen3.8 27B1409
Gemini 3.7 Flash1371
Grok 4.51370
GPT-6 Luna1367
GLM-5.21358
K2 Horizon 375B A23B1349
Opus 4.71338
GPT-5.51336
Agnes 2.5 Pro Beta1311
Gemini 3.6 Flash1265
Loading Atlas data…