Atlas

Benchmarks

← All benchmarks

GDP.pdf (Artificial Analysis, All-pass)

Professional Work

Artificial Analysis' independent GDP.pdf implementation: 100 professional document tasks, five attempts per task, LiteParse text plus page images where supported, and GPT-5.6 Luna Medium judging. All-pass credits an attempt only if every atomic criterion passes; this harness differs from Surge's implementation.

Top models (higher is better)

ModelScore
GPT-6 Astra32.2
Claude Opus 5.528.8
Claude Fable 5.128.0
GPT-6 Sol28.0
GPT-5.6 Sol27.8
Muse Spark 1.326.6
GPT-5.6 Terra24.6
Claude Fable 524.0
GPT-5.6 Luna24.0
Gemini 3.7 Flash23.6
Grok 4.723.2
Opus 4.822.8
Gemini 3.8 Flash22.8
Qwen3.8 Max (0902)22.8
GPT-5.522.6
Kimi K322.0
Claude Opus 521.6
GPT-6 Luna20.4
GPT-5.5 Instant (2026-06-25 hosted snapshot)20.2
Qwen3.8-Max20.2
Gemini 3.5 Flash19.8
MiMo-V2.6-Pro19.2
Grok 4.518.8
Gemini 3.1 Pro Preview17.8
Grok 4.617.8
Gemini 3.6 Flash17.4
Muse Spark 1.217.4
Qwen3.8 27B16.6
Sonnet 4.615.8
Qwen3.8-Flash-Next15.6
GLM-5.3 Flash15.4
Qwen3.8 2.4T A95B15.0
Step 5 Preview14.8
Muse Spark 1.114.4
Gemini 3.5 Flash-Lite13.6
DeepSeek-V4-Pro13.4
GPT-5.4 Mini13.4
Sonnet 513.2
Kimi K2.613.0
DeepSeek V4.1 Flash12.8
Loading Atlas data…