GDP.pdf (Artificial Analysis, All-pass)
Professional Work
Artificial Analysis' independent GDP.pdf implementation: 100 professional document tasks, five attempts per task, LiteParse text plus page images where supported, and GPT-5.6 Luna Medium judging. All-pass credits an attempt only if every atomic criterion passes; this harness differs from Surge's implementation.
Top models (higher is better)
| Model | Score |
|---|---|
| GPT-6 Astra | 32.2 |
| Claude Opus 5.5 | 28.8 |
| Claude Fable 5.1 | 28.0 |
| GPT-6 Sol | 28.0 |
| GPT-5.6 Sol | 27.8 |
| Muse Spark 1.3 | 26.6 |
| GPT-5.6 Terra | 24.6 |
| Claude Fable 5 | 24.0 |
| GPT-5.6 Luna | 24.0 |
| Gemini 3.7 Flash | 23.6 |
| Grok 4.7 | 23.2 |
| Opus 4.8 | 22.8 |
| Gemini 3.8 Flash | 22.8 |
| Qwen3.8 Max (0902) | 22.8 |
| GPT-5.5 | 22.6 |
| Kimi K3 | 22.0 |
| Claude Opus 5 | 21.6 |
| GPT-6 Luna | 20.4 |
| GPT-5.5 Instant (2026-06-25 hosted snapshot) | 20.2 |
| Qwen3.8-Max | 20.2 |
| Gemini 3.5 Flash | 19.8 |
| MiMo-V2.6-Pro | 19.2 |
| Grok 4.5 | 18.8 |
| Gemini 3.1 Pro Preview | 17.8 |
| Grok 4.6 | 17.8 |
| Gemini 3.6 Flash | 17.4 |
| Muse Spark 1.2 | 17.4 |
| Qwen3.8 27B | 16.6 |
| Sonnet 4.6 | 15.8 |
| Qwen3.8-Flash-Next | 15.6 |
| GLM-5.3 Flash | 15.4 |
| Qwen3.8 2.4T A95B | 15.0 |
| Step 5 Preview | 14.8 |
| Muse Spark 1.1 | 14.4 |
| Gemini 3.5 Flash-Lite | 13.6 |
| DeepSeek-V4-Pro | 13.4 |
| GPT-5.4 Mini | 13.4 |
| Sonnet 5 | 13.2 |
| Kimi K2.6 | 13.0 |
| DeepSeek V4.1 Flash | 12.8 |